clef-flash-MLX-8bit

MLX conversion of Cloudflare/clef-flash for mlx-vlm. Clef is a decision model: it returns a probability for every option of every question in one forward pass and does not generate text.

Clef support is not yet in an mlx-vlm release:

pip install "git+https://github.com/Lazarus-931/mlx-vlm.git@feat/clef"
from mlx_vlm import load, predict

model, processor = load("nativ-community/clef-flash-MLX-8bit")
result = predict(model, processor, "Please refund my duplicate charge", {
    "department": {
        "type": "choice",
        "instructions": "Which team should handle this ticket?",
        "criteria": ["billing", "technical", "sales"],
    },
})
print(result["answers"]["department"]["value"])
Field Value
Source Cloudflare/clef-flash
Source revision 17f0b0ad64efb65d273590632833508766b2aae6
Quantization affine 8-bit, group size 64
mlx-vlm Lazarus-931/mlx-vlm@c16f81aa
Verification Identical token ids on 11/11 reference records (text, image, video, image+video, two images, max_pixels, fps, num_frames); same answer as Cloudflare's torch reference (fp32) on 25/25 questions; largest probability gap 0.0321
Downloads last month
37
Safetensors
Model size
10B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for nativ-community/clef-flash-MLX-8bit

Finetuned
Qwen/Qwen3.5-9B
Quantized
(53)
this model