clef-MLX-8bit

MLX conversion of Cloudflare/clef for mlx-vlm. Clef is a decision model: it returns a probability for every option of every question in one forward pass and does not generate text.

Clef support is not yet in an mlx-vlm release:

pip install "git+https://github.com/Lazarus-931/mlx-vlm.git@feat/clef"
from mlx_vlm import load, predict

model, processor = load("nativ-community/clef-MLX-8bit")
result = predict(model, processor, "Please refund my duplicate charge", {
    "department": {
        "type": "choice",
        "instructions": "Which team should handle this ticket?",
        "criteria": ["billing", "technical", "sales"],
    },
})
print(result["answers"]["department"]["value"])
Field Value
Source Cloudflare/clef
Source revision ed3eed331870db2eff4b0db01237128ede8a00ce
Quantization affine 8-bit, group size 64
mlx-vlm Lazarus-931/mlx-vlm@c16f81aa
Verification Identical token ids on 11/11 reference records (text, image, video, image+video, two images, max_pixels, fps, num_frames); same answer as Cloudflare's torch reference (bf16) on 25/25 questions; largest probability gap 0.0349
Downloads last month
25
Safetensors
Model size
27B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for nativ-community/clef-MLX-8bit

Base model

Qwen/Qwen3.8-27B
Finetuned
Cloudflare/clef
Quantized
(25)
this model