KAT-Coder-V2.5 β€” CMF coding specialist (expert-defragmented MoE)

A 34.7B-A3B MoE coder in a single 12.7 GB file β€” 35% smaller than the full quantized model, Γ—1.8 faster on a 24 GB MacBook, +2.8% code perplexity. This is Kwaipilot/KAT-Coder-V2.5-Dev (Qwen3.6-35B-A3B architecture: 40 layers, 30 of them GatedDeltaNet linear attention, 256 routed experts top-8 + a shared expert), quantized to 4-bit tiles and then physically stripped of the experts that code generation never routes to.

file size held-out code ppl decode, M4 24 GB prefill
full q4t model 19.6 GB 5.058 7.6 tok/s 6.1 tok/s
this file 12.7 GB 5.198 (+2.8%) 13.7 tok/s (Γ—1.8) 20.0 tok/s (Γ—3.3)

The speedup is not a kernel trick: the full model does not fit a 24 GB machine and pages from disk on every token, while the specialist resides in memory entirely. On machines with plenty of RAM the two decode at similar speed and you simply save the 7 GB.

How to use

CMF is a single-file LLM format with a small pure-Rust runtime β€” no torch, no CUDA install, no Python. Install the CLI (one command, GPU backends included) and run:

cargo install cortiq-cli   # or a release binary: github.com/infosave2007/cmf/releases

cortiq run KAT-Coder-V2.5-CMF.cmf \
    --prompt "Write a Python function that checks if a number is prime." --max-tokens 300

cortiq serve KAT-Coder-V2.5-CMF.cmf --port 8080   # OpenAI-compatible API + dashboard
cortiq bench KAT-Coder-V2.5-CMF.cmf               # measure on your hardware

The tokenizer and chat template are embedded in the file β€” run is a real chat turn out of the box.

GPU: CMF_GPU=1 enables the GPU path. On discrete Vulkan/DX12 cards the entire decode β€” GatedDeltaNet recurrence, attention, the MoE router, the on-device top-k expert selection and every selected expert β€” executes as one GPU submit per token (the full-model variant of this pipeline decodes at 32.8 tok/s on an RTX 5090 vs 14.4 on its 32-core host CPU). On Apple silicon a runtime probe arbitrates Metal against CPU per operation and keeps whichever wins. Long contexts: --o1 all converts the 10 softmax-attention layers into a constant-memory streaming operator (KV+state at 4K context: 238 β†’ 83 MB).

The technology

MoE expert usage turns out to be strongly task-conditional. Measured on the full KAT-Coder: the top-64 expert sets selected for code vs for natural-language prose overlap with a Jaccard index of just 0.25 β€” near-disjoint working sets β€” and a code-derived expert mask captures only ~39% of prose routing mass. A model serving one task therefore carries hundreds of experts it never routes to.

The pipeline that produced this file (two commands, reproducible with cortiq β‰₯ 0.5.27):

  1. Record the routing field. A teacher-forced pass over a representative code corpus with CMF_MOE_STATS=stats.json records per-layer expert-selection frequencies β€” an empirical routing field over the expert lattice (the "B-field" of the CMF patent family's claim 12).
  2. Drop what the task never uses. cortiq moe-defrag model.cmf --stats stats.json --cover 0.95 keeps, per layer, the smallest top set of experts reaching 95% of the recorded routing mass (mean 158 of 256 per layer here; 3,906 experts dropped in total), renumbers the kept experts into a contiguous prefix, slices the router's rows to match, and writes a standard .cmf. At inference the router's softmax renormalizes over the kept set β€” no retraining, weights byte-identical to the quantized originals.

The same restriction can be previewed at runtime without rewriting the file (CMF_MOE_MASK=stats.json CMF_MOE_MASK_COVER=0.95), which is how the perplexity gate above was measured before the physical cut β€” the two are mathematically identical.

Scope, honestly. This is a code specialist: calibration was Rust-heavy source code, and quality outside the calibrated task degrades by design (prose routes to experts that are no longer there). For general-purpose use, take the full model through the step-by-step guide β€” it covers GGUF β†’ CMF conversion, Vulkan and Metal, and carving your own specialist for any task with your own corpus.

Provenance

Ecosystem

  • Engine / converter / format spec: https://github.com/infosave2007/cmf β€” the pure-Rust runtime this file targets (cargo install cortiq-cli).
  • CMF Mobile β€” a Flutter app on the same runtime: local on-device chat, or the phone as an OpenAI-compatible server. This 12.7 GB specialist wants a desktop's RAM; on phones pick a smaller CMF build (e.g. Bonsai-1.7B or Nanbeige 4.2 3B).
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for infosave/KAT-Coder-V2.5-CMF

Finetuned
(2)
this model