k3f_exp_ckpts β€” DSpark draft checkpoints from the GB300 (g3*) experiments

Final-iteration draft weights for six single-variable arms trained against moonshotai/Kimi-K3-Flash on the GB300 cluster. Staged for batch export + evaluation on another cluster: each folder holds the model/ shards, the run's own config, and the draft architecture JSON.

optimizer/ state is deliberately excluded β€” it is ~80% of the bytes and only needed to resume training.

Every arm's model/ needs all 8 shards (__0_0.distcp … __7_0.distcp plus .metadata, which names all eight). The 2Γ—4 trainer wrote ranks 0-3 and ranks 4-7 to different nodes, so an assembly that reads only one node yields a 4-of-8 set that cannot be loaded at all. Verify the count, not just that model/ exists β€” an earlier revision of this repo shipped four arms with half their shards and passed an existence check.

The arms, and what each one changes

The chain is single-variable: g3d β†’ g3e β†’ g3f β†’ {g3g, g3h, g3j}. Each arm differs from its parent by one thing, so adjacent arms are differenceable.

g3j is the one arm that varies initialisation rather than architecture: it has g3f's draft JSON byte-for-byte but starts from g3e's converged weights instead of scratch, so g3f vs g3j isolates warm-start. Note it loads block-5 weights and then trains at block 3 β€” the load was clean (no shape warnings in the trainer log), but serve it with num_speculative_tokens 3, not 5.

arm final iter fc_norm block taps engines changes vs parent
g3c 6001 false 3 [1,29,57] 16 all-turn corpus (not single-variable vs g3d)
g3d 651 true 5 [1,29,57] 16 baseline β€” v3d recipe on GB300
g3e 4581 false 5 [1,29,57] 16 fc_norm true→false
g3f 4581 false 3 [1,29,57] 16 block_size 5β†’3
g3g 4581 false 3 [46,57,59] 16 tap placement
g3h 2751 false 3 [1,29,57] 12 4β†’3 producer engines
g3j 4581 false 3 [1,29,57] 16 warm-start from g3e iter 4581 (continual_training)

Shared by every arm: 3-layer MLA draft, enable_confidence_head: true, markov_head_type: vanilla, target depth 61, max_seq_length: 65536, last_turn_loss_only: true, v3 blended corpus (except g3c).

Read these before comparing numbers

  • g3e, g3f, g3g and g3j all completed the same 4580-step schedule and are step-aligned with each other.
  • g3d stopped at iteration 651. It is an early checkpoint, not a converged model β€” comparing its acceptance against the 4581-step arms at face value is wrong.
  • g3h stopped at 2751, also short of the full schedule.
  • g3h is not a modelling change. Same draft JSON path as g3f; only producer capacity differs (4 engines β†’ 3). It answers "was g3f producer-bound", nothing about architecture.
  • g3c changes the corpus, so it is not on the single-variable chain.
  • Step noise on these runs is 5-9%; single-step reads are not resolvable. Use windowed medians.

Using a checkpoint

These are TorchSpec distcp shards, not an HF model folder. Convert first (arthur/scripts/export_k3flash_dspark_vllm.sh, wrapping TorchSpec tools/convert_to_hf.py --vllm), then pass the result as the draft in vLLM's --speculative-config.

Two gates that fail silently if skipped:

  1. num_speculative_tokens must equal the arm's block_size (5 for g3d/g3e, 3 for the rest). A mismatch drafts a different width and reads as low acceptance rather than as an error.
  2. The target must be 61 layers deep. Against a different depth, vLLM rescales the aux capture ids and logs only a WARNING, after which the draft reads layers it never trained on β€” again indistinguishable from a weak draft.

Serving K3-Flash on agent traffic additionally needs --enable-auto-tool-choice --tool-call-parser kimi_k3; without them every tools-carrying request 400s before inference and the run looks like spec decode is off.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support