k3f_exp_ckpts β DSpark draft checkpoints from the GB300 (g3*) experiments
Final-iteration draft weights for six single-variable arms trained against
moonshotai/Kimi-K3-Flash on the GB300 cluster. Staged for batch export +
evaluation on another cluster: each folder holds the model/ shards, the run's own
config, and the draft architecture JSON.
optimizer/ state is deliberately excluded β it is ~80% of the bytes and only needed
to resume training.
Every arm's
model/needs all 8 shards (__0_0.distcpβ¦__7_0.distcpplus.metadata, which names all eight). The 2Γ4 trainer wrote ranks 0-3 and ranks 4-7 to different nodes, so an assembly that reads only one node yields a 4-of-8 set that cannot be loaded at all. Verify the count, not just thatmodel/exists β an earlier revision of this repo shipped four arms with half their shards and passed an existence check.
The arms, and what each one changes
The chain is single-variable: g3d β g3e β g3f β {g3g, g3h, g3j}. Each arm differs from its parent by one thing, so adjacent arms are differenceable.
g3j is the one arm that varies initialisation rather than architecture: it has g3f's
draft JSON byte-for-byte but starts from g3e's converged weights instead of scratch,
so g3f vs g3j isolates warm-start. Note it loads block-5 weights and then trains
at block 3 β the load was clean (no shape warnings in the trainer log), but serve it
with num_speculative_tokens 3, not 5.
| arm | final iter | fc_norm | block | taps | engines | changes vs parent |
|---|---|---|---|---|---|---|
g3c |
6001 | false | 3 | [1,29,57] | 16 | all-turn corpus (not single-variable vs g3d) |
g3d |
651 | true | 5 | [1,29,57] | 16 | baseline β v3d recipe on GB300 |
g3e |
4581 | false | 5 | [1,29,57] | 16 | fc_norm trueβfalse |
g3f |
4581 | false | 3 | [1,29,57] | 16 | block_size 5β3 |
g3g |
4581 | false | 3 | [46,57,59] | 16 | tap placement |
g3h |
2751 | false | 3 | [1,29,57] | 12 | 4β3 producer engines |
g3j |
4581 | false | 3 | [1,29,57] | 16 | warm-start from g3e iter 4581 (continual_training) |
Shared by every arm: 3-layer MLA draft, enable_confidence_head: true,
markov_head_type: vanilla, target depth 61, max_seq_length: 65536,
last_turn_loss_only: true, v3 blended corpus (except g3c).
Read these before comparing numbers
- g3e, g3f, g3g and g3j all completed the same 4580-step schedule and are step-aligned with each other.
- g3d stopped at iteration 651. It is an early checkpoint, not a converged model β comparing its acceptance against the 4581-step arms at face value is wrong.
- g3h stopped at 2751, also short of the full schedule.
- g3h is not a modelling change. Same draft JSON path as g3f; only producer capacity differs (4 engines β 3). It answers "was g3f producer-bound", nothing about architecture.
g3cchanges the corpus, so it is not on the single-variable chain.- Step noise on these runs is 5-9%; single-step reads are not resolvable. Use windowed medians.
Using a checkpoint
These are TorchSpec distcp shards, not an HF model folder. Convert first
(arthur/scripts/export_k3flash_dspark_vllm.sh, wrapping TorchSpec
tools/convert_to_hf.py --vllm), then pass the result as the draft in vLLM's
--speculative-config.
Two gates that fail silently if skipped:
num_speculative_tokensmust equal the arm'sblock_size(5 for g3d/g3e, 3 for the rest). A mismatch drafts a different width and reads as low acceptance rather than as an error.- The target must be 61 layers deep. Against a different depth, vLLM rescales the aux capture ids and logs only a WARNING, after which the draft reads layers it never trained on β again indistinguishable from a weak draft.
Serving K3-Flash on agent traffic additionally needs
--enable-auto-tool-choice --tool-call-parser kimi_k3; without them every
tools-carrying request 400s before inference and the run looks like spec decode is
off.