Home › Evidence › Records › con-0009
⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.

con-0009

open
citable URL: https://halobench.com/records/con-0009/ — this address never moves; the anchor /records/#con-0009 keeps resolving
kind report · upstream Nathanw1014/strix-halo-llamacpp
opened 2026-08-22 · https://github.com/Nathanw1014/strix-halo-llamacpp/releases/tag/v0.6.10
Nathanw1014/strix-halo-llamacpp (the Strix Halo Vulkan fork tracked in con-0002) shipped STABLE v0.6.10 (2026-08-22T12:03:16Z) that REVERSES the v0.6.8/v0.6.9 MTP rollback-EXACTNESS tradeoff recorded in con-0007/clm-0105. Token-exact MTP rollback returns on hybrid GDN (qwen35moe) targets (e.g. Qwen3.6-35B-A3B). Root cause of the v0.6.9 deadlock was NOT the state-save path: the server was RE-VERIFYING replayed draft tokens after a checkpoint restore, which livelocked the slot. Fix lands in two commits — "server: do not re-verify replayed draft tokens after a checkpoint restore" (9c5d899) and re-apply "common: use full checkpoints for MTP rollback" (f25eefe). With the re-verify eliminated, MTP rollback through full sequence-state checkpoints is token-exact again and the long-run stall is gone (smoke: Qwen3.6-35B-A3B MTP Q6_K 800-token repro 13.9 s no stall; 3000-token run 46.4 s / ~65 tok/s past the former stall horizon). v0.6.10 ALSO adds DSpark speculative-decode support for bailingmoe3 (Ling 3.0), cherry-picked from upstream llama.cpp PR #27508 (merged 2026-08-22T09:19:49Z, btw616). Runtime side is complete; the reworked bailingmoe3 forward pass was A/B verified byte-identical to the prior tip on Ling-3.0-tiny (greedy, 96 tokens), but the DSpark path is NOT exercisable end-to-end until Ling-3.0-flash draft GGUFs are published (inclusionAI/Ling-3.0-flash-dspark has model.safetensors up, GGUFs not out as of 2026-08-22T16:19Z). Also gates the RADV coopmat LDS pad of 2 to RADV >= 25.3 (older drivers violated VUID-08986). Payload built from Nathanw1014/llama.cpp@2586f6ed (branch strix-halo-vulkan). Speed figures are single runs, explicitly NOT the BENCHMARKS.md protocol; a pre-existing CPU-only divergence in one 800-token exactness matrix cell (token 776) predates v0.6.8 and is unrelated to the rollback path. Treat as an upstream/fork release-note record only: single-payload fork-vendor verification on gfx1151 Vulkan, not a HaloBench benchmark result, and not inherited into any HaloBench number.

Cited by — computed at build time, never stored

claims clm-0105 clm-0107
candidate gate history qwen36-35b
docs benchmark-protocol