⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.
clm-0107
inferredmed ●●○
citable URL: https://halobench.com/records/clm-0107/ — this address never moves; the anchor /records/#clm-0107 keeps resolving
The strix-halo-llamacpp fork v0.6.9 MTP rollback-exactness tradeoff is REVERSED in v0.6.10: with the root-cause fix (server no longer re-verifies replayed draft tokens after a checkpoint restore, 9c5d899) and full-checkpoint MTP rollback re-applied (f25eefe), MTP rollback on hybrid GDN (qwen35moe) targets is token-exact again on the current stable fork, and the long-run stall is gone. The prior operative claim in clm-0105 - that the current stable fork does NOT guarantee token-exact MTP rollback after a rejected draft - is superseded as of v0.6.10 (2026-08-22T12:03:16Z). This does NOT relax the house varied-prompt/raw-token MTP-invariance requirement: the v0.6.9 tradeoff was a fork-rollback-path issue, whereas the #25618 quantized-target divergence (clm-0094/clm-0104/clm-0106) and #26750 spec-path acceptance collapse (clm-0102) are separate upstream issues on the verify/target side that still stand.
Sixth arm of the MTP-correctness dossier, superseding the operative claim of clm-0105 (fork rollback exactness). Fork-vendor release-note context only (single-payload verification on gfx1151 Vulkan; speed figures are single runs, NOT the BENCHMARKS.md protocol), so nothing inherits into a HaloBench number. This changes NO comparison boundary: HO-009 plain-vs-n2 results (clm-0097, clm-0101, cfg-0141/cfg-0142) were measured on UPSTREAM ggml-org/llama.cpp commit 7077abb on ROCm0/gfx1151, which never ran the fork's snapshot-plane rollback path, so v0.6.10 does not alter their interpretation; the active HO-009 n2 tau2 lane (t_ca9ef5d9) is untouched. Also flagged as LINK-UP for the ling-30-flash candidate screen: v0.6.10 adds bailingmoe3/Ling 3.0 DSpark spec support (upstream PR #27508, merged 2026-08-22T09:19:49Z). Runtime complete; NOT exercisable end-to-end until Ling-3.0-flash draft GGUFs are published (inclusionAI/Ling-3.0-flash-dspark has model.safetensors up, GGUFs not out as of 2026-08-22T16:19Z). When the DSpark GGUFs land, Ling-3.0-flash (already benched, clm-0086: MTP n_max 1/2/3 engagement real but median decode ratios 1.007x/0.968x/0.850x below the 1.10x promotion threshold, so plain decode stayed production default) can be re-examined against the DSpark draft path once a registered fork build ships it. Keep clm-0105 for the v0.6.9 historical fork record.