⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.
clm-0110
measured-herehigh ●●●
citable URL: https://halobench.com/records/clm-0110/ — this address never moves; the anchor /records/#clm-0110 keeps resolving
On the v0.6.10 strix-halo fork under Vulkan/RADV (build 2586f6edd), Qwen3.6-35B-A3B-MTP UD-Q4_K_M with reasoning OFF recovers MTP n2 capability that build 7077abb-upstream-on-ROCm0 destroyed. This is a STACK effect, not a build-only change: AB ran ggml-org 7077abb on ROCm0/gfx1151 (cfg-0144/0145, run-0568/0569/0570/0571), while this retest ran the Nathanw1014 fork on Vulkan/RADV (cfg-0149/0150, run-0575..0578) — repo, backend and build all changed together, so the comparison is strictly "v0.6.10-fork-on-Vulkan vs 7077abb-upstream-on-ROCm0", never a build-only claim. Within v0610, native-MTP n2 (--spec-type draft-mtp --spec-draft-n-max 2) scored 17/26 (mean 0.65385) vs plain 22/26 (mean 0.84615); Fisher exact two-sided n2-worse p=0.1994 — n2 is NOT significantly worse than plain on this stack. Against the ORIGINAL AB n2 failure, n2 17/26 vs AB n2 0/26 (run-0571) p≈2.8e-07 (Fisher two-sided) — MTP n2 capability destroyed on 7077abb is recovered on v0.6.10. The empty-turn/agent-spam mechanism is resolved at the mechanism level: all 26 v0610 n2 terminations were user_stop with 0 too_many_errors, 0 max_steps, 0 empty-assistant and 1 empty-argument tool call (tool_call_messages 192), versus AB n2's 415 tool-call + 21 too_many_errors + 2 max_steps spam on run-0571. Scope is strictly qwen35moe target, UD-Q4_K_M, reasoning-off, n_max=2 only. NO DFlash2, stock-MTP, n_max>=4, reasoning-on, alt-model, alt-backend, or production-throughput claim is made or inherited.
Capability-recovery claim resolving the HO-009-AB negative verdict (clm-0108) for the v0.6.10 fork/Vulkan stack. Values are RECORD values from the raw tau2 summaries (plain tool_call 215 / empty-arg 1; n2 tool_call 192 / empty-arg 1), NOT the over-reported parent metadata (315/2). MTP draft acceptance range 0.728-1.00 per the n2 guard (run-0577). Fisher p-values: n2-vs-plain within v0610 = 0.1994 (two-sided, not significantly worse), n2 17/26 vs AB n2 0/26 ~2.8e-07, plain v0610 22/26 vs AB plain 6/26 = 1.7e-05. Energy join: eng-0235 (plain) / eng-0236 (n2). The mechanism-level empty-turn recovery supersedes the AB reasoning-off capability loss ONLY on this exact stack+config scope; it does NOT relax the house varied-prompt/raw-token MTP-invariance requirement (clm-0102/clm-0104/ clm-0106 still stand) and does not alter the plain-remains-capability-safe upstream-7077abb recommendation of clm-0108.