⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.
clm-0113
measured-herehigh ●●●
citable URL: https://halobench.com/records/clm-0113/ — this address never moves; the anchor /records/#clm-0113 keeps resolving
On the v0.6.10 strix-halo fork under Vulkan/RADV (build 2586f6edd), Qwen3.6-35B-A3B-MTP UD-Q4_K_M: raising native-MTP draft n_max from 2 to 4 (reasoning OFF) is FLAT — n4-off full-26 scored 17/26 (mean 0.653846), identical to n2-off (17/26, run-0578); Fisher exact two-sided n4-vs-n2 p=1.000000. Neither n2 nor n4 is significantly worse than plain (plain 22/26, mean 0.846154, run-0576): n4-vs-plain p=0.199350, n2-vs-plain p=0.199350 — matching the documented HO-009-v0610 baseline. So n_max=4 does NOT narrow the 5-correct deficit to plain, does NOT degrade, and does NOT collapse (EOS-cliff sentinel PASS 0/8 at n_max=4; the clm-0055 accumulation trigger did not reproduce). The reasoning-ON path is BLOCKED, not just degraded: the plain-rea-ON content sentinel FAILED under G6 (reasoning_content_len=1846 but assistant_content_len=0, empty assistant content, rc=2) — the same empty-content/content-drop defect lineage that forced reasoning OFF in HO-009-AB re-emerges under -rea on on this fork/stack. Because the failure is on the PLAIN reasoning-ON control (no MTP contribution required), the entire reasoning-ON lane is do-not-score: n2-rea-on tau2 skipped, gated n4-rea-on cancelled. Within the admitted reasoning-OFF capability matrix, plain remains the capability winner (22/26). Scope is strictly single-model/single-build MTP n2|n4 × reasoning on|off on this fork stack: cfg-0160..0163, build 2586f6edd, UD-Q4_K_M sha 0b21525e, f16/f16 KV, c32768, airline@668d3bcd seed 42 tasks 0-25 trials 1 agent temp0 max_tokens4096 judge claude-haiku-4.5 temp 0. NO cross-model, KV-quant, DFlash2, n>4, degradation- threshold, or production-throughput claim is made or inherited.
Capability-matrix claim from HO-013-REVIEW (t_ebb2e334) ADMIT verdict: Cell-0 EOS sentinel PASS, Cell-1 n4-off full-26 ADMITTED (17/26, comparable + complete), reasoning-ON REJECTED under G6 (do-not-score) with n2-ON tau2 skipped and gated n4-ON cancelled. Values are RECORD values from the raw cell1-n4-off/tau2-full26-summary.json (tool_call 211 / empty-arg 1, verified on-disk by review) and sentinel reasoning-content files (reasoning_content_len = 1846, NOT 2182 as parent metadata incorrectly stated — provenance corrected in the review verdict). Fisher p-values recomputed by review (n4-vs-plain 0.199, n4-vs-n2 1.0, n2-vs-plain 0.199). Energy join: eng-0251 (57.17 Wh, wh/task 2.199) added this ingest; compare arms eng-0235/eng-0236. This closes the MTP n_max investigation: n2 and n4 are statistically identical, both trail plain. Reasoning-on remains a production-relevant blocker for the config that must be reasoning-ON; a DIAG card mirroring HO-004-DIAG is recommended for the reasoning-ON empty-assistant content-drop lineage on this fork. The plain- remains-capability-safe recommendation of clm-0108/clm-0110 is unchanged within reasoning-OFF scope; this claim does NOT relax the house varied-prompt/raw-token MTP-invariance requirement (clm-0102/clm-0104/clm-0106 still stand).