Home › Evidence › Records › clm-0108
⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.

clm-0108

measured-herehigh ●●●
citable URL: https://halobench.com/records/clm-0108/ — this address never moves; the anchor /records/#clm-0108 keeps resolving

HO-009-AB measured plain-vs-n2 capability A/B on Qwen3.6-35B-A3B-MTP at production agentic depth (c32768), reasoning OFF on both arms (build 7077abb, ROCm0/gfx1151, UD-Q4_K_M, f16/f16 KV). Plain scored 6/26 (mean 0.2308) vs native-MTP n2 (--spec-type draft-mtp --spec-draft-n-max 2) scoring 0/26 (mean 0.0). Fisher exact one-sided n2-worse p=0.0113, two-sided p=0.0226: with the tau2 reasoning_content-drop confound removed, MTP at n_max=2 is NOT capability- safe on this build — n2 is a real, statistically significant capability loss vs plain on the same task set, not noise-around-plain. Both arms clean: 0 infra / 0 empty-assistant / 0 empty-arg tool calls, no single-token cliff (min asst completion_tokens=14, n2). Wilson 95% CIs: plain [0.11,0.42], n2 [0.00,0.13]. n2 terminates via 21 too_many_errors / 2 max_steps / 3 user_stop with 415 tool call messages -> agent-side validation-error spam, a capability loss in practice.

verified 2026-08-22 · volatility medium
evidence run-0568 run-0569 run-0570 run-0571

Note — the record's own working

Plain-vs-n2 capability comparison only (assessed as-run on build 7077abb). This does NOT change the n2 decode-throughput context of clm-0097: n2 still improves server /completion decode throughput through d204800, but that throughput benefit carries a capability cost, so plain remains the capability-safe option. No n_max>=4, KV-quant, alt-backend, or model-change claim is made or inherited. No n2 quality-equivalence, production recommendation, or throughput claim beyond this plain-vs-n2 capability comparison on this card. The reasoning-off requirement itself costs capability on this model (plain 6/26 vs the 122B f16 10/26 / q8_0 15/26 range), consistent with the DIAG confound fix; both arms ran identical reasoning-off protocol. The v0.6.10 MTP rollback re-land is a separate future-facing question tracked in the MTP-correctness dossier (clm-0105/clm-0107); it does not change this as-run verdict on upstream 7077abb.