clm-0036
measured-herehigh ●●●
citable URL: https://halobench.com/records/clm-0036/ — this address never moves; the anchor /records/#clm-0036 keeps resolving
τ²-bench at 5 tasks cannot resolve the differences drawn from it in this project's capability matrix. The IDENTICAL stock q8_0 configuration scored 0.60, 1.00 and 0.60 across three independent runs — a 0.40 spread from noise alone. That is the same size as most gaps in the capability matrix, so those gaps are not established. Three tasks always pass and two are coin-flips, which makes the 5-task mean effectively two Bernoulli trials. The domain has 50 tasks available; five were used.
verified 2026-08-09 · volatility low
Note — the record's own working
**This claim retracts the KV half of clm-0035 and puts every ranking in it in doubt.**
**Core finding:** repeating the identical stock q8_0 configuration three times scored 0.60, 1.00, 0.60 — the patched-vs-stock "quality fix" that prompted the repeat lay entirely within that spread, and would have been published as a confirmed result without it. The failures are not scattered: in both 0.60 runs the same two tasks failed, both at the step-limit boundary, which means the 5-task mean is really two Bernoulli trials (values 0.60/0.80/1.00 only) — and the step limit itself was a 40-turn budget chosen for wall-clock convenience, cutting tasks off right where they naturally finish rather than scoring them wrong.
**What survives:** categorical completion failures, not reward differences — the 122B deadlocking with thinking ON (`clm-0033`) and gpt-oss's stale-answer bug (`clm-0025`). **What does not survive:** q8_0-vs-f16 quality, whether the dequant patch restores it, and the Nemotron thinking-on result from `clm-0035` — all fall inside this noise band.
The full worked diagnosis — the run table, the Bernoulli-trial argument, the fix (the airline domain ships 50 tasks; five were used) and the methodological lesson about confirming results that look *too* clean — is written up as a standing rule in `docs/methodology-lessons.md` §1.
Cited by — computed at build time, never stored
model pages qwen35-122b
candidate gate history nemotron3-super