clm-0035
measured-herelow ●○○retracted
citable URL: https://halobench.com/records/clm-0035/ — this address never moves; the anchor /records/#clm-0035 keeps resolving
RETRACTED: the per-model reward rankings and quantised-KV cost reported by this run do not hold — they came from 5-task tau2 arms whose ~0.40 run-to-run noise and 40-step cap bias were only characterised afterward (clm-0036), so the reward numbers below are not usable. The corrected 122B score is clm-0037's; the corrected cross-model comparison is clm-0039's. What survives is categorical, not scored: the 122B fails to terminate some tasks with thinking on, and gpt-oss fails the domain outright.
superseded by clm-0037 clm-0039 — the corrected statement lives there; this record keeps its URL and full text
verified 2026-08-09 · volatility medium
Note — the record's own working
⛔ **RETRACTED IN LARGE PART — see clm-0036 before reading any number below.**
The identical stock q8_0 configuration was subsequently run three times and scored 0.60, 1.00, 0.60. The run-to-run noise of this 5-task harness is ~0.40, which is the same size as nearly every gap reported here. Specifically:
· **The KV quality finding (f16 1.00 vs q8_0 0.60) is retracted.** f16 is 2 for 2 and
stock q8_0 is 1 for 3 — suggestive, not significant. Underpowered, not disproven.
· **The Nemotron thinking result (0.60 → 1.00) is not established.** It is exactly one
noise-width, and was the headline conclusion here.
· **The per-model rankings are not established** for the same reason.
What survives is the categorical outcomes rather than the scores: the 122B's failure to TERMINATE with thinking on (2 of 5, reproduced across four runs in clm-0033), and gpt-oss's 0.00 total failure alongside its independently reproduced staleness bug.
Confidence dropped medium → low. The text below is left unedited as written so the reasoning that produced it stays inspectable.
---
τ²-bench airline, 5 tasks, `--max-concurrency 1`, `--max-steps 40`, 60-min backstop per arm. aihydra, ROCm 7.1.0, stock llama.cpp 3653e6d, `-fa on`, f16 KV unless stated, `--parallel 1`. Thinking gated server-side with `-rea on|off`, and every ON arm probed for `reasoning_content` before running — all four models produce it, so none was skipped.
## The thinking matrix
| model | thinking OFF | thinking ON | wall OFF / ON | |---|---|---|---| | Qwen3.5-122B-A10B | **5/5 @ 1.00** | 3/5 @ 1.00 (2 deadlock) | 10 / 90+ min | | Nemotron-3-Super-120B-A12B | 5/5 @ 0.60 | **5/5 @ 1.00** | 15 / 45 min | | Qwen3.6-35B-A3B | 5/5 @ 0.40 | 1/1 @ 1.00 (backstop) | 4 / 60+ min | | gpt-oss-120b | 5/5 @ **0.00** | 5/5 @ 0.60 | 15 / 10 min |
⚑ **THIS CORRECTS THE FRAMING OF clm-0033.** That claim measured the 122B and concluded thinking is a net negative — correctly, for that model. But I treated it as the likely general case and flagged Warden's `thinking=high` cron for reversion on that basis. **The opposite is true for the other three.** Nemotron goes from 0.60 to a perfect 1.00 with thinking on, completing every task. gpt-oss is unusable without it (0.00) and merely poor with it (0.60).
So the lever is not "thinking good" or "thinking bad" — it is **per-model, and must be measured per-model**. Any placement decision that sets it globally is wrong for three models out of four whichever way it is set.
**Nemotron-3-Super with thinking on is the standout of the whole set**: 5/5 completed at reward 1.00, the only configuration to match the 122B's best while also terminating on every task. It is the slowest model measured (17.51 tok/s decode, clm-0025) — which makes it a genuine speed-versus-reliability trade rather than a dominated option.
⚠ The 35B's ON arm completed only 1 task before the 60-minute backstop, so its 1.00 is over n=1 and means very little. It does show the same deadlock tendency as the 122B.
## q8_0 KV costs task success — clm-0022's open half, closed
Same model, same tasks, thinking off, ONLY the cache type differing:
| KV | completed | mean reward | failures | |---|---|---|---| | f16 | 5/5 | **1.00** | — | | q8_0 | 5/5 | **0.60** | tasks 2 and 4 hit `max_steps` at 42 and 41 turns, scored 0.0 |
**Quantised KV is not free.** clm-0022 established it costs nothing in SPEED once the dequant patch is applied (+70.3% recovery at production depth). This is the other half: it costs **40% of task success** in this sample, and the failure mode is specific and legible — the two failed tasks did not produce wrong answers, they **failed to terminate**, running to the step limit exactly as the deadlock cases do.
That is a more useful result than a perplexity delta would have been. The KL route was dead on this build (clm-0029, PPL 532 on repetitive English), and τ² turned out to measure the thing that actually matters — whether the model finishes the job.
## Scope, stated plainly
· **One domain, five tasks.** Every cell is n=5. A 0.40 swing is two tasks. · **Self-play** — agent and user simulator are the same model throughout. · **Deterministic scoring only** — DB-state match and tool-call trajectory, no LLM judge. · **Read paths only**; no write actions exercised. · The q8_0 result is on the STOCK build. Whether the dequant patch (clm-0022) also
restores quality, or only speed, is **untested and is now the obvious next run**.
Cited by — computed at build time, never stored
candidate gate history nemotron3-super