⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.
The core measurement is a clean A/B on agentic work (multi-turn jobs where the model calls tools and acts, rather than answering once): same model, same tasks, thinking gated server-side and verified per arm clm-0033. Off, the model solves every task at reward 1.0 in ~10 minutes. On, it argues in circles on the two adversarial-pressure tasks — the ones where the right move is to refuse or escalate — and never terminates, running 4h15m and 2h10m before being cut. Where both arms complete a task, reward is identical: thinking produced fewer, more expensive turns, and no measurable accuracy.
The first write-up generalised that result, and the record corrects it in public. A four-model matrix pointed the other way for every other model tested — then the noise audit landed: the identical configuration scored 0.60, 1.00 and 0.60 on repeat, a 0.40 spread from noise alone, the same size as nearly every gap in the matrix, and the per-model rankings were retracted clm-0035 clm-0036. The categorical outcomes survive — the 122B's thinking-on deadlocks reproduced across four runs — while the score-based case that thinking helps the others remains suggestive, unproven.
The operational rule this leaves is unglamorous: treat thinking as a per-model lever that is never assumed safe and never set globally. The 122B's corrected baseline, thinking off, is 0.545 ±0.208 at n=22 under the full step budget clm-0037 — decisively below the 1.00 its easy-task subset had reported, which is its own lesson in sampling. A mechanism for the deadlocks — extended reasoning giving the model more room to rationalise continuing instead of stopping — is plausible and explicitly unestablished clm-0033; what is measured is that the behaviour flips cleanly with the lever.
generated from the citations above — each entry shows the claim's current state, so this page cannot silently rest on withdrawn evidence
RETRACTED: the per-model reward rankings and quantised-KV cost reported by this run do not hold — they came from 5-task tau2 arms whose ~0.40 run-to-run noise and 40-step cap bias were only characterised afterward (clm-0036), so the reward numbers below are not usable. The corrected 122B score is clm-0037's; the corrected cross-model comparison is clm-0039's. What survives is categorical, not scored: the 122B fails to terminate some tasks with thinking on, and gpt-oss fails the domain outright.
corrected by clm-0037 clm-0039 · cited here deliberately: cited as the retracted cross-model matrix — the retraction is the finding this page reports, not a defect in it
τ²-bench at 5 tasks cannot resolve the differences drawn from it in this project's capability matrix. The IDENTICAL stock q8_0 configuration scored 0.60, 1.00 and 0.60 across three independent runs — a 0.40 spread from noise alone. That is the same size as most gaps in the capability matrix, so those gaps are not established. Three tasks always pass and two are coin-flips, which makes the 5-task mean effectively two Bernoulli trials. The domain has 50 tasks available; five were used.
The 122B's real τ²-bench airline score is 0.545 +/-0.208, not the 1.00 reported by 5-task runs, which sampled only the easiest tasks in the domain. At n=22 it passes 12 and fails 10. Nothing was cut at the 200-step budget, and tasks run a median of 22 turns and a maximum of 36 — so the model's long tasks are long in TOKENS per turn (~10k), not in turns.