The core measurement is a clean A/B on agentic work (multi-turn jobs where the model calls tools and acts, rather than answering once): same model, same tasks, thinking gated server-side and verified per arm clm-0033. Off, the model solves every task at reward 1.0 in ~10 minutes. On, it argues in circles on the two adversarial-pressure tasks — the ones where the right move is to refuse or escalate — and never terminates, running 4h15m and 2h10m before being cut. Where both arms complete a task, reward is identical: thinking produced fewer, more expensive turns, and no measurable accuracy.
The first write-up generalised that result, and the record corrects it in public. A four-model matrix pointed the other way for every other model tested — then the noise audit landed: the identical configuration scored 0.60, 1.00 and 0.60 on repeat, a 0.40 spread from noise alone, the same size as nearly every gap in the matrix, and the per-model rankings were retracted clm-0035 clm-0036. The categorical outcomes survive — the 122B's thinking-on deadlocks reproduced across four runs — while the score-based case that thinking helps the others remains suggestive, unproven.
The operational rule this leaves is unglamorous: treat thinking as a per-model lever that is never assumed safe and never set globally. The 122B's corrected baseline, thinking off, is 0.545 ±0.208 at n=22 under the full step budget clm-0037 — decisively below the 1.00 its easy-task subset had reported, which is its own lesson in sampling. A mechanism for the deadlocks — extended reasoning giving the model more room to rationalise continuing instead of stopping — is plausible and explicitly unestablished clm-0033; what is measured is that the behaviour flips cleanly with the lever.