Corrections
Every claim this lab has withdrawn, kept at its citable id with its full original text — struck through nowhere, hidden never. Each entry names what replaced it; the worked diagnosis of what went wrong lives on the record itself. These entries also appear inline in the Log, badged, in the same timeline as every other learning: this page is a generated filter of that stream, not a shrine.
3 corrections on record · a headline result that gets retracted is the process working — catching it before it became policy is the retraction's whole value
SUPERSEDED by clm-0042's per-task measurement, which found the true energy cost roughly 10x lower — this run's figures are whole-arm totals padded by model loading and non-scoring tasks, not the model's actual energy per answer, and must not be used to rank models. As measured here: a correct τ² answer cost 78.0 Wh on the 122B with f16 KV, 100.1 Wh on Nemotron, and 115.3 Wh on the 122B with q8_0 KV, but only 8-27% of each arm's wall time fell inside a scored task. aihydra's power envelope stands on its own: idle 10.1 W, 150-168 W under inference, peaking at 218 W.
corrected by: clm-0042 — the withdrawn record keeps its URL and full text; the successor carries the number to cite
SUPERSEDED: the Pass^1 = 1.000 reported by this run came from a 3-task subsample biased toward the domain's easiest tasks — the other two of the original five never terminated and were excluded as infrastructure errors. The sustained score across a realistic sample is clm-0037's 0.545 (n=22), which is the number to cite for this model on tau2 airline. This run was nonetheless the project's first genuine capability measurement rather than a throughput number, and it proved the harness and scoring path work.
corrected by: clm-0037 — the withdrawn record keeps its URL and full text; the successor carries the number to cite
RETRACTED: the per-model reward rankings and quantised-KV cost reported by this run do not hold — they came from 5-task tau2 arms whose ~0.40 run-to-run noise and 40-step cap bias were only characterised afterward (clm-0036), so the reward numbers below are not usable. The corrected 122B score is clm-0037's; the corrected cross-model comparison is clm-0039's. What survives is categorical, not scored: the 122B fails to terminate some tasks with thinking on, and gpt-oss fails the domain outright.
corrected by: clm-0037 clm-0039 — the withdrawn record keeps its URL and full text; the successor carries the number to cite