citable URL: https://halobench.com/records/clm-0048/ — this address never moves; the anchor /records/#clm-0048 keeps resolving
Simulator identity changes what tau2 measures. On identical seeded tasks with the same 122B agent: self-play and a Haiku-4.5 simulator produce identical rewards (9/9) but Haiku lengthens conversations (162 to 210 total turns) — the self-play customer volunteers what its twin needs, the clm-0043 correlated-blind-spot effect now measured. A Sonnet-4.5 simulator failed the agent on 2 of 9 tasks; forensics split them — one simulator persona violation (discounted), one genuine agent bug (clm-0049). The registered prediction that an off-box simulator would cut wall time FAILED: self-play was faster per task (median 121s vs 173s) and cheaper in local energy (4.4x less per task), because the independent simulator's longer conversations mean more agent computation — not because of wait-state draw, which resident-quiet measurement rules out (13.6 W).
Haiku pin CONFIRMED as primary: adequate on rewards, independent by family, and the cheapest option with evidence. Sonnet proved less persona-constrained rather than more discriminating (task 5: its customer pivoted to demanding the cancellation the scenario explicitly forbids). Periodic Sonnet secondary passes remain worth considering — its unconstrained pressure DID surface clm-0049's genuine bug, even if by accident. The self-play speed/energy advantage is real but buys measurements of an easier, flattered benchmark: correctness of measurement beats cost here. Provider routing is account-guardrailed to Anthropic (verified first-probe leak to Bedrock, then dashboard- confirmed). Wall-time comparison cannot separate cache effects from API latency — per-turn agent-side timing would; registered and closed as a failed prediction.