Home › Evidence › Records › clm-0047

clm-0047

measured-heremed ●●○
citable URL: https://halobench.com/records/clm-0047/ — this address never moves; the anchor /records/#clm-0047 keeps resolving

Nemotron's quant confound resolves cleanly: on 12 identical seeded tasks, UD-Q4_K_M and UD-IQ4_XS produce IDENTICAL reward on every task (0.583 both), but IQ4_XS takes +11% more total turns (630 vs 567). Quantisation cost this model efficiency, never correctness — prior cross-model comparisons at mismatched quant understated Nemotron's speed, not its quality. At fair quant its turn median (25) still trails the 122B's (~20) on the same subset: the incumbent's efficiency lead narrows but survives.

verified 2026-08-11 · volatility medium
evidence run-0107 run-0108

Note — the record's own working

Stock 3653e6d, seeded subset (tasks 0-6,8-12), --max-steps 200, self-play (sound for a within-model lever, clm-0043). Energy followed duration, not draw: E1 ran 1.86x as long as E2 at near-identical wattage (eng-0061/0062). One task cut per arm. AUDIT ANNOTATION (2026-08-12 re-execution spot-check): the absolute 0.583 mean is partly a simulator artifact and MUST NOT be compared against pinned-simulator runs of other models. Re-running the E1 arm on tasks 0-3 under the protocol-pinned Haiku simulator flipped tasks 1 and 2 from 0.0 to 1.0 (trajectories: the self-play user talked the agent into DB-mutating policy violations the independent simulator never induces; total turns 158 -> 76). The within-configuration quant equivalence claimed here is unaffected. Cross-model comparison at fair quant requires re-running both arms under the pinned simulator. See audits/repeatability-2026-08.md.

Cited by — computed at build time, never stored

model pages nemotron3-super