clm-0134
measured-heremed ●●○
citable URL: https://halobench.com/records/clm-0134/ — this address never moves; the anchor /records/#clm-0134 keeps resolving
The pwilkin candidate cost 10.26 Wh per correct answer on its tau2 0-25 run (205.2 Wh whole-session across 20 correct, mean 187 W, 0.31 pence at 30.3 p/kWh; eng-0280) — below the worker's 15.18 (eng-0278). The direction is not a clean A/B: the pwilkin stack decodes faster so its window was shorter (66 vs 106 min -> less total energy), even though it answered fewer tasks correctly (20 vs 22). So it is cheaper per correct answer but on a lower capability base — the same speed/quality trade seen everywhere in this comparison, expressed in energy.
Note — the record's own working
Wh-per-correct = wh_total / tasks_passed, whole-session, same method as eng-0081/eng-0278 for comparability. Energy differenced from the aihydra HA cumulative-kWh counter over the recorded window. Confidence medium: the comparison to eng-0278 is confounded (quant, spec, build all differ) and rests on the canonical 10.1 W idle floor. The efficiency direction is robust (a faster decode over a fixed idle floor is fewer Wh per answer), but it buys that efficiency at the 8-point capability cost of clm-0132.
Cited by — computed at build time, never stored
model pages qwen38-27b
candidate gate history qwen38-27b