Home › Evidence › Records › clm-0131
⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.

clm-0131

measured-heremed ●●○
citable URL: https://halobench.com/records/clm-0131/ — this address never moves; the anchor /records/#clm-0131 keeps resolving

The Qwen3.8-27B worker config cost 15.18 Wh per correct answer on the tau2 0-25 run (333.87 Wh whole-session across 22 correct, mean 189.6 W, 0.46 pence at 30.3 p/kWh; eng-0278) — well under the 38.29 Wh per correct of the earlier single-slot Q8_0 run (eng-0081). Two things drive the direction, neither a clean A/B: serving --parallel 4 amortises the box's ~190 W draw across ~4 concurrent tasks, and the run answered more tasks correctly (22 vs 15). Quant, build, reasoning and parallelism all differ between the two windows, so this is an efficiency observation about the whole worker profile, not an attribution to any single lever.

verified 2026-09-06 · volatility medium
evidence eng-0278 run-0658

Note — the record's own working

Wh-per-correct = wh_total / tasks_passed, whole-session (includes user-simulator wait and inter-task gaps), the same method as eng-0081 so the two are comparable. Energy differenced from the aihydra HA cumulative-kWh counter across the recorded run window. Confidence is MEDIUM because the comparison to eng-0081 is confounded (four fingerprint fields differ) and because the delta_w rests on the canonical 10.1 W idle floor while a spot idle reading this session was ~20 W — a fresh post-fan-control idle-baseline would refine delta_w (not wh_per_task). The multi-slot amortisation effect itself is robust: a fixed ~190 W spread over more concurrent useful work is fewer Wh per answer, which is the general reason a worker that actually receives concurrent tasks is cheaper per task than one driven single-stream. See clm-0130 for the throughput side.

Cited by — computed at build time, never stored

model pages qwen38-27b
candidate gate history qwen38-27b