eng-0146
window recorded
citable URL: https://halobench.com/records/eng-0146/ — this address never moves; the anchor /records/#eng-0146 keeps resolving
22.73 Whmean 54.69 W − 10.1 W idle floor (box-idle-all-empty) = 44.59 W attributable
window 2026-08-17T06:37:19Z → 2026-08-17T07:02:15Z · provenance recorded
method wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the run-meta window) · no contention
per task 4.5457 Wh · 0.1377 p at 30.3 p/kWh
covers run-0372
aihydra deep-thought-posttrain FULL 26-task tau2 airline run (run-0372), window 06:37:19Z-07:02:15Z (1496s) matching run-meta.jsonl's tau2-full-26task line exactly. Counter interpolated from raw ~10s-resolution HA history pulled fresh this session: window-start bracket 24.8130654543638276/24.8131226748228076 kWh (06:37:10.533Z/06:37:20.529Z) -> interpolated 24.8131139 kWh at 06:37:19Z; window-end bracket 24.8358238190412576/24.8358655422925976 kWh (07:02:10.545Z/07:02:20.551Z) -> interpolated 24.8358424 kWh at 07:02:15Z. Delta 0.0227285 kWh = 22.7285 Wh over 1496s = 54.69 W mean, 44.59 W over the 10.1 W box-idle baseline — a real, elevated draw (the raw samples show a sustained faster climb from ~06:39Z through ~06:56Z, tracking the busiest run of tasks in the log, settling toward the baseline slope in the final ~6 minutes as the remaining tasks resolved quickly or errored near-instantly).
wh_per_task IS ARITHMETICALLY DEFINED HERE (5 tasks scored reward 1.0 out of 26 — not the zero-correct-answers case that would make this a literal division by zero) but is published with a hard caveat, not as a clean comparator: run-0372's own guard (run-0363) and this arm's own tool_call_messages count (0 across all 26 simulations, 233/233 assistant messages) place this run squarely inside the SMOKE validity gate (protocol.json comparison_rules / benchmark-protocol.md §11) — a mean computed where the agent never calls a tool is INVALID, because the airline domain hands out an untouched-database point and a vacuous COMMUNICATE point to any agent that never acts, which is mechanically how a model that only ever says "42" collected 5 nominal "passes". 4.5457 Wh / 0.1377p per NOMINALLY-correct answer is therefore reported for completeness and arithmetic honesty only, and is explicitly NOT comparable to any other model's Wh-per-correct-answer figure on this site (ornith-35b 6.4834, laguna-s-21 7.2588, deepseek-v4-flash 24.29, nemotron3-super 32.43, qwen38-27b 38.29) — those all cleared the SMOKE gate with a nonzero tool_call_messages count; this one did not, and its "correct answers" are not evidence of task completion. wh_total (22.7285 Wh, the undiluted figure) is the number this run actually supports.
Wh, never joules — read as a cumulative-counter difference across the window; delta over the idle baseline is the only figure that means anything
Cited by — computed at build time, never stored
model pages deep-thought-posttrain
candidate gate history deep-thought-posttrain