eng-0168
window recorded
citable URL: https://halobench.com/records/eng-0168/ — this address never moves; the anchor /records/#eng-0168 keeps resolving
67.05 Whmean 103.69 W − 10.1 W idle floor (box-idle-all-empty) = 93.59 W attributable
window 2026-08-17T19:09:14Z → 2026-08-17T19:48:02Z · provenance recorded
method wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the run-meta window) · no contention
per task 4.7896 Wh · 0.1451 p at 30.3 p/kWh
covers run-0433
aihydra qwen3-coder-next-80b-fullbench tau2 26-task run window (19:09:14Z-19:48:02Z, 2328s), matching run-0433 exactly (excludes the preceding 51s server-load window, per the deepseek-v4-flash/laguna-s-21/ ornith-35b convention). Counter interpolated from raw ~10s-resolution HA history samples pulled fresh this session: start 26.265702 kWh, end 26.332756 kWh. wh_per_task is Wh-per-CORRECT-ANSWER: wh_total / tasks_passed (14 of 26, run-0433.metrics) = 4.7896 Wh, 0.1451 pence at the standing 30.3 p/kWh tariff. WHOLE-SESSION energy -- includes simulator (user_llm) wait time end to end, not an agent-only or decode-only figure, directly comparable to ornith-35b's eng-0134 (6.4834 Wh), laguna-s-21's eng-0124 (7.2588 Wh), deepseek-v4-flash's eng-0107 (24.29 Wh), nemotron3-super's eng-0116 (32.43 Wh) and qwen38-27b's Wh figure (38.29 Wh) on the same denominator. THIS IS THE NEW LOWEST Wh-per-correct-answer among full standard-26-task tau2 comparisons on this box, ahead of ornith-35b's previous-leading 6.4834 Wh -- notable because this candidate has, at the SAME TIME, the LOWEST capability score of the same comparison set (0.5385 mean reward, below qwen38-27b's 0.577): the efficiency lead is real and measured on the identical protocol, but it does not come with a capability lead, and the verdict must not imply otherwise. Screened against every other benched model's Wh-per-correct-answer before being claimed (the qwen35-122b 5.71 Wh/correct figure remains a 5-task SMOKE window predating the standard-26-task convention and is not a like-for-like comparator, same caveat every prior leader's page already carries).
Wh, never joules — read as a cumulative-counter difference across the window; delta over the idle baseline is the only figure that means anything
Cited by — computed at build time, never stored
model pages qwen3-coder-next-80b