Home › Evidence › Records › inc-0005

inc-0005

blockedunresolved
citable URL: https://halobench.com/records/inc-0005/ — this address never moves; the anchor /records/#inc-0005 keeps resolving

Abrupt power loss under sustained load; board never POSTed again — RMA submitted

detected
2026-07-23 · window 2026-07-23T09:57Z → unresolved · by smart-plug-statistics-retrospective · lag Warden itself noticed in ~71 seconds (failover at 09:58:24 after the 09:57:13 attempt). The warm-lane canary did not flag it until 06:33 the NEXT DAY — a ~20 hour gap during which nothing escalated. Warden knew; nothing told anyone.
signature: Hourly statistics show active operation (mean 41.3 W, peak 178.4 W) at 09:00 UTC, then mean 1.0 W / max 3.7 W at 10:00 UTC. Five hours at 0.5-0.6 W, then ~10 W from 15:00 UTC onward, flat ever since and slowly declining to ~5.2 W by 5 Aug.
should have caught it
Nothing alerted. The box simply stopped answering and the outage was noticed by its absence. A plug-power threshold alarm — "AI Power fell below 5 W while a model was supposed to be resident" — would have fired within minutes, and the same sensor that diagnosed this retrospectively could have raised it live. The data existed the whole time; nobody was looking at it.
diagnosis
initial: same failure as the 2026-07-21 OOM panic
actual: A DIFFERENT failure. The 21 July event wedged at a flat 45 W — powered, hung after POST. This one collapsed to ~0.5 W, which is BELOW the ~10 W standby the board draws today, so even standby power was absent for five hours. Recovery to ~10 W at 15:00 UTC is consistent with a plug cycle clearing a latched state back to standby. The board has not POSTed since. Most consistent with a power-delivery or protection latch under load rather than a software or configuration fault.
false-path cost: none — diagnosed remotely before travelling
falsified by: Home Assistant long-term statistics for sensor.hardware_ai_power_power. Raw history is ~10 day retention and had expired; statistics have permanent retention.
blast radius
configs affected: cfg-0002
lesson
Two lessons. First: the mitigations applied after 21 July (--ctx-checkpoints 4, desktop stack removed, swap 16G) target dynamic memory growth and are almost certainly IRRELEVANT to this failure — a patch chosen for the wrong incident buys false confidence. Second and more useful: a smart plug is an availability instrument, not just an energy one. The power curve distinguished "wedged after POST" from "lost power entirely" retrospectively, from another country, with no access to the machine. That signal should be an alarm, not an archaeology tool.
state
UNRESOLVED · class hardware

Cited by — computed at build time, never stored