Instrumentation gaps — structured-record audit
Refreshed 2026-08-19 from the repository’s structured records. Counts below describe what is recorded, not what may survive only in raw artifacts.
The repository now contains 459 run records and 188 energy records. Energy is no longer an absent collection: 314 runs reference an energy record, five carry a structured reason why an honest join cannot be made, and other unjoined records remain outside the published model-page tripwire. The reference check currently reports zero model-page energy warnings.
Record types: 5 of 12 missing
| § | type | status | what its absence costs |
|---|---|---|---|
| 2.1 | node |
✅ built | |
| 2.2 | config |
✅ built | |
| 2.3 | run |
✅ built | |
| 2.4 | energy |
✅ built 2026-08-10 | was absent while 99 runs were recorded |
| 2.5 | idle_baseline |
✅ built 2026-08-10 | |
| 2.6 | workload |
❌ missing | no first-class drilldown dimension — suite strings and run prose carry the workload identity |
| 2.7 | telemetry |
❌ missing | production behaviour and benchmark behaviour are not separable; Warden’s real usage is invisible |
| 2.8 | suite |
❌ missing | suites are named as free strings (tau2-bench-airline@v1.0.1) with nothing defining or versioning them |
| 2.9 | decision |
❌ missing | no browsable log of what was chosen and why — the reasoning lives in claims and commits instead |
| 2.10 | claim |
✅ built | carrying more weight than intended, precisely because decision is absent |
| 2.11 | contribution |
✅ built | |
| 2.12 | edge |
❌ missing | cross-host coordination unrepresented; cross_host_pair is a bare string with nothing to point at |
suite is the most quietly damaging. Every run references a suite by name, and nothing
anywhere says what that suite is — so a suite could change underneath a series and no
record would show it. That is the same class of hazard as the chat-template drift the
config schema explicitly guards against.
Run metrics: the headline one is still in a single run
The design doc’s §2.3 example lists seven metrics. Across all 459 runs:
| metric | present in | note |
|---|---|---|
decode_tps |
198 / 459 | primary performance metric |
prefill_tps |
187 / 459 | |
tool_call_success |
5 / 459 | specified as a gate; many guard outcomes are instead represented by task totals and guard links |
empty_arg_calls |
4 / 459 | |
ttft_ms |
2 / 459 | |
turns_before_degradation |
1 / 459 | |
time_to_correct_answer_s |
1 / 459 | the project’s own headline metric remains scarcely captured |
structured energy.ref |
314 / 459 | energy lives in linked eng-* records, not inside metrics |
⚑ time_to_correct_answer_s is what docs/14-model-backend-benchmark.md names as the
headline — “time-to-correct-result + loops-to-done, with tool-call success as a gate” —
and what the operator independently identified as the thing that actually determines whether a
model is pleasant to use. It is populated once.
The consequence is visible in this session’s own reasoning: model comparisons were argued
on mean_score for hours before turn counts and wall-clock were pulled out of the raw
simulation files by hand. That analysis should have been a query against recorded metrics,
not archaeology.
Energy coverage and limits
- 314 runs are joined through
energy: {ref: eng-*}. - Five runs are explicitly unjoinable through
energy_unjoined_reason; this includes pre-aihydra work with no defensible wall-meter window and combined windows that cannot be apportioned without inventing data. - 188 energy records exist. A single invocation window may legitimately cover several run rows, so run links and energy-record counts are not expected to match.
- An energy total is not automatically Wh/correct, Wh/token or model-only energy. Each claim must preserve the window and denominator actually measured.
Implemented controls
energycollection is built to §2.4, with build-time enforcement:delta_wmust equalmean_w_active − baseline_w, a non-recordedwindow must explain its derivation, and joules are rejected outright.idle_baselinecollection is built to §2.5, with aihydra’s first record. Two of the three deltas the doc asks for are markedmeasured: falserather than estimated — notablyigpu-resident-quiet, which is what priceskeep_alive, andnpu-resident-quiet, which is the question an always-on NPU triage loop turns on.runsgained optionalstarted_at/ended_at.bench/runmeta.shinstalled and wired into all seven bench scripts via an EXIT trap, so every invocation now emits its UTC window, duration, exit code and acontentionflag to~/run-meta.jsonl. Verified on a deliberately failing run — failed runs emit too, which is the case most likely to be silently dropped.
What remains
- Capture
time_to_correct_answer_sandtool_call_successas standard, not incidentally. - Measure
igpu-resident-quiet— a model loaded and idle. It is one server start and ten minutes of doing nothing, and it unlocks thekeep_alivecost question entirely. - Build
suite, so a suite name means something checkable. - Decide whether
workload,telemetry,decisionandedgeare still wanted, or whether the design has moved on. Four unbuilt record types may be a plan that outgrew itself rather than a backlog — but that should be an explicit decision, not a silent omission.