Docs › instrumentation-gaps
⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.

Instrumentation gaps — structured-record audit

Refreshed 2026-08-19 from the repository’s structured records. Counts below describe what is recorded, not what may survive only in raw artifacts.

The repository now contains 459 run records and 188 energy records. Energy is no longer an absent collection: 314 runs reference an energy record, five carry a structured reason why an honest join cannot be made, and other unjoined records remain outside the published model-page tripwire. The reference check currently reports zero model-page energy warnings.


Record types: 5 of 12 missing

§ type status what its absence costs
2.1 node ✅ built
2.2 config ✅ built
2.3 run ✅ built
2.4 energy ✅ built 2026-08-10 was absent while 99 runs were recorded
2.5 idle_baseline ✅ built 2026-08-10
2.6 workload ❌ missing no first-class drilldown dimension — suite strings and run prose carry the workload identity
2.7 telemetry ❌ missing production behaviour and benchmark behaviour are not separable; Warden’s real usage is invisible
2.8 suite ❌ missing suites are named as free strings (tau2-bench-airline@v1.0.1) with nothing defining or versioning them
2.9 decision ❌ missing no browsable log of what was chosen and why — the reasoning lives in claims and commits instead
2.10 claim ✅ built carrying more weight than intended, precisely because decision is absent
2.11 contribution ✅ built
2.12 edge ❌ missing cross-host coordination unrepresented; cross_host_pair is a bare string with nothing to point at

suite is the most quietly damaging. Every run references a suite by name, and nothing anywhere says what that suite is — so a suite could change underneath a series and no record would show it. That is the same class of hazard as the chat-template drift the config schema explicitly guards against.


Run metrics: the headline one is still in a single run

The design doc’s §2.3 example lists seven metrics. Across all 459 runs:

metric present in note
decode_tps 198 / 459 primary performance metric
prefill_tps 187 / 459
tool_call_success 5 / 459 specified as a gate; many guard outcomes are instead represented by task totals and guard links
empty_arg_calls 4 / 459
ttft_ms 2 / 459
turns_before_degradation 1 / 459
time_to_correct_answer_s 1 / 459 the project’s own headline metric remains scarcely captured
structured energy.ref 314 / 459 energy lives in linked eng-* records, not inside metrics

⚑ time_to_correct_answer_s is what docs/14-model-backend-benchmark.md names as the headline — “time-to-correct-result + loops-to-done, with tool-call success as a gate” — and what the operator independently identified as the thing that actually determines whether a model is pleasant to use. It is populated once.

The consequence is visible in this session’s own reasoning: model comparisons were argued on mean_score for hours before turn counts and wall-clock were pulled out of the raw simulation files by hand. That analysis should have been a query against recorded metrics, not archaeology.


Energy coverage and limits

  • 314 runs are joined through energy: {ref: eng-*}.
  • Five runs are explicitly unjoinable through energy_unjoined_reason; this includes pre-aihydra work with no defensible wall-meter window and combined windows that cannot be apportioned without inventing data.
  • 188 energy records exist. A single invocation window may legitimately cover several run rows, so run links and energy-record counts are not expected to match.
  • An energy total is not automatically Wh/correct, Wh/token or model-only energy. Each claim must preserve the window and denominator actually measured.

Implemented controls

  • energy collection is built to §2.4, with build-time enforcement: delta_w must equal mean_w_active − baseline_w, a non-recorded window must explain its derivation, and joules are rejected outright.
  • idle_baseline collection is built to §2.5, with aihydra’s first record. Two of the three deltas the doc asks for are marked measured: false rather than estimated — notably igpu-resident-quiet, which is what prices keep_alive, and npu-resident-quiet, which is the question an always-on NPU triage loop turns on.
  • runs gained optional started_at / ended_at.
  • bench/runmeta.sh installed and wired into all seven bench scripts via an EXIT trap, so every invocation now emits its UTC window, duration, exit code and a contention flag to ~/run-meta.jsonl. Verified on a deliberately failing run — failed runs emit too, which is the case most likely to be silently dropped.

What remains

  1. Capture time_to_correct_answer_s and tool_call_success as standard, not incidentally.
  2. Measure igpu-resident-quiet — a model loaded and idle. It is one server start and ten minutes of doing nothing, and it unlocks the keep_alive cost question entirely.
  3. Build suite, so a suite name means something checkable.
  4. Decide whether workload, telemetry, decision and edge are still wanted, or whether the design has moved on. Four unbuilt record types may be a plan that outgrew itself rather than a backlog — but that should be an explicit decision, not a silent omission.