Benchmark protocol
How this test bed measures things, and what it refuses to measure. Written 2026-08-04, before the boxes are back — deliberately, because a protocol designed after the first results is contaminated by them.
The governing metric is unchanged from Phase 3C: time to a correct result, with tool-call success as a gate rather than a tiebreaker. Raw tokens/second is a means. A model that decodes faster but reasons in circles is slower where it counts.
0. Preconditions — pin these or the results mean nothing
Not axes. Get one wrong and the run measures the mistake instead of the subject.
| Precondition | Why |
|---|---|
preserve_thinking verified end-to-end |
Qwen3.5/3.6 templates discard prior-turn thinking blocks without it, and the documented failure signature is tool calls arriving with empty argument objects. Tool-selection is our stated failure mode — the same signature. An uncontrolled run measures a template bug. Acceptance test: a 5-10 turn tool loop with zero empty-arg calls, before any suite runs. |
| Chat-template version recorded per result | Template version has been a larger source of 2026 agent variance than model choice. A result without it is not comparable across dates. |
| Host health verified | Kernel ≥6.18.4 with CWSR exported, firmware not linux-firmware-20251125, no amdgpu-dkms. Below that ROCm reports wrong VGPR counts and crashes. Benchmarks on a sick host are noise. |
| Quiesce protocol | No other model resident, no background crons, llama-swap quiesced for GPU-borrow windows, page cache state known. Our own decode collapse (0.12 tok/s on a 99.4%-warm turn) was caused by memory pressure from previous activity. |
| Output sanity gate | llama-bench measures throughput and never checks the output is valid. A backend can be 24% “faster” while emitting a single garbage character — observed on gfx1151 Vulkan with Qwen3.6 (clm-0014). Before recording ANY throughput number: one real generation per (model, backend, build), eyeballed for coherence. Fail it and the throughput is discarded, not footnoted. HIP/gfx1151 chunked-input addition (2026-08-25): for future or changed low-ubatch performance work, this is not satisfied by a short generation alone. Run an exact-fingerprint chunked-input forward control (PPL/raw-logit preferred) and a real generation before admitting the number. Upstream #25863 (con-0016) reports input corruption under chunking on a different Qwen model; it triggers this prospective control, not a retroactive defect finding for any HaloBench record. |
--load-mode none on unified memory |
With mmap the model file sits in page cache while the tensors are allocated again. On a discrete GPU those are separate pools; on shared memory they are the same DIMMs, so a 73 GiB model on a 122 GiB box goes to swap or dies with thousands of amdgpu: SVM mapping failed, exceeds resident system memory limit. This defect was hit FOUR separate times before being fixed properly — each occurrence patched at one call site while the others stayed broken. It is now enforced twice: LLAMA_ARG_LOAD_MODE=none on the host as a floor, and an explicit flag in every script so the setting is visible in the fingerprint. An env var alone is invisible to a reader; a flag alone is easy to omit in a new call site. |
| Scatter is a symptom, not an error bar | Repetitions are not merely a way to average noise away. On a broken code path the same deep cell varies 5-10% between identical repetitions; on a correct one it reproduces within ~0.3% (clm-0017, where that scatter retroactively explained a 20% standard deviation that had gone unexplained for months). A cell whose repetitions disagree by more than ~3% at depth is evidence of a defective path and must be investigated before its mean is published. sweep.sh computes per-cell coefficient of variation and flags this; a clean mean over dirty samples is the most confident way to publish a wrong number. |
| Varied prompts under speculation | Draft-free speculation (ngram-mod) shares an n-gram pool across slots within one llama-server. Run identical or near-identical content on multiple slots and they feed each other’s cache: throughput reflects the pool, not the model. A community Strix Halo run (clm-0016) measured 302-305 tok/s across 4 identical streams and 430 tok/s on same-prompt repeat, and correctly discarded both as artifacts — its trustworthy sustained figure, 121 tok/s, came from 500 varied prompts. Any multi-slot run with speculation enabled must use varied prompts, and must report the no-speculation floor alongside the speculated number — the ratio is the finding; the absolute is not comparable across configs. |
| Build provenance | build_commit, backend, and runtime.tree on every row. With two forks now in play (con-0001, and possibly the quantized-KV fix) “which binary” is not a detail. |
| User-simulator identity pinned and recorded | For any simulated-user suite (τ²) the simulator is half the measurement. Self-play (agent == simulator) confounds every cross-model comparison — the twin volunteers what its counterpart needs (clm-0043, clm-0048). Pin it (claude-haiku-4.5), record it structurally (sim_self_play), and flag + exclude a self-play arm from cross-model axes — never silently average it in. |
| Output budget floored and shown non-binding | max_tokens is a capability lever wearing a serving-detail disguise: too low and it truncates tool calls and reasoning, suppressing the score. A bench-tier capability run carries max_tokens ≥ 2048; a lower cap is a screening-only choice and travels labelled as one. For a reasoning-mode comparison, pin the same budget on both arms and size it for the longest mode’s reasoning plus its response. Audit every final timing/usage row for structural cap hits, and record reasoning versus visible output separately where the endpoint exposes them. Any finish_reason=length, exact-cap token count, truncated tool call, or missing terminal response makes that cell harness-limited and ineligible for a capability comparison. Raise the cap only under a declared new config; preserve the strict fixed-budget result and report any intended-config or recovery result separately rather than pooling them. |
| Instrument can express the intended config | Before recording a performance number, confirm the tool can actually run the config the number claims to represent. llama-bench silently rejects --kairic-edge and has no speculation path; a run that drops the flag to satisfy the tool is measuring a different model. If the intended config doesn’t fit the instrument, change the instrument (§1d) — never the config. |
1. Two measurements, not one — capability and performance
Everything below rests on this split, so it comes before the axes.
Capability is what a model can do: retrieval at depth, tool calling, blind-judged reasoning. Performance is how fast it delivers that: prefill, decode, TTFT, and the cache and parallelism machinery around them. A full benchmark establishes capability, then finds the configuration that delivers the best performance without breaking it.
The two differ enormously in cost. A capability suite is hours; a llama-bench row is
seconds. That asymmetry is the whole reason to separate them — it lets capability be
measured rarely and performance often. What makes that legal is knowing precisely which
levers can break capability, and how.
1a. Levers, classified by how they touch capability
| Class | Levers | Behaviour | Consequence |
|---|---|---|---|
| neutral | speculation, mtp, ngram, cache_ram, slot_restore, disk_warmth | Provably capability-preserving | Sweep freely. Capability may be inherited. |
| binary | backend, build | Fine or catastrophic, nothing between | Cheap guard catches it in seconds |
| lossy | weight_quant, kv_quant | A real slope | Needs KL divergence + a full capability run |
| hazard | parallel, chat_template, preserve_thinking | Silently corrupts correctness | Never inheritable, never assumed safe |
Neutral is a proof, not an optimism. Correctly implemented speculative decoding
reproduces the base model’s distribution exactly — the verification step guarantees it.
So an unchanged quality score under MTP or ngram is the expected result, and mainly
evidences that the implementation is not lossy. This is why clm-0016’s IFEval-strict
78.6% under ngram-mod is reassuring about the implementation rather than surprising about
the technique; the more interesting half of that result is that it also held under
ROCmFP4, which is a lossy lever.
Hazard is the class that breaks the clean separation, and it is not hypothetical:
llama.cpp #25992 leaks responses across concurrent requests on gfx1151 HIP at
--parallel > 1. That is a throughput setting silently corrupting output. Note that
the community’s headline 121 tok/s sustained figure was measured at parallel 4 on HIP on
exactly that silicon, and an IFEval score cannot clear the concern — contamination would
surface as scattered per-prompt failures, indistinguishable from ordinary model error at
that granularity.
Capability is a function of depth, not a scalar. A model fluent at 8k can fail at
200k. Every capability result carries depth_tokens; one without it is not a result.
1b. Three tiers
- Capability suite — expensive, thorough, run per (model, quant, template).
- Guard (
bench/guard.sh) — cheap cliff-detector, run at every performance config. Coherence, a tool call with non-empty arguments, a shallow needle at depth, and cross-request isolation when parallel > 1. Minutes, not hours. - Performance sweep — run freely once the guard passes.
The guard is what makes “best performance without breaking capability” tractable. It does
not certify capability; it detects cliffs. A lossy lever cannot be cleared by a guard,
and the schema enforces that: inheriting capability across a quant or backend change is a
build error, not a judgement call.
1c. Publish the frontier, not the winner
“Best performance without breaking capability” implies a threshold, and thresholds are judgement calls that age badly. Every config is published carrying both its capability basis and its throughput, with risk flagged rather than the row suppressed — the same pattern already settled on for OOM. A schema that refuses to record a fast-but-degraded config is the same schema that refuses to record reality.
1d. The comparable instrument and the intended-config instrument are different tools
llama-bench is the floor instrument: raw, non-speculative decode/prefill, one
identical harness for every model. Its number is the cross-model comparable — and it is
only the floor. It cannot express speculation (ggml-org/llama.cpp #22947, closed
not-planned), a vendor runtime, or sidecars, and must not be coerced into pretending to.
The intended production config — MTP, a vendor fork, .pfs sidecars, a
non-autoregressive paradigm — is measured on a server-driven depth ladder:
llama-server in the deployed config, a script reading the server’s own
timings.predicted_ms/prompt_ms at each depth. It is the only instrument that turns
speculation on, pins a depth, and emits a clean tok/s together.
Rules.
- Both numbers are published, with the ratio. Floor (llama-bench, or server spec-off) and intended (server, spec/sidecars on). A speculated number without its floor is not a result — the ratio is the finding, the absolute is not comparable.
- “Best config” is the intended config, not the floor. A model whose deployable config is MTP/sidecars is not represented by its llama-bench floor. Where they diverge, the floor is labelled a floor and the row’s headline performance is the intended config. A config is “best” only once its speculation/runtime status is resolved — assessed-and-better, assessed-no-gain, or architecturally-impossible — never merely “the part the floor instrument could reach.”
- The server sweep obeys every anti-inflation rule. Varied, realistic prompts
(repetitive input inflates draft acceptance — the HK-001 126 t/s echo;
clm-0016), the same prompt set across spec-on/spec-off, server timings not wall-clock, and a couple of nearby depths per rung (MTP acceptance collapses at KV-slot boundaries).
This resolves the §1a/§2 inconsistency: MTP stays a neutral capability lever (output
identical, sweep freely for quality and inherit it), but it is a first-class performance
lever measured on the server, never on llama-bench. §2’s “quant × MTP as a ratio” is
therefore a server measurement, and always so.
1e. Novel models — never null, never forced-misrepresentation
A model that needs a custom runtime, sidecars, or a non-autoregressive paradigm is neither rejected for not fitting the harness nor forced into a harness that misrepresents it. It is onboarded by degrading gracefully:
- Run the intended config (validity), on the instrument that can express it (§1d).
- Enforce the comparability invariants that can be held (§2a — simulator, task set, output budget, no repetitive/self-play tricks).
- Compare on paradigm-neutral outcomes — reward, time-to-correct, Wh-per-correct — never on a mechanism the paradigm doesn’t share.
- Disclose the intrinsic rest.
LLaDA2.2-flash is the worked precedent. Block-diffusion decode tok/s is not comparable to autoregressive by construction, so it entered the field on Wh-per-correct and reward — a valid, honest entry from a model the standard throughput harness cannot describe. The screen → bench gate is the onboarding ramp: a novel model gets a fast, provisional read in intended config at the screen (5-task, marked smoke), and graduates to the citable bench once the invariants are enforced over the full task set. For a speculation-capable model the screen is run twice — spec-on and spec-off — which is enough to establish working speed with and without speculation before committing the full arm.
2. Axes — and why full factorial is the wrong instinct
Five axes interact: backend × context × quant × KV-quant × MTP. Full factorial is hundreds of runs, most of them uninformative.
Sample deliberately instead:
- Backend × context is the primary grid. Not a fixed choice — ROCm was picked in
June because Vulkan
vk::DeviceLost-crashed at 128K, but the community quantized-KV work is Vulkan-side for prefill and ROCm-side for decode. The right backend may now differ by phase and by context, which is exactly the kind of thing an inherited assumption hides. - Quant × MTP must be measured together, as a ratio. Published Strix Halo numbers show the MTP multiplier moving opposite to intuition — 2.44x at Q8_0 against 1.81x at Q4_K_M — because decode is bandwidth-saturated, so fewer memory passes pay more the heavier the weights. Measuring either alone gives the wrong answer about both. The transferable number is the ratio, not the absolute.
- KV-quant is now a first-class axis, not a memory-saving footnote. Community
verification reports stock llama.cpp repeatedly dequantizing during inference on this
silicon, with a fix worth +75% to +203% ROCm decode and q8_0 running faster than
f16. If true here, “quantize the cache to save memory” becomes “quantize the cache to
go faster”, which inverts the trade. See
clm-0006. - Cold vs warm is a dimension, not noise. Every row carries
served: warm|cold. A suite that measures only cold prefill misrepresents production; one that measures only warm hides the cold-start cost that dominates a first turn.
2a. Capability config — pinned, disclosed, guarded
Capability has as many knobs as performance and the same tension: pin everything and you misrepresent models that legitimately differ; pin nothing and you destroy comparability. Sort the knobs.
| Class | Knobs | Handling |
|---|---|---|
| Comparability invariant — pinned, guarded | simulator identity, task set + seed + trials, output budget, scoring method | Held constant across the field. A run that breaks one is flagged and excluded from cross-model axes — structurally, like a guard failure, not by footnote. |
| Model-intrinsic — disclosed, never normalized | thinking/reasoning mode, recommended sampling, quant, runtime/backend | Part of the model’s identity. Forcing them uniform yields an invalid number, not a comparable one — a thinking model forced silent, a model held off its recommended temperature. Run as intended; record on the metric card. |
| Speed-only — floored, not on the quality axis | speculation/MTP, batching, backend fp noise | Do not move the reward (speculation is distribution-exact). Comparable on capability regardless; on the cost axis they are marked or floored. (The qwen38-27b τ² arm ran with MTP while the field ran plain — its reward is comparable, its time-to-correct carries the spec marker.) |
The consequence — compare on outcomes, not mechanisms. Reward, time-to-correct, and Wh-per-correct survive any intrinsic-config difference. Tok/s, MTP acceptance, and decode rate do not — they are per-model diagnostics with floors, never the cross-model ranking axis.
3. Tiers — cheap and broad first, expensive and narrow last
Tier 0 — capability probe. Does it load, what is the max fit context, what is peak
memory. Minutes. Kills candidates before they consume eval time — and produces
observed_peak_gb, the field clm-0001 says we are missing. Capture peak during
Tier 0 rather than as a separate exercise; it is nearly free here and expensive later.
Tier 1 — throughput sweep. TTFT, prefill and decode tok/s across 4K → 128K → 262K, per backend. Throughput on this hardware is deterministic (N=3 stddev ≈ 0), so N=3 at ≤32K as published variance evidence, N=1 above where the number does not move.
Depth ladder, extended to production/max context (added 2026-08-17). The house
throughput ladder was d0 / d32768 / d131072 (N=3 at d0/d32768, N=1 at d131072). It now
extends to d0 / d32768 / d131072 / d204800 / d<model-max> where fit allows —
d204800 because that is the depth band this lab actually operates production turns
in, not a round number, and d<model-max> because a model’s advertised context window
is a claim until something has actually been pushed to it. The three deepest cells
(d131072, d204800 and d<model-max>) run at N=1 — at this depth the question is
“does it complete, and roughly how fast,” not variance, and scatter of dirty repetitions
from a run that already costs tens of minutes is not worth tripling the cost to detect.
The two shallower cells keep N=3. Additive, not a retcon: every historical three-cell
matrix (d0/d32768/d131072) remains a valid, comparable result on its own terms — a
model without deep cells is a model not yet backfilled, not a model with an invalid
result. d<model-max> is genuinely per-model (qwen38-27b/ornith-35b/
qwen36-27b-mtp top out at 262144; nemotron3-super/laguna-s-21 claim a
YaRN-scaled 1048576 that will very likely fail the memory-fit check before it fails
anything else — record the fit-projection refusal as a result, not a silent skip).
Tier 2 — cache capability trace. The operator’s addition, and the part nobody else publishes (§3).
Tier 3 — agentic quality. Expensive, so run it only on the configuration Tier 1-2 selected, at the context band we actually operate in (>100K). For blind-judged single-response tasks, N≥3 minimum, because this is the only place run-to-run variance is real: use a separate, fresh judge per response seeing only the task, the rubric and one anonymised answer.
τ² Airline full-domain house convention — one complete pass, not three repeats. The
house command shape runs tasks 0-25 once with num_trials=1, a fixed seed and the pinned
independent simulator. Scoring mode is part of the configuration: deterministic judge-null
and separately labelled LLM-judged arms are distinct evidence and must not be pooled as if
they used one scorer. This is the established HaloBench method, not a new exception or a
relaxation from N=3; the generic blind-judged rule above was simply too broadly worded. No
retained full-domain record declares more than one trial, although an omitted trial field is
an evidence gap rather than proof of its value. Here N=1 means one trial for each of 26
varied tasks, not one trajectory or one judged response. Report the exact task count, seed,
simulator, scorer/judge mode, harness revision and output budget, retain every trajectory,
and make no run-to-run variance or confidence-interval claim. A Low/XHigh comparison is two
paired configurations, not two repetitions: use the same 26 tasks, seed, simulator,
scorer/judge mode, harness and demonstrably non-binding output budget. Repeat the full
domain only for an infrastructure-invalid run or a declared new hypothesis. This
native-tool domain arm runs at its explicitly pinned serving context and does not
establish >100K agentic-context quality; that requires separate long-context evidence.
The R1 rule. A prescriptive scorer once ranked the eventual winner 0/15 while it was producing the highest-value output in the field — the gate was measuring entity-name luck, not reasoning. No pass/fail gate is trusted until transcripts have been read. Scores propose; transcripts decide.
4. Cache capability trace — the original contribution
Everyone publishes tokens/second. Nobody publishes what their cache actually does, and on this box cache behaviour is the difference between an 8-minute turn and a 10-second one. A complete trace per config:
| Capability | What to record | Why it matters |
|---|---|---|
| Slots | --parallel n, per-slot context, whether MTP forces -np 1 |
Doubling slots doubles KV. Almost never stated publicly, and it changes every memory figure. |
| In-slot prefix reuse | cache_n / prompt_n on a repeat turn |
The 97% figure that makes the warm lane work. Measure at realistic bootstrap scale — a small-prefix test gave a misleading 40% and cost a day. |
| Cross-session reuse | reuse when the prefix differs after the system prompt | Known to fail: isolated sessions diverge immediately and cold-prefill. Worth measuring, not assuming. |
--cache-ram |
size, hit rate, and its memory cost | A host-side cache that competes with the model for the same pool. |
| Disk save/restore | restore wall time, tokens reused, file + sidecar size | ~105 ms restore, 3017 tokens reused with the checkpoint sidecar and 0 without — the whole reason con-0001 exists. |
| Checkpoint count | --ctx-checkpoints, resident cost, save-spike size |
Defaulting to 32 was ~4.5 GB resident and a contributing cause of the OOM panic. |
| Page-cache dependence | restore time cold vs warm page cache | Our 105 ms restore probably relies on the file being cached. A real NVMe read is ~190 ms. Unmeasured. |
This is also where the cache-wall hypothesis (clm-0007) gets tested: a measured
GPU-cache knee at 32-40 MB with a ~4x read-speed drop past it, offered as a candidate
explanation for our 354 → 144 tok/s prefill curve. Distinct mechanism from pool
utilisation — do not conflate the two.
5. What every run must carry
Non-negotiable, or the result is not publishable: config id (never a naked number),
build_commit, backend, runtime.tree, chat-template version, served: warm|cold,
N and the variance where N>1, and the suite version. Plus incidents: [] — a run
during a thermal event or a competing process is contaminated, and that must be visible
at the point of use rather than buried in a postmortem.
Suite versioning: patch = comparable, minor = comparable with a note, major = breaks
the series — stated explicitly in comparable_with, never inferred from the number.
The grader is a dependency: changing the judge model silently re-baselines all
history, so it is a major bump even when the test content is identical.
6. Cross-host runs — the comparison nobody has
With two boxes, run the same suite and config on both and set cross_host_pair. That
isolates hardware as the only variable. It is the comparison people actually want when
deciding whether a second box is worth buying, and it is almost never available because
most people replace rather than overlap. Run it during the overlap window or lose it.
7. Build on existing work — what to adopt rather than invent
Checked before writing anything bespoke. Most of Tiers 0-1 already exist and are better standardised than anything we would write.
| Tier | Adopt | Why |
|---|---|---|
| 1 — throughput | llama-bench -o json — gated on the output sanity check above |
Its JSON already carries build_commit, build_number, gpu_info, model_*, n_batch, n_ubatch, n_depth, type_k/type_v, n_gpu_layers, flash_attn, avg_ts, stddev_ts and the per-repetition samples — i.e. almost exactly our required provenance set, including the context-depth and KV-quant axes. Use it, do not reimplement it. It is the floor instrument, not the ceiling — see §1d. |
| 1 — served path | llama-benchy — adopt as-is |
llama-bench measures internal C++ timing and cannot see proxy overhead, queueing or warm-lane restore. llama-benchy (MIT, actively maintained) measures the client-perceived side against any OpenAI-compatible endpoint and deliberately mirrors llama-bench’s output shape so the two sit side by side. Its --enable-prefix-caching mode is the exact primitive we need: a two-step measurement separating the cold context load from the cached-prefix reuse. Run both, publish the gap — that difference is a finding, not an inconvenience. Two caveats: it speaks only /v1, not llama.cpp’s Anthropic-compatible /v1/messages, so the Claude Code path needs a thin second script; and client-side timing conflates its own precision with our proxy hops, so sanity-check a few readings against the proxy’s own /_stats before publishing any delta. Effort: hours. |
| 1 — determinism | temperature 0, fixed seed — but do not rely on it | The homebench convention, and worth setting. ⚠ It does not hold on Vulkan (clm-0011): same config, temp 0, batch 1, no speculation, three different outputs. Two consequences: one trajectory cannot establish output invariance, and speculative losslessness cannot be verified by diffing outputs — which is the obvious way one would check that MTP or ngram speculation changes nothing. Use a distributional instrument (KL divergence) instead. This does not turn the explicitly bounded τ² full-domain N=1 convention above into an invariance claim: that arm is a 26-task capability observation with no run-to-run variance claim. Unverified on ROCm; check rather than assume it transfers. |
Single-trajectory exactness is not invariance. Upstream #25618 now has a
Strix Halo Vulkan report (con-0004, distilled as clm-0094) where the original
F16-V prose control reproduces byte-exactly, but five varied prompts under the
same target/draft/runtime/template/completion path produce 0/5 exact-token parity,
and a separate fixed-sampler replay of those same prompts also stays 0/5. The
protocol consequence is narrow but important: one exact MTP trajectory is real
evidence for that trajectory, not a proof that the target path and speculative
path are generally invariant. Any future MTP-losslessness dossier must carry
varied prompt fixtures and raw token/probability evidence, not only one passing
prompt.
Spec-specificity check (upstream #26750, clm-0102). Upstream llama.cpp
#26750 provides a second independent reason the MTP-losslessness dossier must
keep the varied-prompt/raw-token requirement: the draft-mtp acceptance rate is
spec-path- and prompt-dependent in a way that is byte-reproducible. On a
CUDA/GB10 box (Grace Blackwell, aarch64, 128 GB unified) a same-box build pair
moved prose acceptance from 84.0% to 51.3% (b10261 to b10344, -26.8% wall),
and later still 66.67% to 52.38% at b10532, while with speculative decoding
fully disabled the two builds are byte-identical in raw llama-bench prefill
and decode (+0.24%/tg256 +0.26%, flat within noise) — i.e. the entire loss
lives in the speculative path, not in general engine or kernel drift. The
collapse is strongly prompt-specific: enumerable, highly-predictable output
masks it entirely, so any regression suite built around canned enumerable
output will miss it. This is upstream context, not a HaloBench benchmark gain
(our surface is ROCm/Vulkan aihydra, not CUDA), and it leaves the protocol
requirement intact: any MTP-losslessness or speculation-correctness dossier
here must carry varied prompts and raw token evidence, and separate
spec-on/spec-off baselines.
Drafter-independence + KV-lossy axis (upstream #25618, clm-0104). A
2026-08-22 #25618 follow-up on RDNA4 (R9700/gfx1201, Vulkan, Qwen3.8-27B
Q6_K_XL, greedy) adds a third upstream confirmation, with two consequences.
First, the spec-path output divergence is DRAFTER-INDEPENDENT — the code
prompt diverges from a --spec-type none baseline at the same first-diff
byte 204 for the built-in MTP head and for an external DFlash2 drafter at
Q4_K_M / Q8_0 / BF16, with identical acceptance (draft 347, accepted 283) and
byte-identical DFlash2 output across drafter quants. So drafter precision is
NOT a fixable variable: switching or quantizing the drafter will not restore
parity. Any local MTP-correctness check must therefore hold the verify/target
side and the KV-quant axis the loss sources, not the drafter. Second,
KV-cache quantization itself is output-moving even with NO speculation: a
reasoning prompt lossless under f16 KV diverges at byte 155 under q8_0 KV,
and --spec-type none differs between f16 and q8_0 KV at byte 561, so
q8_0 KV is never output-preserving and compounds the spec-path divergence.
This corroborates the output-parity caveat on our q8_0-KV adoption
(see Qwen3.5-122B / HO-003, clm-0103) WITHOUT invalidating that
capability/energy conclusion — the tau2 tasks recovered and the 29.92 vs
42.06 Wh/correct energy join were measured directly and are unchanged by this
upstream context. Reported surface is RDNA4, not our gfx1151 surface, so this
is upstream context only, and the house varied-prompt/raw-token requirement
stands intact.
Rollback-exactness on hybrid GDNs — v0.6.9 tradeoff and its v0.6.10 reversal
(fork clm-0105/clm-0107, con-0007/con-0009). A fourth arm of the
dossier comes not from an upstream issue but from a stable fork release.
v0.6.8 introduced a one-line change forcing MTP rollback through full
sequence-state checkpoints; v0.6.9 (2026-08-22T04:09:43Z) reverted it because
on Vulkan with a hybrid GDN target (qwen35moe, e.g. Qwen3.6-35B-A3B) that path
deadlocked deterministically a few hundred tokens into a long response
(futex_do_wait, GPU idle), returning rollback to fast snapshot-plane restore
that the vendor stated “can diverge slightly from a no-draft run after a
rejected draft” — a deliberate availability-vs-exactness tradeoff. That
tradeoff was REVERSED in v0.6.10 (2026-08-22T12:03:16Z): the deadlock’s root
cause was the server RE-VERIFYING replayed draft tokens after a checkpoint
restore (livelocking the slot), not the state-save path; with “server: do not
re-verify replayed draft tokens after a checkpoint restore” (9c5d899) and
full-checkpoint MTP rollback re-applied (f25eefe), rollback on hybrid GDN
targets is token-exact again. So a future MTP-LOSSLESSNESS claim on the current
stable fork should no longer assume a deliberately-diverging rollback path
after a rejected draft. This does NOT relax the house varied-prompt/raw-token
MTP-invariance requirement: the v0.6.9 tradeoff was a fork-rollback-path issue,
whereas the upstream #25618 quantized-target divergence (clm-0094/clm-0104/
clm-0106) and #26750 spec-path collapse (clm-0102) are separate verify/target-
side issues that still stand. Neither does it change any HO-009 comparison
boundary: the Qwen3.6-35B-A3B-MTP plain-vs-n2 results (clm-0097, clm-0101, and
the Qwen3.5-122B incumbent of the clm-0103 series) were measured on UPSTREAM
ggml-org/llama.cpp 7077abb on ROCm0/gfx1151, which never used the fork’s
snapshot-plane rollback path, so this tradeoff never directly applied to those
numbers. All four are fork-vendor release-note records (single-payload
verification, speed figures not BENCHMARKS.md-protocol); nothing inherits into
a HaloBench number, and the house varied-prompt/raw-token MTP-invariance
requirement stands intact.
Quantized-target divergence is now corroborated across unrelated model
families, OSes and GPUs (upstream #25618, clm-0106). A fifth arm of the
dossier comes from an independent non-Qwen3.8, non-ROCm/AMD report: on
NVIDIA Vulkan/Windows 11 (llama.cpp b10566 / bb4caa754, GTX 1660 SUPER)
the official LiquidAI LFM2.5-2.6B pair diverges from vanilla on a Q8_0
target, while the SAME Q8_0 draft against an F16 target produces
byte-identical greedy output to the F16 vanilla target (40/64 draft tokens
accepted). The differential control isolates target quantization as the
causal variable: the Q8_0 target consistently differs from vanilla but
loading the draft with --spec-draft-p-min 1 (zero draft tokens) restores
byte-exact match, and the mismatch persists at --spec-draft-n-max 1 with
target KV in F16 (neither a larger speculative block nor quantized target
KV). It also diverges at n_max=1, unlike the Qwen3 boundary earlier in
#25618 where n_max=1 was reported lossless — so the narrow n_max=1 boundary
is not a general safety property across targets. It does NOT change the
house varied-prompt/raw-token MTP-invariance requirement; it strengthens it.
The drift is now evidenced on multiple unrelated families, so the target
quantization axis is a live loss source regardless of drafter or model
family, and a local varied-prompt/raw-token control remains required before
any baseline/MTP invariance claim on our surface. Reported surface is
NVIDIA Vulkan/Windows, not our ROCm/Vulkan Strix Halo gfx1151 aihydra box,
so this is upstream context only and inherits nothing into a HaloBench number.
Draft-depth ceiling + CUDA multi-GPU MTP lockup (upstream #27122 clm-0114 /
con-0010; adaptive PR #27210 clm-0115 / con-0011). Two more upstream
corroborations reinforce the standing shallow/adaptive draft-depth guidance and
the “no deep blind sweep” stance; neither changes a comparison boundary or any
measured number here.
-
A high fixed n_max CEILING carries a fixed cost, independent of adaptive logic (llama.cpp PR #27210,
clm-0115): an author held draft depth at fixed 3 while keeping max depth 10 and measured merely having max depth 10 as a fixed ~2.6% performance penalty; the follow-on fix (limit the full MTP buffer scan on truncated/short drafts) recovers ~2-3% for--spec-draft-n-max 10 --spec-draft-p-min >0.5configs. An independent RTX 5090 run confirms depth changing does not slow CUDA-graph decode (GGML_CUDA_DISABLE_GRAPHS deltas small) and that adaptive 3..10 + p-min is the only config with a positive recall delta (+2.6%) — so the cost is the static ceiling, not adaptive switching. This meaning is exactly the standing rule: prefer a shallow-plus-adaptive pairing (open question whether the high ceiling pays for deep-climb on dense-target gains) and treat a deep static n_max as a dead cost. -
The MTP spec-path is fragile under multi-GPU tensor split (llama.cpp #27122,
clm-0114): a CUDA 4x-RTX-A4000 build hard-crashes 6/6 on a Qwen3.8 MTP deep-context prefill under--split-mode tensorwhile the identical workload on 2 GPUs (internal 2-device AllReduce path) is stable with MTP on;LLAMA_GRAPH_REUSE_DISABLE=1is a validated workaround consistent with PR #24549’s graph-reuse mechanism. Crash frequency scales with n-max. This is CUDA multi-GPU; our surface is single-GPU ROCm/Vulkan gfx1151, so it changes no HaloBench boundary — but it is a further independent confirmation that the spec-path, not the engine, is where MTP fragility lives.
Both arms are upstream community/PR context only; nothing inherits into a
HaloBench number, and the house varied-prompt/raw-token MTP-invariance
requirement (clm-0102/clm-0104/clm-0106/clm-0114/clm-0115) stands intact. The
standing shallow/adaptive + no-deep-blind-sweep guidance is unchanged.
| 3 — agentic | — build a minimal one ourselves | Checked 2026-08-05: there is no public repo. The task set, hidden graders and traces are deliberately private to keep them out of training corpora, and the work carries no OSS licence — so it cannot be adopted, only requested. (Beware a name collision: harness-benchQihoo360/harness-bench is an unrelated academic benchmark.) Its shape is still the right one and worth copying: 5-8 tasks, isolated workspace/ per cell, grading by a hidden test.sh the agent never sees. Its author also reports a real contamination case — one harness read the hidden tests in 14 runs — which is a fairness trap to design against from the start. Our question is narrower than his anyway: he compares models across harnesses; we want raw-engine versus proxy-fronted behaviour for a fixed model. Estimated 1-2 days. |
| 3 — token economics | end-to-end token accounting | Current work shows the harness — how context is assembled and tools exposed — dominates token cost, with large reductions available from how tools are surfaced rather than which model is used. Relevant directly: our own tool-schema curation cut ~17.5K tokens from every prompt. Measure tokens per completed task, not just tokens per second. |
The interoperability prize. Emitting llama-bench-compatible JSON makes our numbers
directly comparable to everyone else’s on this silicon, and directly contributable —
the standing community request is literally “reply with the numbers, with build_commit
and the fa mode from the bench JSON attached”. A bespoke format would forfeit that for
no gain.
What stays ours, deliberately: the cache capability trace (§3) — nobody publishes it, and on this box it is the difference between an 8-minute turn and a 10-second one; energy per correct answer, wall-metered with a stated baseline; lived production telemetry alongside controlled benchmarks, kept strictly separate; and the decision log — what displaced what, and what was rejected.
8. Task sets — adopt the standard ones, keep the bespoke ones for a reason
Distinct from §6, which is about plumbing. This is about content: which task sets to run. The rule is the same one the site is built on — the scores in the public commons are unreliable; the tasks are not. Adopting a task set while generating our own numbers against our own config fingerprint is exactly the point. It is the missing fingerprint we criticise, not the questions.
| Capability | Adopt | Notes |
|---|---|---|
| Long-context recall | lm-eval’s RULER port (NOT NVIDIA/RULER) + llama-perplexity --kl-divergence |
Four task families where NIAH has one. But verified 2026-08-05, and the plan changed: NVIDIA/RULER’s standalone pipeline is deprecated in its own README, will not import without NeMo, and its OpenAI client is hardcoded (no base_url, and a model2length lookup that KeyErrors on any local model name). It also silently dropped answer_prefix from every prompt between Jan 2025 and 21 Jul 2026 — scores from that window are not comparable to anything. Use lm-evaluation-harness’s RULER port instead: all 13 tasks, generated on the fly, points at any OpenAI endpoint with no code changes. ~2-4 hours.⚠ RULER cannot answer the question we wanted it to answer, on this hardware. Published KV-quant effect sizes are ~1 point (NVFP4 on Ruler-64K: 95.6 / 95.5 / 94.6 for BF16 / FP8 / NVFP4; xKV 4-bit: 88.85 → 87.64). At n=25 the binomial SE near p=0.9 is ~6pp; at n=125, ~2.7pp. Any sample count we can afford resolves 5-10pp, not 1pp. And the cost is brutal: the maintainer quotes ~2 h for 128K × 500 samples on 8×H100 with vLLM batching; we have one slot, strictly serial, no batching, and every RULER sample is a distinct long prompt so prefix caching buys nothing. Full suite ≈ 6,500 sequential 128K prefills per KV config — weeks. So use two instruments. Primary: llama-perplexity --kl-divergence, built into llama.cpp — record f16-KV logits once, replay q8_0 and q4_0 against them, and get mean KLD, ΔPPL, Δp percentiles and top-1 agreement with uncertainty bars, over hundreds of thousands of tokens, in a single forward pass with no generation. Prefill is the fast half of this box, so it is the cheapest signal-per-hour available. Secondary: a RULER-lite cliff detector — ruler_vt, ruler_cwe, ruler_fwe, niah_multikey_3 at two lengths, 30-50 samples. Those four are the KV-sensitive ones; niah_single_* sits near ceiling and absorbs noise. Decision rule: if KLD(q4_0) is the same order as KLD(q8_0) and no cliff appears, ship q4_0; if KLD is 10×+ or vt/cwe fall away, we have the answer without needing resolution RULER could never give us. |
| Tool calling | τ²-bench — adopt. MCP-Bench — do NOT adopt the harness | τ²-bench (verified 2026-08-05): MIT, genuinely active (v1.0.1 Jul 2026, pushed 4 Aug), and — checked in source, not docs — it calls LiteLLM with a real tools= array, i.e. native function calling, not a text-prompted imitation. Scoring is mostly deterministic (DB state-hash diff, tool-call match against a reference trajectory, substring checks) with one LLM-judged component whose model is overridable — so hold both the judge AND --user-llm fixed across comparisons or you re-baseline silently. Effort ~2-4 hours. Domains are small: 14 / 17 / 13 tools.MCP-Bench — rejected, on a structural finding. It does not use the native tools= parameter at all: it serialises each tool’s schema as text in the prompt and parses a JSON plan itself. So running it against llama-server never touches --jinja’s grammar-constrained path, and it cannot reproduce or validate a fix for our HTTP-400 grammar-size ceiling — that failure only occurs when a real tools=[...] array is sent. It is also unmaintained (no commits since Oct 2025, no releases), has no LICENSE file at all, and hardcodes o4-mini as judge, so it needs an OpenAI key even for an all-local test. |
| Function-calling breadth | BFCL v4 — with care | Broader agentic categories now (web search, memory read/write, format sensitivity). Two cautions: v4 scores are not yet broadly published, and v3 and v4 are not comparable — mixing them silently is the exact failure this site criticises. Record the version in the suite id or do not run it. |
| Knowledge / reasoning | MMLU-Pro, GPQA — deliberately deprioritised | Not because they are bad, but because they do not discriminate on our decision axis. Every candidate that reaches our shortlist is competent at recall; none of them fail on knowledge. Chasing these would spend eval time on a tier that has never changed a decision here. Stated so the omission is a choice, not an oversight. |
| Creativity / judgement | stays bespoke — C1-C8 | No standard set covers what Warden is actually for: proactive insight, sensor-conflict resolution, ambiguous intent, restraint under an attention budget, emotional attunement. These were written against the real job and blind-judged. Keeping them is not NIH — it is that the standard sets do not test the thing. |
| Coding / agent harness | harness-bench shape (§6) | Sandboxed cells, hidden-test grading. |
⚑ Neither benchmark tests our actual failure mode, for opposite reasons. τ²-bench
uses native tool-calling but its domains are far too small (13-17 tools) to approach the
ceiling; MCP-Bench has large authentic schemas but bypasses native tool-calling entirely.
So the grammar-ceiling probe stays bespoke — and the cheap win is to lift MCP-Bench’s
mcp_servers/ corpus as raw fixture data (28 real third-party MCP schemas) and feed
combinations of them into a native tools=[...] request against llama-server, scaling
until HTTP 400. That extends the manual probe from 2026-07-05 (55 tools = 200, 60 = 400)
with authentic schema diversity, for about an hour’s work and none of their machinery.
⚠⚠ VERIFY BEFORE SPENDING A NIGHT ON THIS — the hybrid confound. Our 122B is a
hybrid/recurrent architecture; that is precisely why the checkpoint bug (con-0001)
existed at all. On such models -ctk/-ctv only affect the minority of layers that use
full attention — the linear-attention layers hold no KV cache to quantise. llama.cpp
issue #21385 reports q4_0 KV token-identical to f16 (BLEU 1.000) on a Qwen3.5 hybrid,
attributed to only 8 of 32 layers using full attention. That both predicts a tiny delta
and dilutes any measurement of it. Count how many layers actually hold a quantised KV
cache before designing the matrix — if it is a small minority, the honest finding is
“q4_0 is near-free on this architecture, and that is an architecture result, not a
quantisation result.”
Two further traps: on ROCm only matching -ctk/-ctv types hit the fused
flash-attention kernel — mismatched types fall back silently to a slower path, so keep
K and V identical or you are measuring something else. And the f16 arm may not fit at
128-200K beside a 122B; if f16 only fits at 64K, the three arms are not comparable at the
top length and the matrix must be redesigned around what actually loads.
Operational note: run all of this against the raw llama-server, never through
llama-swap or the slotpin proxy. A multi-hour sweep will evict the production warm-lane
cache and wedge the single slot.
The compounding benefit: RULER at three KV quantisation levels answers two questions at once — our own “is q8_0 safe at 200K”, and the community’s open “what does cache compression cost in quality”. Same runs, two audiences, and the second one is a standing request from people who published their half of it.
8a. Live Home Assistant: schemas yes, actuation no
No benchmark touches a live Home Assistant. Not as a safety rule but a measurement one: HA state is mutable and shared, so a suite that toggles real helpers changes the world it is measuring and cannot be re-run from the same starting state. It would also silently import HA version, MCP server version, integration availability and entity naming into every result as uncontrolled variables.
We have already paid for this lesson. The R1 task in Phase 3C appeared to measure
energy-dispatch reasoning. Reading 60 transcripts showed it measured whether a model
happened to toggle an entity literally named switch.dishwasher and whether the mock
recognised the entity IDs it guessed — zero genuine reasoning failures across the set,
and it nearly selected the wrong model. Environment coupling is how a benchmark starts
scoring the environment.
But schemas are not actuation, and one probe genuinely needs the real ones. The grammar ceiling is a schema-size limit; synthetic stubs are shorter and more shallowly nested than real MCP tools and would flatter the result. So:
Capture the schemas, freeze them, mock the execution.
Export the tool definitions from the HA and Apple MCP servers once, commit them as static
JSON fixtures, and replay them into a real tools=[...] array (bench/grammar-ceiling.sh --schemas <dir>). Authentic size and shape, no live dependency, repeatable — and
versioned, so a later HA upgrade appears as a deliberate fixture change rather than as
silent drift in a result nobody re-derived.
If HA-shaped tasks are wanted, mock the environment and treat the mock as a suite dependency: versioned, frozen, and bumped like the grader. R1’s real lesson was not “do not mock” but that the mock’s fidelity became the thing being scored.
Where live HA does belong: telemetry, not benchmarks. §2.7 already separates them — benchmarks are controlled and reproducible, production telemetry is observed and workload-dependent, and mixing them destroys both. Real Warden-touching-real-HA behaviour is a counters-and-distributions question.
9. When a run fails — which levers are safe, and which invalidate the result
Benchmarks fail mid-suite. The dangerous moment is not the failure, it is the twenty seconds afterwards when you change one flag to make it complete. Some changes cost nothing but a footnote; others silently produce a number that cannot be compared to anything, including our own earlier runs.
The rule
If the change would alter the config fingerprint, it is a NEW CONFIG, not a retry.
That is not a metaphor. The fingerprint is computed over exactly these fields —
build_commit, model_filename, type_k, type_v, n_gpu_layers, flash_attn,
n_batch, n_ubatch (see scripts/ingest-llama-bench.mjs). The fingerprint fields
ARE the invalidation boundary. Touch one to rescue a run and you have not rescued it;
you have started a different experiment that happens to share a suite name.
Safe levers — declare them, then carry on
These change how much we measure, not what:
| Lever | Cost | Condition |
|---|---|---|
Reduce repeats (-r) |
Wider variance | State N and the stddev. N=1 is legitimate where the metric is deterministic and for the established τ² Airline full-domain convention above; it is not legitimate for a blind-judged single-response claim. |
| Drop the deepest cell | Shorter series | The remaining cells stay comparable. Log what was dropped — silent truncation reads as “we covered everything”. |
| Raise timeouts | None | A cold prefill on the 122B is minutes. A timeout is a harness limit, not a property of the model. |
| Quiesce and re-run | Time | Removes contamination without touching the config. Always the first move, not the last. |
| Increase warmup runs | Time | Affects variance, not the metric, provided it is stated. |
| Split the matrix across sessions | Bookkeeping | Same config, same build. Record the dates; check for drift if the gap is long. |
Invalidating levers — a new config, or nothing
Reach for any of these and the old and new numbers must never appear in the same series:
Fingerprint fields: quant · KV quant (-ctk/-ctv) · -ngl / offload split ·
flash attention · batch and ubatch · build commit · backend · model file or revision.
Beyond the fingerprint, equally fatal:
- The grader model. Changing it silently re-baselines every judged result in history. A major suite bump even if the task content is byte-identical.
- The task set. Any edit is a suite version bump; whether it stays comparable is
stated in
comparable_with, never inferred from the number. preserve_thinking. A precondition, not a setting. Flip it and agentic results measure a template defect.- System prompt or bootstrap. Changes the prefix, which invalidates every cache-behaviour metric outright — reuse is measured against that prefix.
- MTP on/off. A different decode path.
The grey zone — permitted, but the caveat travels with the number
Sometimes the only way to measure a thing at all is to compromise. That is allowed. What is not allowed is letting the compromise sit in a footnote while the number travels alone.
| Situation | Verdict |
|---|---|
| Higher quant does not fit, so test lower | Permitted, not quant-equivalent. This is the Nemotron case: it was forced to IQ4_XS while peers ran Q4_K-class, and came last. The caveat is a first-class claim (clm-0005) precisely so it cannot be separated from the result. |
| Reduce context to make it fit | Permitted; not comparable to a full-context result. A 32K number is not a small 128K number. |
| Another model resident during the run | contention: true. A real-world figure, not a clean per-engine one. |
| Warm vs cold | Both valid, never mixed. Every row carries served. |
| Thermal throttling suspected | Do not publish the number. Log an incident, link it via run.incidents, re-run cold. |
When to stop rather than push through
Three failures where continuing produces confident nonsense:
- Empty-argument tool calls. Stop. That is the
preserve_thinkingsignature, and every agentic number collected past it measures the template, not the model. - HTTP 400 on grammar build. Our tool surface exceeded the schema-size ceiling and the fallback is silent — the run appears to succeed against a different model entirely. Fix the tool count, do not record the run.
- OOM anywhere in the sweep. On unified memory this takes the whole box down and
needs physical access. Stop the matrix, record the ratio at failure — that is the
calibration data
clm-0001says we lack — and only then decide what to change.
10. What we deliberately do not measure
- Completeness. Decision-relevance is the bar: what is live, what displaced what, what was rejected. A run that cannot change a decision does not need to exist.
- Correctness in production. It is not observable there. Production yields throughput, cache behaviour and cost — not quality. Say so on the page rather than implying otherwise.
- Output quality under KV quantisation — currently. This is a known gap in the community work too, and the one place we already hold an answer worth contributing: NIAH at q8_0 KV, perfect retention to ~205K including the multi-needle tier.
Compute surfaces and the NPU lane (added 2026-08-13)
The NPU lane follows the SAME benchmark format as GPU/CPU work — same tiers, same record types, same energy discipline. Screening tier for NPU candidates: FIT = loads-and-serves under FastFlowLM on the NPU (device evidence required: NPU lock, power-state transition); GUARD = the standard probes over FLM’s OpenAI-compatible endpoint; SMOKE = the standard 5-task tau2 set, pinned simulator, cliff-detector semantics unchanged.
Marking rule (operator, 2026-08-13): every result is labelled with its compute surface
and backend (config engine + runtime.backend). Cross-surface comparison of the
same model is welcome IN COMMENTARY — “NPU prefill is x% faster, decode slower” is
exactly the interesting contrast — but purpose-fit verdicts are per-surface, and no
table mixes surfaces without an explicit surface column. NPU-lane candidates never
inherit GPU metrics (schema-enforced).
11. Energy join is part of publishing (added 2026-08-15)
A bench suite is not publishable until its energy joins are computed — or explicitly recorded as unjoinable, with a reason. This is not a nice-to-have appended after the numbers land: wall-metered energy is this lab’s stated differentiator, and a bench that ships without it is missing the one figure nobody else publishes.
The incident this closes. The qwen38-27b screen (run-0245..0271) published
on 2026-08-14 with a full run matrix and zero energy joins — every run carried
energy: null. The HA windows existed the entire time in
aihydra:~/run-meta.jsonl; the join was simply never run before the page went
live. It was caught and backfilled the next day (eng-0065..eng-0081), but
retroactive is strictly worse than on-time: Home Assistant’s raw history
retention is ~10 days, so a suite left unjoined for too long loses the data
permanently, not just the convenience of computing it early.
The habit. Before a model page or a suite is treated as published:
- For every run window that has a
started_at/ended_at, compute the wall-meter join (counter-difference on the cumulative kWh sensor, §“What every run must carry”) and write thecontent/energy/eng-*.yamlrecord, wired via the run’senergy: {ref: eng-*}field. - Where a window genuinely cannot be joined — the sensor has no data for the
period, the run predates energy instrumentation, or (as with a suite whose
own run records have not yet been ingested) there is nothing to wire the
join to yet — record that explicitly in the run’s structured
energy_unjoined_reasonfield, rather than leavingenergy: nullto read as “not yet gotten to it”. A run may carry an energy reference or an unjoined reason, never both; comments do not count. An honest gap is a published fact; a silent one is a debt someone else discovers on a stale clock. - Batch sensibly. One broad history pull covering the whole suite’s span, sliced locally into per-cell windows, is cheaper and more reliable than one HA query per cell — and it is the only practical way to join a suite with dozens of cells before the retention window closes.
The guard. scripts/check-references.mjs runs a non-fatal check on every
build: for each page: true model, if any run on that model’s own configs
(best.serve_config, best.bench_config, best.fit_config, chart series
configs) carries neither an energy reference nor a nonblank structured
energy_unjoined_reason, it prints a warning naming the model and the unjoined run
ids. It never fails the build — an unjoined run is a legitimate,
recordable state per point 2 above — but it means the gap shows up in every
build log instead of staying invisible until someone goes looking for it.
Campaign-time ownership and close gate (added 2026-08-30). Every new campaign
predeclares the cumulative meter entity, counter-difference method and authorized
join-owner profile. The runner never receives Home Assistant credentials merely to
make that convenient: it always writes exact UTC started_at/ended_at edges and
hands them to the declared authorized reader. Before final evidence review, the
campaign energy handoff must be either joined with raw and derived receipt hashes,
or accepted-unjoined with a nonblank reason and reviewer identity. pending is not
a terminal state.
scripts/check-energy-handoff.mjs makes that boundary executable. It rejects a
campaign with valid time edges when its authorized handoff is still pending, rejects
hbrunner as the Home Assistant join owner, requires receipt hashes for a completed
join, and scans text artifacts for credential-like material. The v1 contract is
halobench.campaign-energy-handoff.v1; npm run test:energy-handoff exercises joined,
accepted-unjoined, pending and credential-leak cases. This closes inc-0011: the
full-STRIX runner correctly lacked credentials, but no authorized handoff was routed
before review sealed the unjoined disposition.
SMOKE validity gate (added 2026-08-14)
A tau2 SMOKE result with zero tool-call messages is invalid — on any compute surface,
for any model. The airline domain awards an untouched-database point and a COMMUNICATE
point to an agent that simply never acts, so a model whose tool surface fails silently can
post a perfect score while being incapable of the task. This is not hypothetical: the
first NPU screen produced npu-lfm2 at a vacuous 1.00 across 5/5 tasks with zero tool
calls (FastFlowLM dropped the tools array entirely — HTTP 200, no warning, measured as
a 0-token prompt delta), which would have out-ranked every GPU model screened on this box.
Therefore: tool_call_messages is a required recorded field for every SMOKE run, and a
mean reported without it is refused. The same silent-failure path exists on the GPU lane
whenever tool schemas break quietly (the HTTP-400 grammar-ceiling class), so this gate is
lane-independent.
Verdict readability is part of publishing (added 2026-08-17)
A verdict that reads as one dense wall of text is a publishing defect, not just a
style preference — six qwen38-27b/ornith-35b/deepseek-v4-flash/nemotron3-super/
laguna-s-21 verdicts each grew past 6-9 tightly-packed sentences with no break at
all. Before a model page ships: split the verdict into 2-4 short paragraphs by
story (lead, backend, capability, energy, hazards) using a blank line in the YAML
> scalar — full rule and mechanism in docs/site-design-v2.md §3. Words don’t
move between claims and their [clm-XXXX] chips; only where the paragraph breaks
fall changes.
Guard substitutions must be schema-visible (added 2026-08-16)
A guard requirement is never satisfied by a comment explaining why one guard stands in for
another — it is satisfied by a run record, wired through the guard (or guard_waived)
field with its own id. guard takes one reference; where a config’s guard picture needs more
than one (e.g. a bespoke cliff-detector plus the house 4-item capability guard), point guard
at the primary one and name the rest in guard_waived’s text — both are schema fields
src/lib/guard.ts actually reads, so both resolve at build time. A prose note that never
touches either field is not a guard reference; it is an explanation the badge cannot see.