Benchmark runbook — arrive and run
The ordered list.
benchmark-protocol.mdsays why; this says what, in what order, and what to do when it breaks. Written 2026-08-05 while the boxes were away, so that the day they boot is spent collecting rather than designing.
Everything here runs against the raw llama-server, never through llama-swap or the
slotpin proxy — except the one step that deliberately measures the proxy. A multi-hour
sweep through production wedges the single slot and evicts the warm lane.
Step 0 — Preflight (minutes, blocking)
./bench/preflight.sh --endpoint http://<host>:5802
Checks kernel ≥6.18.4 with CWSR exported, firmware, absence of amdgpu-dkms, quiesce
state, free memory, endpoint health — and runs the preserve_thinking acceptance
test. Exit 1 means do not benchmark yet.
If the tool-call probe reports empty-argument calls: stop. That is the template defect, and every agentic number collected past it measures the template rather than the model. Nothing downstream is salvageable.
Step 1 — Capability probe (Tier 0, ~1 h)
./bench/capability-probe.sh --model <gguf> --kv f16,q8_0,q4_0 --ctx 32768,131072,200000
Loads at increasing context per KV arm until it fails, stopping each arm at its first failure. Records max fit context and
observed_peak_gb — the field clm-0001 says is missing, and the reason our OOM
metric has never predicted an OOM. Capture peak here; it is nearly free now and expensive
to retrofit.
Do this per KV-quant arm, because it also answers the f16-fit question below.
Step 1b — Rebuild llama.cpp FIRST (possibly the cheapest win available)
A community report puts the upstream decode overhaul landing between roughly b9747 and b10218 at ~1.7x on this hardware — “if you’re on an old build of anything llama.cpp-based, rebuild”. Our production binary dates from June.
That is potentially larger than MTP, larger than the ubatch sweep, and costs a compile.
Sequencing tension, decide deliberately: rebuilding invalidates comparability with every existing number (build_commit is a fingerprint field). But the existing numbers are already on a stale build, so they are a baseline for that build and nothing else. The honest move is to measure once on the current binary, rebuild, measure again — that delta is itself a publishable result and nobody has quantified it for gfx1151.
⚠ Also rebuild the patched binary (con-0001 sidecar), or the comparison confounds a
version change with a fork change.
Step 1c — Capability guard (minutes, run at EVERY config)
./bench/guard.sh --endpoint http://aibeast:5802 --depth 8000 --parallel <n>
Run this before every sweep, not once at the start. It is the cheap tier that makes “best performance without breaking capability” affordable — four checks, minutes not hours:
- Coherence — catches a
binarycliff, where a backend is faster and emits garbage. - Tool call with non-empty arguments — catches the
preserve_thinking/template failure, whose signature is a call that looks successful becausetool_callsis populated while every argument object is empty. - Needle at depth — catches a retrieval collapse. Detects a cliff, not a slope; it is not a substitute for RULER.
- Cross-request isolation — only at
--parallel > 1, targeting llama.cpp #25992 response leakage on gfx1151 HIP. Pass--parallelmatching the config under test, or this check silently does not run.
A pass means no cliff detected, not capability intact. It cannot clear a lossy
lever: if this config differs from its capability baseline by weight or KV quant, run
Step 5 and set capability_basis: measured. The schema rejects inheriting capability
across a quant or backend change (protocol §1a), so this is enforced rather than trusted.
Record the guard as its own run (kind: guard) and point the performance run at it via
guard: run-XXXX. A performance run with neither a guard nor a written waiver fails the
build.
Step 2 — Throughput sweep (Tier 1, hours)
./bench/sweep.sh --model <gguf> --backend rocm \
--depths 0,4096,32768,131072 --kv f16,q8_0 --ub 512,2048 --fa 0,1 --reps 3
node scripts/ingest-llama-bench.mjs bench/results/<ts>/<cell>.json \
--host aibeast --backend rocm --out content
One llama-bench invocation per (KV × ubatch) cell — those are fingerprint fields and
must not be mixed inside one output file. Depth is not, so depths sweep together.
Sweep --fa 0,1, do not assume it. On a stock upstream build a dense model on this
silicon loses >90% of prefill at depth with flash attention on; with the unmerged
contiguize patch, FA wins every cell (clm-0017). Which side of that patch our build sits
on changes the correct setting, so it is a lever to measure, not a default to inherit.
Read the scatter line, not just the means. The sweep flags any cell whose repetitions disagree by more than 3%. High scatter at depth is the signature of a defective code path — treat a flagged cell as a finding to chase, not an error bar to publish.
Covers three open questions at once: the ubatch sweep we have never run, the
backend × context grid, and the quant × MTP ratio (add a --kv arm and an
MTP-off build to complete it — the transferable number is the ratio, not the absolute).
Step 3 — Cache capability trace (Tier 2, ~1 h) — the part nobody else publishes
./bench/cache-trace.sh --endpoint http://<host>:5802 --tokens 20000
In-slot reuse, cross-session divergence, slots, and disk save/restore. Do not shrink
--tokens: measured at 857 tokens once, reuse looked like 40%; at realistic bootstrap
scale it was ~97%. That artefact cost a day.
Includes the page-cache dependence measurement we have never made — drop caches between save and restore. If the cold and warm restore times match, our published ~105 ms figure was never touching disk and is wrong.
Step 4 — Served-path delta (~2 h)
llama-benchy (MIT, maintained) against the same model twice: once at the raw server,
once through slotpin → llama-swap. Its --enable-prefix-caching mode separates the cold
context load from cached reuse, and its output mirrors llama-bench so the two sit side
by side. Publish the gap — that difference is the cost of our own architecture, and
nobody else can report it. Sanity-check a few readings against the proxy’s /_stats
before trusting the delta.
Step 5 — KV-quant quality (overnight) — the contribution
./bench/kv-quality.sh --model <gguf> --text <corpus.txt> --kv q8_0,q4_0 --ctx 131072
Primary: llama-perplexity --kl-divergence. Records f16-KV logits once, replays q8_0
and q4_0 against them. The script verifies the flags against your binary before starting,
because flag names have moved between releases and a wrong one wastes an overnight run. Mean KLD, ΔPPL, Δp percentiles, top-1 agreement — with uncertainty
bars, single forward pass, no generation. Set -c to the full context so the whole cache
is exercised. Note the .kld file is large (tens of GiB).
Secondary: RULER-lite via lm-evaluation-harness (not NVIDIA/RULER — see protocol §8):
ruler_vt, ruler_cwe, ruler_fwe, niah_multikey_3 at 32K and 128K, 30-50 samples.
A cliff detector, not a precision instrument.
Read protocol §8 before starting. RULER at any affordable sample count resolves 5-10pp, while the published effect size is ~1pp — so it can only tell us “catastrophic or not”. And the hybrid-architecture confound may mean there is very little to measure.
Step 6 — Tool calling (~2-4 h)
./bench/tool-calling.sh --setup # one-off: clone + uv sync
./bench/tool-calling.sh --smoke --endpoint http://<host>:5802 # DO THIS FIRST
./bench/tool-calling.sh --domain airline --limit 10 --endpoint http://<host>:5802
Run --smoke before any real run. tau2-bench reaches models through LiteLLM, and
pointing LiteLLM at a local OpenAI-compatible server uses a pattern that is not spelled
out in tau2-bench’s own docs — it is the single most likely thing to fail. The smoke test
checks it twice: a direct HTTP call with a real tools= array, then the same call through
LiteLLM exactly as tau2-bench makes it.
τ²-bench, MIT and active, via LiteLLM against our endpoint — real native tools= calling
on our own jinja template. Hold both --user-llm and the NL-assertion judge model
fixed across comparisons, or results re-baseline silently.
./bench/grammar-ceiling.sh --endpoint http://<host>:5802 --max 120 [--schemas <dir>]
Binary-searches the ceiling in ~7 requests and reports it in bytes as well as count,
because bytes are the actual limit. Then the probe stays bespoke because neither τ²-bench (13-17
tools) nor MCP-Bench (bypasses native tool-calling entirely) exercises it. Lift
MCP-Bench’s mcp_servers/ corpus as raw schema fixtures and scale a real tools=[...]
array until HTTP 400, extending the manual 2026-07-05 probe (55 = 200, 60 = 400).
Step 7 — Cross-host overlap (when both boxes are up) — run it or lose it
Same suite, same config, both machines, cross_host_pair set. Isolates hardware as the
only variable. Almost nobody has this comparison because most people replace rather than
overlap, and the window closes when aibeast is repurposed.
Step 8 — Close the energy handoff before evidence review
The campaign declaration names the cumulative meter entity, counter-difference method and authorized join owner before hardware starts. Hbrunner records exact UTC edges for every arm but does not receive Home Assistant credentials. The authorized reader joins those edges and returns hashed raw/derived receipts, or hbreviewer explicitly accepts a structured unjoined reason.
Run node scripts/check-energy-handoff.mjs <campaign-energy-handoff.json> --scan <artifact-root> before final evidence review. A pending handoff is unfinished work,
even when both time edges are valid; a token or authorization header in the artifact
tree is a hard failure. See protocol §11 and inc-0011.
Open questions — for the operator
Is the 122B’s KV cache mostly linear-attention?— PARTLY ANSWERED 2026-08-05. The Qwen3.6-35B-A3B model card gives the layout explicitly: 40 layers as 10 x (3 x Gated DeltaNet -> MoE, then 1 x Gated Attention -> MoE) — so only 10 of 40 layers (25%) hold a KV cache at all, and those use just 2 KV heads. That matches the independently reported ~1.6 GB of KV+state per 128k tokens for this family. Consequence:-ctk/-ctvtouch a quarter of the layers. A near-zero quality delta would be an ARCHITECTURE result, not a quantisation one, and must be published as such. Still to confirm: the 122B’s own layer ratio, which is likely similar but unverified. Step 5 is probably worth ONE night to characterise, not a campaign.- Does the f16 arm fit at 128-200K? If not, the three KV arms are not comparable at the top length and the matrix needs redesigning around what loads. Step 1 answers it.
- Do we adopt the community quantized-KV fix before or after baselining? Before gives the best numbers; after gives a clean A/B of the fix itself. Recommendation: after — the A/B is the publishable result, and it needs an unpatched baseline to exist.
- Email Minervini for harness-bench access? No public repo, no licence. Near-zero cost to ask; unknown response time. Meanwhile Step 6 stands alone.
- MTP-off build for the ratio? Completing the quant × MTP matrix needs one. Worth the rebuild, or accept absolutes only?