The enforced protocol
This page renders bench/protocol.json — the machine-readable constraint file, written after four constraint violations in one day, each one a rule the project had already written down and then not checked before acting. Prose is read once and remembered badly; checks fire every time. The prose protocol, with the full reasoning, is at docs › benchmark-protocol.
τ²-bench pins
| pin | value | why |
|---|---|---|
| user simulator | openrouter/anthropic/claude-haiku-4.5 | Pinned independent simulator per clm-0043 and methodology-lessons 3: different family from every local candidate, dated slug, temperature 0. Served via OpenRouter (PAYG, operator account, 2026-08-11); first probe routed to Amazon Bedrock, so routing is enforced account-side: the EZAxis Provider Limit guardrail in OpenRouter allows only Anthropic (set by the operator 2026-08-11, verified by probe: provider now reports Anthropic where the first probe hit Bedrock). Supersedes the 122B self-play pin, which was only ever satisfiable when the agent WAS the 122B. Adequacy to be verified by the sim-sensitivity pair (haiku vs sonnet on the seeded anchor subset) before the cross-model series is trusted. |
| simulator temperature | 0 | determinism — the simulator is scaffolding, and a noisy simulator is a noise source in its own right (methodology lesson 3) |
| max steps | 200 | tau2 default. A lower cap sits above one model's whole turn distribution and slices through another's, biasing comparison toward terse models (clm-0039). |
| max concurrency | 1 | llama.cpp #25992 leaks responses across requests on gfx1151 HIP at --parallel > 1; and tau2 defaults to 3, which produces infrastructure_error against a --parallel 1 server. |
Headline metrics
docs/14-model-backend-benchmark.md: 'time-to-correct-result + loops-to-done, with tool-call success as a gate'. Raw reward is NOT the headline; it is binary per task and hides turn-count differences entirely.
time_to_correct_answer_s loops_to_done tool_call_success
Required run fields
Without a window, per-run energy must be reconstructed from file mtimes, and the recorder keeps 10-second history for only ~10 days.
started_at ended_at
Energy units
design doc 2.4 - joules are not a home-energy unit and do not map to /kWh tariffs.
| allowed | forbidden |
|---|---|
| Wh · kWh · mWh · pence · W | joule · joules · J/token · kJ |
Comparison rules
- Matched task sets. Bound-limited arms reach different depths; comparing different prefixes compares different work (clm-0037).
- Same quantisation across compared models, or the comparison measures the quant (Nemotron IQ4_XS vs peers Q4_K_M).
- Paired over identical tasks beats unpaired means at these sample sizes (clm-0038).
- State the denominator. Energy per correct answer is not energy per task is not agent-only energy (clm-0040, clm-0042).
- Compute surfaces never mix silently. NPU/CPU/GPU results follow the same benchmark format and MAY be compared for the same model across surfaces — that contrast is commentary gold (NPU prefill vs GPU, decode trade-offs) — but every displayed result carries its surface and backend (config engine + runtime.backend), purpose-fit verdicts stay within-surface, and no table row mixes surfaces without a surface column. NPU-lane candidates never inherit GPU metrics (existing schema rule).
- SMOKE VALIDITY GATE (added 2026-08-14, from the first NPU screen). A tau2 result with ZERO tool-call messages is INVALID and must never be recorded as a mean, on any surface. An agent that never acts scores the untouched-DB and COMMUNICATE points by default and can outscore a competent tool-user: npu-lfm2 scored a vacuous 1.00 this way, beating every GPU model screened, because FastFlowLM silently dropped the tools array (HTTP 200, no warning). Same do-nothing path opens whenever a tool surface fails quietly - including the HTTP-400 grammar-ceiling class. Record tool_call_messages alongside every SMOKE mean; refuse a mean without it.
Quantisations expected in a series
protocol_check_build refuses a quant matching none of these. Same quantisation across compared models is a comparison rule (the Nemotron IQ4_XS confound, clm-0047); a stray quant appearing in a series is either a typo or an unplanned comparison, and both should stop the run. UD-IQ4_XS stays listed because the IQ4_XS-vs-Q4_K_M pair is itself a measured comparison. Extend the list deliberately when a new quant enters the programme.
UD-Q4_K_M UD-Q4_K_XL UD-IQ4_XS Q4_K_M · added 2026-08-12 repeatability audit
Builds — pinned, anchored, never bumped mid-series
Minimum-required-version per model, never a fleet-wide bump mid-series. Each build is a PINNED commit with its own config records (runtime.commit); a model's minimum build lives in its candidate record. Introducing a build REQUIRES an anchor pair: the anchor workload run on both old and new binary in the same session, so the build delta is measured and nearest-comparable-config comparisons are calibrated, not assumed. Operator decision, 2026-08-10.
anchor workload: 122B UD-Q4_K_M, tau2 airline --task-ids 0 1 2 3 4 5 6 8 9 --seed 42 (subset only), plus llama-bench -p 512 -n 128,512. ~30 min. · fleet default: 3653e6d
| build | what it is |
|---|---|
| 3653e6d | stock baseline, all pre-Aug-10 runs |
| ce7689f | kvfix — stock + KV dequant patch (clm-0022) |
| min-62bf73d | muse-glimmer support (llama.cpp PR 26841); building |
| 77a9a66eb | dual-backend HIP+Vulkan single binary with PR 26856 bf16-FA cherry-pick — the clm-0046 backend-comparison build. Added 2026-08-12 (repeatability audit): it had produced published numbers without being a known build. |
| 3be50ccc | Nathanw1014/strix-halo-llamacpp v0.6.1 portable tarball (fork branch strix-halo-vulkan, Vulkan only, bundled RADV Mesa 26.3.0-devel + libdrm 2.4.134) - backend-decider arm C. Calibration: matched pp1024/tg256 depth cells vs stock 3653e6d ROCm and Vulkan run in the same session (backend-decider suite) serve as the anchor pair for this Vulkan-only build; the 122B tau2 anchor does not apply. Added 2026-08-12. |
| ed89854 | stew675/llama.cpp branch rdna-boosts, head ed89854b2aeb0e333dd61424f14af2aedaca126e (2026-08-13T21:09Z), built on-box 2026-08-14 with the stock ROCm flags (-DGGML_HIP=ON -DAMDGPU_TARGETS=gfx1151 -DCMAKE_BUILD_TYPE=Release -DLLAMA_CURL=ON), like-for-like with 3653e6d except the branch kernel-fusion changes; community challenger to clm-0050 (claims ROCm decode deficit ~3.2% at d32768 and ROCm prefill +36.5% past Vulkan with BF16 KV). Anchor pair: matched pp1024/tg256 depth cells vs stock 3653e6d ROCm and stock 3653e6d Vulkan run in the same session (rdna-boosts-challenger suite). Added 2026-08-14. |
| a94d563 | upstream llama.cpp a94d563ed801d1da1b8c2432946de07d0231bb3d (2026-08-13, PR #27026), the exact commit stew675's rdna-boosts branch is rebased on, built on-box 2026-08-14 with the same ROCm flags as every other arm. Exists solely to de-confound the rdna-boosts-challenger matrix: ed89854 sits six days of upstream ahead of the 3653e6d baseline, so a94d563-vs-3653e6d measures upstream drift and ed89854-vs-a94d563 measures stew675's 23 kernel-fusion commits alone. Anchor pair: matched pp1024/tg256 depth cells run in the same session (rdna-boosts-challenger suite). Added 2026-08-14. |
| baf0025 | Nathanw1014/strix-halo-llamacpp v0.6.2 portable tarball (fork branch strix-halo-vulkan, Vulkan only, bundled RADV Mesa 26.3.0-devel + libdrm 2.4.133, same pinned shaderc as v0.6/v0.6.1) - two upstream cherry-picks on top of 3be50ccc (v0.6.1): Muse Glimmer support (llama.cpp PR #26841) and a tool-call-after-EOM parsing fix (#26879). Release notes state no backend or shader changes vs v0.6.1, so the vendor claims it compares like-for-like with 3be50ccc; this lab did NOT run its own anchor pair for this specific increment (recorded as a gap, not silently assumed) — the two builds differ only by model-support cherry-picks, not by anything on the throughput-relevant path. Used to re-screen muse-glimmer-30b (candidate history: prior gate blocked on implementation immaturity on min-62bf73d, RE-SCREEN TRIGGER named there). Tarball sha256 3f16b6671dfb349b088ec158d1775b476cd551a6442842609b4855034846d2c8 (verified against the GitHub release asset digest before use). Added 2026-08-15. |
| ece963f | upstream llama.cpp master ece963f41b0b02d7a0d61436ae365762c073a4c8 (2026-08-15T20:52:55Z, HEAD at clone time), Vulkan-only shallow-clone build on-box 2026-08-16 (GGML_VULKAN=ON, GGML_HIP=OFF, Release, GNU 15.2.0 -- same compiler as the fleet baseline) built solely for the eos-cliff filing-gap ITEM B still-present-on-master check (exact failing baseline vs run-0277's 3653e6d result); no anchor pair run -- bug-reproduction check, not a throughput/capability comparison. Added 2026-08-16. |
| 7077abbe | post-BailingMoE3-merge master. Upstream llama.cpp 7077abbe14c510cb829c93a1328c2815b5805ebd (2026-08-17T11:14:33Z, HEAD at clone time, shallow depth 50), which includes 'model : BailingMoE3 Support (#26608)' commit 373336672029b12e09f272bc027cc801345a3fd6 merged 2026-08-17T07:49:49Z -- verified in-tree (git log shows the commit; src/models/bailingmoe3.cpp and the bailingmoe3 arch string present in llama-arch.cpp). Built on-box 2026-08-17 in TWO variants from the same clone/commit: build-rocm (-DGGML_HIP=ON -DAMDGPU_TARGETS=gfx1151 -DCMAKE_BUILD_TYPE=Release -DLLAMA_CURL=ON) and build-vulkan (-DGGML_VULKAN=ON -DGGML_HIP=OFF -DCMAKE_BUILD_TYPE=Release -DLLAMA_CURL=ON), both Release/GNU default compiler, matching the fleet's existing per-backend flag conventions (3653e6d ROCm arm, 3be50ccc/ece963f Vulkan arms). No anchor pair run: this build exists solely to unblock ling-30-flash's 2026-08-13 FIT-FAILED 'unknown model architecture: bailingmoe3' gap -- the architecture did not exist in any prior registered build, so there is no matched-config baseline to anchor a delta against (same no-anchor rationale class as ece963f). Added 2026-08-17. |
| 2586f6ed | Nathanw1014/strix-halo-llamacpp v0.6.10 portable payload (fork branch strix-halo-vulkan, Vulkan only, bundled RADV Mesa d18d598e + libdrm 2.4.133 + glslc/shaderc v2026.3-dev; payload commit 2586f6eddae19bef3dd21a1a0b109cca7bcf1c32 'model : support DSpark for bailingmoe3 (#27508)', matched by llama-server --version 2586f6ed). Downloaded from GitHub release v0.6.10 (asset strix-halo-llamacpp-vulkan-portable.tar.gz, published 2026-08-22T12:03:16Z); tarball sha256 3219732039cb9bceb36427ae7108b36992a3696529fdc04bc7a165502e5119eb verified against the release asset digest (matches card G0), MANIFEST.txt sha256 f4cb938858caebdb8e6e63959008fbe45e0e3ebaccb62de4cce1bf90cf4d0493 verified, staged at aihydra /home/aihydra/src/strix-halo-v0610/ with pristine binaries (vulkan/bin/llama-bench, llama-server, llama-cli all executable, Vulkan device Radeon 8060S detected). No anchor pair run in this staging scope (card is non-box / no-benchmark): the builds.policy anchor pair is a recorded necessary gap, not silently assumed, to be run by the HO-004-DIAG box-ready card before any HO numbers are trusted. Added 2026-08-23. |
rendered from bench/protocol.json at build time — this page cannot disagree with the check that fires