Now
Generated 2026-08-11 from
content/claims, candidate gate histories and the git log — not maintained by hand. This page regenerates with the record: rebuild it whenever the claims, candidates or change stream move, and it cannot drift from what actually happened. If a line here looks stale, the record it was built from moved and this page was not rebuilt against it — that is a build problem, not an editorial one.
Running now
- Stage E — Nemotron quant re-run, on aihydra: Q4_K_M against the IQ4_XS it has run at
since acquisition. Every Nemotron-vs-peer comparison to date is confounded by this quant
mismatch (
candidates/nemotron3-super,clm-0039); this run is the fix. - aihydra is the sole benchmark host. aibeast’s board draws standby but never latches
or POSTs — a board-level short, RMA open with GMKtec (
inc-0005). - The failed host’s production telemetry and configs are archived — the only copy of the MTP production metrics is secured off the dead unit.
Landed this week (2026-08-09 → 2026-08-11)
- The 122B’s real τ²-bench score is 0.545 at n=22, not the 1.00 that five-task arms
reported — those sampled only the easy end of the domain (
clm-0037). - Turns, not reward, are the real headline. On tasks it gets right, the 122B needs a
median 19 turns / 2.0 min against Nemotron’s 26 / 3.7 — reward alone called them
indistinguishable (
clm-0039). - Every cross-model τ² comparison ran an unpinned user simulator — the model under
test played both sides. 122B-vs-Nemotron and thinking-on/off comparisons are now flagged
provisional; same-model KV comparisons are unaffected
(
clm-0043). - Energy re-measured per task, matched task sets. 122B f16 costs 6.81 Wh / 0.21p per
correct answer, the cheapest of anything measured — the earlier arm-level figures had
over-counted by roughly 10x (
clm-0042). - The KV-dequant patch closes most of quantised KV’s agentic cost, not just its speed
cost. Stock q8_0 cost +39% turns; patched, that falls to +9.3%
(
clm-0045). - A second independent community source narrows the Vulkan-vs-ROCm decode gap from an
earlier confounded 55% down to ~10.5% clean, and surfaces an unmerged BF16
flash-attention PR relevant to the KV-quality thread
(
clm-0044). - Laguna S 2.1 screened on stock, no fork required after all: guard 4/4, smoke 0.80
over 5/5 tasks — the blocker had already resolved upstream
(
candidates/laguna-s-21). - Muse Glimmer 30B screened on its minimum build: guard 3/4 — a needle-retrieval
failure at depth 8000 — and a meandering smoke (median 108 turns). The evidence points at
day-one upstream arch immaturity, not the model: an independent run holds retrieval to
832K tokens, and same-week upstream bugs sit in exactly this arch’s attention metadata.
Re-screens when the upstream fixes land (
candidates/muse-glimmer-30b). - Wall-energy instrumentation backfilled across the board: 34 energy records added, 92 of 99 runs now linked, and a units bug (joules crept into notes) was caught and fixed before it shipped.
- The candidate field was reconstructed from scratch after two models had been quietly
dropped on assumption rather than measurement: 17 candidates, every gate transition now
carries a reason and, where rejected, evidence (
docs/candidates.md).
Queued
- Nemotron Q4_K_M — running now, see Running now above.
- Five screened candidates await a powered τ² slot: Laguna S 2.1, Nemotron-Cascade 2, Gemma 4 26B-A4B, GLM-4.7-Flash, LFM2-24B-A2B — each has a fit probe, guard and smoke test, none has a full arm.
- Qwen3.6-27B-MTP — regressed to
acquired; zero rows in the aihydra matrix, needs re-screening on this box rather than inheriting its old-box pass. - Qwen3.8-27B — watched, not acquired. No 27B GGUF exists yet, only FP8/NVFP4; a daily two-tier watch on aihydra flags the first one that lands.
- Cross-model τ² comparisons need re-running under a pinned, independent user
simulator before the 122B-vs-Nemotron and thinking-on/off results stop being
provisional (
clm-0043). - Whether the patched q8_0 +9.3% turn residual is real or noise — n=11 pairs, one
seed; needs repeats before it is a number rather than a direction
(
clm-0045).
Pending decisions & external dependencies
- Independent pinned user-simulator — now provisioned (Haiku 4.5 via OpenRouter, provider-locked); the simulator-sensitivity pair gates the cross-model series.
- Muse Glimmer upstream issue — the ROCm-side needle failure is characterised; an issue dossier awaits filing upstream.
- RMA in progress on the failed benchmark host’s board (
inc-0005); the replacement lab host carries the full programme meanwhile.
Recent decisions
- aihydra is now the benchmark host of record, not a stand-in for aibeast. Nothing in the protocol or runbook was aibeast-specific, so the measurement programme itself carried over intact.
- 5-task τ²-bench arms are retired as a ranking instrument. A 5-task mean is
effectively two Bernoulli trials; the 200-step default and matched task sets replace it
(
clm-0036,clm-0037). - The energy denominator is fixed to per-task, matched-task windows. Whole-arm totals
over-stated cost by about 10x because most of an arm’s wall time was model loading or
tasks that never scored (
clm-0042). - Two candidate exclusions made on assumption — Maple-Preview, Gemma-4-12B — were
withdrawn. Nothing leaves the candidate field without a measurement against it
(
docs/candidates.md). - Quant equivalence is now a precondition, not a footnote. A candidate may not be compared against the field at a different quantisation without the mismatch stated in the result itself — the rule Stage E exists to satisfy for Nemotron.