Docs › now

Now

Generated 2026-08-11 from content/claims, candidate gate histories and the git log — not maintained by hand. This page regenerates with the record: rebuild it whenever the claims, candidates or change stream move, and it cannot drift from what actually happened. If a line here looks stale, the record it was built from moved and this page was not rebuilt against it — that is a build problem, not an editorial one.

Running now

  • Stage E — Nemotron quant re-run, on aihydra: Q4_K_M against the IQ4_XS it has run at since acquisition. Every Nemotron-vs-peer comparison to date is confounded by this quant mismatch (candidates/nemotron3-super, clm-0039); this run is the fix.
  • aihydra is the sole benchmark host. aibeast’s board draws standby but never latches or POSTs — a board-level short, RMA open with GMKtec (inc-0005).
  • The failed host’s production telemetry and configs are archived — the only copy of the MTP production metrics is secured off the dead unit.

Landed this week (2026-08-09 → 2026-08-11)

  • The 122B’s real τ²-bench score is 0.545 at n=22, not the 1.00 that five-task arms reported — those sampled only the easy end of the domain (clm-0037).
  • Turns, not reward, are the real headline. On tasks it gets right, the 122B needs a median 19 turns / 2.0 min against Nemotron’s 26 / 3.7 — reward alone called them indistinguishable (clm-0039).
  • Every cross-model τ² comparison ran an unpinned user simulator — the model under test played both sides. 122B-vs-Nemotron and thinking-on/off comparisons are now flagged provisional; same-model KV comparisons are unaffected (clm-0043).
  • Energy re-measured per task, matched task sets. 122B f16 costs 6.81 Wh / 0.21p per correct answer, the cheapest of anything measured — the earlier arm-level figures had over-counted by roughly 10x (clm-0042).
  • The KV-dequant patch closes most of quantised KV’s agentic cost, not just its speed cost. Stock q8_0 cost +39% turns; patched, that falls to +9.3% (clm-0045).
  • A second independent community source narrows the Vulkan-vs-ROCm decode gap from an earlier confounded 55% down to ~10.5% clean, and surfaces an unmerged BF16 flash-attention PR relevant to the KV-quality thread (clm-0044).
  • Laguna S 2.1 screened on stock, no fork required after all: guard 4/4, smoke 0.80 over 5/5 tasks — the blocker had already resolved upstream (candidates/laguna-s-21).
  • Muse Glimmer 30B screened on its minimum build: guard 3/4 — a needle-retrieval failure at depth 8000 — and a meandering smoke (median 108 turns). The evidence points at day-one upstream arch immaturity, not the model: an independent run holds retrieval to 832K tokens, and same-week upstream bugs sit in exactly this arch’s attention metadata. Re-screens when the upstream fixes land (candidates/muse-glimmer-30b).
  • Wall-energy instrumentation backfilled across the board: 34 energy records added, 92 of 99 runs now linked, and a units bug (joules crept into notes) was caught and fixed before it shipped.
  • The candidate field was reconstructed from scratch after two models had been quietly dropped on assumption rather than measurement: 17 candidates, every gate transition now carries a reason and, where rejected, evidence (docs/candidates.md).

Queued

  • Nemotron Q4_K_M — running now, see Running now above.
  • Five screened candidates await a powered τ² slot: Laguna S 2.1, Nemotron-Cascade 2, Gemma 4 26B-A4B, GLM-4.7-Flash, LFM2-24B-A2B — each has a fit probe, guard and smoke test, none has a full arm.
  • Qwen3.6-27B-MTP — regressed to acquired; zero rows in the aihydra matrix, needs re-screening on this box rather than inheriting its old-box pass.
  • Qwen3.8-27B — watched, not acquired. No 27B GGUF exists yet, only FP8/NVFP4; a daily two-tier watch on aihydra flags the first one that lands.
  • Cross-model τ² comparisons need re-running under a pinned, independent user simulator before the 122B-vs-Nemotron and thinking-on/off results stop being provisional (clm-0043).
  • Whether the patched q8_0 +9.3% turn residual is real or noise — n=11 pairs, one seed; needs repeats before it is a number rather than a direction (clm-0045).

Pending decisions & external dependencies

  • Independent pinned user-simulator — now provisioned (Haiku 4.5 via OpenRouter, provider-locked); the simulator-sensitivity pair gates the cross-model series.
  • Muse Glimmer upstream issue — the ROCm-side needle failure is characterised; an issue dossier awaits filing upstream.
  • RMA in progress on the failed benchmark host’s board (inc-0005); the replacement lab host carries the full programme meanwhile.

Recent decisions

  • aihydra is now the benchmark host of record, not a stand-in for aibeast. Nothing in the protocol or runbook was aibeast-specific, so the measurement programme itself carried over intact.
  • 5-task τ²-bench arms are retired as a ranking instrument. A 5-task mean is effectively two Bernoulli trials; the 200-step default and matched task sets replace it (clm-0036, clm-0037).
  • The energy denominator is fixed to per-task, matched-task windows. Whole-arm totals over-stated cost by about 10x because most of an arm’s wall time was model loading or tasks that never scored (clm-0042).
  • Two candidate exclusions made on assumption — Maple-Preview, Gemma-4-12B — were withdrawn. Nothing leaves the candidate field without a measurement against it (docs/candidates.md).
  • Quant equivalence is now a precondition, not a footnote. A candidate may not be compared against the field at a different quantisation without the mismatch stated in the result itself — the rule Stage E exists to satisfy for Nemotron.