Home › Method
⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.

Method

Why any of this should be believed. Every published number is wall-measured on the lab's own hardware, carries its provenance and confidence, joins to the run and configuration that produced it, and is corrected in public when it falls. The pages below are the machinery: the protocol that constrains how measurements are made, the standing rules extracted from measurement failures, and the corrections the process has already produced.

The protocol

Two measurements, not one: capability is what a model can do — slow, expensive, run rarely; performance is how fast it delivers that — cheap, run often, and only meaningful behind a guard, the cheap cliff-detector run at every performance configuration. What makes that split legal is a classification of every configuration lever by how it can break capability, and a machine-readable constraint file the harness checks before every run — prose is read once and remembered badly; checks fire every time.

fleet default build 3653e6d ·11 known builds, each pinned with an anchor pair · user simulator pinned to openrouter/anthropic/claude-haiku-4.5 ·6 comparison rules enforced

Corrections — the record correcting itself

Corrections live in the Log timeline with every other entry, badged inline — visible, not enshrined. The registry below them is a generated filter of that same stream: every retracted or superseded claim, what was wrong, and what fixed it.

clm-0126supersededverified 2026-08-29

On AI Hydra's admitted HIP/gfx1151/native/Release host boundary, the KingJones R2 Qwen3.8-Flash-Next full-STRIX ROCmFP4 campaign fixed model artifact c6770d7442a06bf1d78edf28cec83e1ec93afdd34664c23ff898807b6b9349fa (121838036032 bytes, publisher revision 069dddb53bab04218d734fa9a771f8a0242ab059), runtime source 36e9acd40e10a87cd3c3ef8ec734668757dc8520, and the independently admitted post-route receipt patch 1f1c3bed910922b415c1be36c9a04c9b7fedc4162aedf504a9da007daebcb4d2. The exact tracked 1,017-byte rotate-bits header was present at blob 75c4881fc322f2e6a6ee9d809e696852531abb8c and the patch clean-applied at +109/-2. The card-authorized targeted llama-server capture build, not an all-target build, then failed because sha256.c could not resolve rotate-bits/rotate-bits.h. Therefore no receipt-backed guard, Tau2 capability, llama-bench floor, served-path, cache, depth or energy result exists for this campaign. This is not a model quality, fit, throughput, capability or source-runtime performance claim. The earlier Unsloth UD-Q4_K_XL/Vulkan records are a separate artifact, runtime and backend series and are not merged with this full-STRIX result.

clm-0124supersededverified 2026-08-27

Qwen3.8-Flash-Next UD-Q4_K_XL (Qwen4-arch preview MoE: 125B MoE / ~6B active + 51B N-gram/PLE tables + vision) on the Unsloth qwen4exp Vulkan build (cfg-0177, commit 250b6144, gfx1151/RADV) — first full benchmark. ADMITTED CLAIM 1 (capability leader): on the standard tau2 airline full-26 suite (seed 42, claude-haiku-4.5 simulator, deterministic reward) it scored mean_reward 0.9231, 24/26 tasks at reward 1.0 (run-0632). This LEADS the tau2 airline board — the prior best was Ornith-1.0-35B UD-Q4_K_XL at 22/26 (0.846). Tool-calling was clean: 184 tool calls, 0 empty-argument calls, 1 tool-error message, 24.7 messages/task. Failures = tasks 7 and 20. ADMITTED CLAIM 2 (energy leader per correct answer): wall-metered join (eng-0275, HA counter-difference) = 256.08 Wh total over 6605 s, 237.55 Wh active above the 10.1 W idle floor, = 9.90 Wh per correct answer (0.300 p @ 30.3 p/kWh). This is ~4x cheaper per correct answer than the ho003 stock-f16 tau2 control (42.06 Wh/correct), because it solves more than twice as many tasks in less wall time. ADMITTED CLAIM 3 (footprint): GTT-resident 76.7 GiB with ~43 GiB free for KV. The 26.8 GiB n-gram/PLE table (per_layer_token_embd, iq4_nl) is placed on CPU automatically by the Vulkan backend regardless of -ot, so it never occupies GTT; KV scales ~48 MiB per 1000 tokens (hybrid QSA attention). Performance: prefill 277-313 tok/s at 5-8K, decode ~22 tok/s short falling to 8.5 at 131K; TTFT ~2.9 s. BLOCKED CLAIMS (explicitly NOT made): single trial (N=1), no variance bound. reasoning_effort=low, not the model's default xhigh — a higher-reasoning result is untested and may differ (both tasks 7 and 20 may be reasoning-recoverable). The runtime is a draft/WIP PR (#27742), not a released/immutable build; the GGUF's advertised LICENSE artifact 404s and conversion-base provenance is unpinned, so this is a LAB result with NO weight redistribution or public/production recommendation. MTP speculative decoding is unavailable (the GGUF lacks MTP head layers) and n-gram speculation did not help this quant. 262K context hits a Vulkan workgroup-count assertion (cap -c <= 262140). No UD-Q2 control or cross-host pair run yet.

clm-0105supersededverified 2026-08-22

SUPERSEDED (operative claim) as of strix-halo-llamacpp v0.6.10. As written against v0.6.9 (2026-08-22): on hybrid GDN (qwen35moe) targets the then-stable Strix Halo Vulkan fork did NOT guarantee token-exact MTP rollback after a rejected draft. v0.6.8 had introduced a one-line change forcing MTP rollback through full sequence-state checkpoints; v0.6.9 (2026-08-22T04:09:43Z) reverted that line because the full-checkpoint save/restore path deadlocked deterministically on Vulkan with a hybrid target (Qwen3.6-35B-A3B stalled a few hundred tokens into a long response, all threads parked in futex_do_wait, GPU idle, GTT flat). With the revert, rollback returned to the v0.6.4-v0.6.7 fast snapshot-plane restore, which the vendor stated "can diverge slightly from a no-draft run after a rejected draft" - a subtle distribution drift after rejection, not garbled output. That was a deliberate availability / exactness tradeoff by the fork, pending a state-save-path fix. v0.6.10 (clm-0107) re-lands token-exact full-checkpoint rollback; the operative tradeoff claim above no longer holds on the current stable fork.

Lessons — standing rules with their scars

Rules extracted from measurement failures, not a diary of them. Each cites the claim it was drawn from, so the worked diagnosis is one click away.

Essays — the long-form record

Instruments

The bench scripts are the instrument set — each enforces part of the protocol rather than trusting a reader to remember it. Energy is read from a wall meter's cumulative kWh counter, differenced across each run's window; never modelled, never joules.

scripts: cache-trace-probe.py cache-trace.sh capability-probe.sh env-lock.sh export-tool-schemas.sh grammar-ceiling.sh guard-probe.py guard.sh ha-energy-join.py ho004-derive.py ho004-dump-evidence.py ho005-revised-energy-join.py ho005-revised-ornith-batch-energy-join.py ho009-v0610-energy-join.py ho012-122b-energy-join.py ho014-energy-join.py kv-quality.sh preflight.sh protocol-check.sh run.py runmeta.sh spec-ab.sh sweep.sh tool-calling.sh
meters: TP-Link smart plug via Home Assistant AI Hydra energy/power sensors via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy; peak/mean drawn from sensor.hardware_ai_hydra_power) aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy) aihydra energy sensor via Home Assistant (5-minute statistics on sensor.hardware_ai_hydra_energy, linearly interpolated to the run window) aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, 5-minute statistics interpolated to the run-meta window) aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s history samples interpolated to the run-meta window) aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the run-meta window) aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to each run-meta window; two windows summed, not a single counter-diff over the outer span) aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the session window) aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the transcript-derived window) aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the recorded window) aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, high-confidence raw ~10s-resolution samples linearly interpolated to the recorded window) aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s samples linearly interpolated to the recorded window) aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the recorded window; power from sensor.hardware_ai_hydra_power where available) aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples linearly interpolated to the recorded window; power from sensor.hardware_ai_hydra_power where available) aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy; raw ~10s-resolution history samples queried from hborchestrator using retained HA history) aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the run window) aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s samples queried by hborchestrator via Nabu Casa API, idle baseline 10.1 W) aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-history samples queried by hborchestrator via Nabu Casa API and linearly interpolated to the recorded window; power from sensor.hardware_ai_hydra_power) aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-history samples queried by hborchestrator via Nabu Casa API and linearly interpolated to the window; power from sensor.hardware_ai_hydra_power) aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-history samples queried by hborchestrator via Nabu Casa and interpolated to the window; power from sensor.hardware_ai_hydra_power) aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, sampled and interpolated to the window; power from sensor.hardware_ai_hydra_power) aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-history samples queried by hborchestrator via the Nabu Casa API and linearly interpolated to the window; power from sensor.hardware_ai_hydra_power) aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-history samples queried by hborchestrator via the HA API and linearly interpolated to the window; power from sensor.hardware_ai_hydra_power) aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-history samples queried by hborchestrator via the HA API and linearly interpolated to the band edges; power from sensor.hardware_ai_hydra_power) aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw history samples queried by hborchestrator via the HA API and linearly interpolated to the window edges; power from sensor.hardware_ai_hydra_power) aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw history queried via ha_get_history and linearly interpolated to the window edges; power cross-checked from sensor.hardware_ai_hydra_power 5-minute statistics) AI Hydra Home Assistant cumulative kWh counter sensor.hardware_ai_hydra_energy, queried by the authorized hborchestrator reader and linearly interpolated to retained exact UTC edges. aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw history queried via ha_get_history at the window edges and a midpoint; power cross-checked from sensor.hardware_ai_hydra_power hourly statistics) · corpus pinned in bench/CORPUS.md

Contributions — upstream work, fork-carrying disclosed

carrying marks a fork in production use but not merged upstream — a reproducibility hazard disclosed on every configuration that depends on it, not a badge of honour