Home › Log

Log

The lab notebook, generated from the record: claims as they land, candidates moving through the gates, run series, incidents and upstream contributions. Retractions and supersessions appear inline in the same timeline, badged — corrections are entries here like any other learning, and a generated filter of this stream collects them. Subscribe: RSS.

87 entries · 3 of them corrections · every entry derives from a dated record — nothing on this page is written by hand

August 2026

  • 2026-08-13claimOn gfx1151 at f16 KV, stock Vulkan beats stock ROCm in EVERY cell of a matched matrix (one binary commit 3653e6d, one model, depths 0 to 131,072): decode +19-21% at every depth, prefill +4% to +20% growing with depth to 65k. clm-0050
  • 2026-08-13claimQuantised KV on gfx1151 splits three ways by build. clm-0051
  • 2026-08-13gate7 candidates → screened: gemma4-12b, glm-45-air, ling-30-flash, llama4-scout, maple-preview, nemotron35-lightning-30b, qwen36-27b-mtp
  • 2026-08-13gatenemotron35-lightning-30b: listed → acquired — Official ggml-org Q4_K_M downloaded to aihydra (25,430,738,944 bytes, exactly the listed size) and sha256-verified against the HF LFS oid (6110e2e2e6cd324e6ee69ddced5a6b34fad6c94ca9827222a1e420fb92e3c90b).
  • 2026-08-12gatedeepseek-v4-flash: listed → listed — Identity now concrete via the r/LocalLLaMA Strix Halo guide thread: DeepSeek V4 Flash 0731, deepseek4 arch, 256 experts/6 active + 1 shared, MIT licence, 1M native context. clm-0031
  • 2026-08-12gatenemotron35-lightning-30b: listed — Surfaced via an operator-shared r/AIDeveloperNews link ("NVIDIA has launched Nemotron 3.5 Lightning") plus a follow-up r/StrixHalo post on a community ROCmFP4 requant with hardware-matched Strix Halo numbers.
  • 2026-08-11claimA second independent Strix Halo source (llama.cpp PR #26856 + its Reddit write-up) reports Vulkan ahead of ROCm on decode at depth by ~10.5% on a clean same-binary comparison — same direction as clm-0031's +55% but a fifth the magnitude, confirming that figure was mostly build-gap and private patches. clm-0044
  • 2026-08-11claimThe KV dequant patch removes most of quantised KV's agentic cost, not just its speed cost: on identical seeded tasks, patched q8_0 takes +9.3% more turns than patched f16 (234 vs 214 over 11 paired tasks) where the stock build cost +39% (clm-0038). clm-0045
  • 2026-08-11claimThe community BF16 flash-attention predictions (clm-0044) reproduce on this hardware under single-binary methodology: at 32k depth, stock Vulkan leads ROCm +17.4% prefill / +17.1% decode at f16 KV, and the PR-26856 patch inverts prefill to ROCm +46.9% with bf16 KV. clm-0046
  • 2026-08-11claimNemotron's quant confound resolves cleanly: on 12 identical seeded tasks, UD-Q4_K_M and UD-IQ4_XS produce IDENTICAL reward on every task (0.583 both), but IQ4_XS takes +11% more total turns (630 vs 567). clm-0047
  • 2026-08-11claimSimulator identity changes what tau2 measures. clm-0048
  • 2026-08-11claimThe 122B has a reproducible rule-precedence bug in policy application: on tau2 airline task 9 it cancels a partially-flown reservation in 7 of 8 trials, every failure with the identical signature — cancel_reservation called without ever checking flight status. clm-0049
  • 2026-08-11gatemuse-glimmer-30b: blocked → screened — Screened on min-62bf73d (its minimum build, anchor-calibrated: +2.8%/+0.9% vs fleet baseline, identical tau2 capability).
  • 2026-08-11runs9 runs landed on cfg-0029, cfg-0030, cfg-0025, cfg-0031, cfg-0028, cfg-0027 (tau2-bench-airline) cfg-0029 cfg-0030 cfg-0025 cfg-0031 cfg-0028 cfg-0027
  • 2026-08-10claimThe 122B's real τ²-bench airline score is 0.545 +/-0.208, not the 1.00 reported by 5-task runs, which sampled only the easiest tasks in the domain. clm-0037
  • 2026-08-10claimPaired on identical tasks, q8_0 KV costs TURN EFFICIENCY: 228 turns against f16's 164 over the same 9 tasks, +39%, taking more turns on 6 of 9 and fewer on 1. clm-0038
  • 2026-08-10claimOn reward the 122B and Nemotron are indistinguishable, but reward is the wrong headline: on tasks both get RIGHT, the 122B needs 19 turns and 2.0 minutes against Nemotron's 26 and 3.7 — 37% fewer loops and 85% less wall-clock to the same correct answer. clm-0039
  • 2026-08-10supersededby clm-0042's per-task measurement, which found the true energy cost roughly 10x lower — this run's figures are whole-arm totals padded by model loading and non-scoring tasks, not the model's actual energy per answer, and must not be used to rank models. clm-0040 clm-0042
  • 2026-08-10claimThe KV dequant patch cuts energy 42% at 200k context — 146.1 Wh unpatched against 85.3 Wh patched for the same throughput benchmark — and patched q8_0 (85.3 Wh) beats f16 (89.6 Wh). clm-0041
  • 2026-08-10claimPer-task energy windows, matched to a common task set across arms, put a correct τ² answer at 6.81 Wh on the 122B with f16 KV, 9.48 Wh with q8_0, and 12.75 Wh on Nemotron. clm-0042
  • 2026-08-10claimEvery τ² arm was run with --user-llm set to the same model as --agent-llm, so cross-model comparisons changed the agent AND the user simulator together — exactly what the runbook forbids ("hold both --user-llm and the judge fixed across comparisons, or results re-baseline silently"). clm-0043
  • 2026-08-10gate7 candidates → acquired: cascade2-30b, gemma4-26b, glm-47-flash, laguna-s-21, lfm2-24b, muse-glimmer-30b, qwen36-27b-mtp
  • 2026-08-10gate5 candidates → screened: cascade2-30b, gemma4-26b, glm-47-flash, laguna-s-21, lfm2-24b
  • 2026-08-10gate16 candidates → listed: deepseek-v4-flash, gemma4-12b, glm-45-air, ling-30-flash, llama4-scout, maple-preview, muse-glimmer-30b, npu-embeddinggemma, npu-embeddinggemma, npu-lfm2, npu-qwen3-4b-thinking, npu-qwen3-4b-thinking, npu-whisper, npu-whisper, qwen38-27b, qwen38-27b clm-0031 clm-0034 clm-0015 clm-0018
  • 2026-08-10gate4 candidates → blocked: laguna-s-21, laguna-s-21, muse-glimmer-30b, muse-glimmer-30b
  • 2026-08-10runs2 runs landed on cfg-0026, cfg-0027 (tau2-bench-airline) cfg-0026 cfg-0027
  • 2026-08-09supersededthe Pass^1 = 1.000 reported by this run came from a 3-task subsample biased toward the domain's easiest tasks — the other two of the original five never terminated and were excluded as infrastructure errors. clm-0030 clm-0037
  • 2026-08-09claimA community DeepSeek-V4-Flash report independently confirms our mmap/GTT double-residency finding, demonstrates a 120 GiB GTT ceiling in production use, and — most consequentially — reports Vulkan BEATING ROCm on 3 of 4 cells including 55% faster decode at depth. clm-0031
  • 2026-08-09claimThe tau2 "runaway" tasks are a failure to escalate — and clm-0033 later established the failure is CAUSED BY THINKING, which this claim wrongly ruled out. clm-0032
  • 2026-08-09claimOn tau2-bench airline, thinking is a NET NEGATIVE for this model: thinking OFF solves 5/5 tasks at reward 1.0 in ~10 minutes, while thinking ON solves 3/5 and deadlocks indefinitely on the other two (4h15m and 2h10m in unbounded runs). clm-0033
  • 2026-08-09claimdomdoss/Warden (unrelated project, coincidental name) implements the multi-model architecture we have been designing toward — a small local orchestrator routing to named specialists with per-agent model selection. clm-0034
  • 2026-08-09retractionthe per-model reward rankings and quantised-KV cost reported by this run do not hold — they came from 5-task tau2 arms whose ~0.40 run-to-run noise and 40-step cap bias were only characterised afterward (clm-0036), so the reward numbers below are not usable. clm-0035 clm-0037 clm-0039
  • 2026-08-09claimτ²-bench at 5 tasks cannot resolve the differences drawn from it in this project's capability matrix. clm-0036
  • 2026-08-09gatedeepseek-v4-flash: listed — Community report worth testing directly. clm-0031
  • 2026-08-09gategemma4-12b: listed — Evidence from a comparable project that a 12B suffices for the orchestrator role. clm-0034
  • 2026-08-09gatenemotron3-super: screened → benched — Ran the tau2 arms. clm-0035 clm-0036
  • 2026-08-09gateqwen36-35b: screened → benched — Throughput, FA, ngram speculation and tau2 arms run. clm-0026
  • 2026-08-09runs1 run landed on cfg-0006 (tau2-bench-airline) run-0098
  • 2026-08-09runs1 run landed on cfg-0006 (tau2-bench-airline-nothink) run-0099
  • 2026-08-09runs1 run landed on cfg-0025 (tau2-bench-airline) run-0100
  • 2026-08-08claimOn aihydra (gfx1151, ROCm 7.1, llama.cpp 3653e6d), running llama-server with `--parallel 4` destroys long-context needle retrieval — 0/8 across controlled trials — while `--parallel 1` on the same build, model and prompt succeeds 8/8. clm-0019
  • 2026-08-08claimOn ROCm/gfx1151 with a stock llama.cpp, flash attention is unambiguously BETTER at depth — at 32k it is worth 1.22x prefill and 1.84x decode on the 122B MoE — which is the opposite of the Vulkan cliff reported in clm-0017. clm-0020
  • 2026-08-08claimThe dense-model flash-attention prefill cliff reported in clm-0017 does NOT exist on ROCm. clm-0021
  • 2026-08-08claimThe community KV-dequantisation fix is real, large, and scales monotonically with depth: one cherry-picked commit recovers +18.3% at 32k, +55.6% at 131k and **+70.3% at 204,800 — production's own context** — while leaving f16 unchanged at every depth. clm-0022
  • 2026-08-08claimQwen3.6-35B-A3B is 2.3x the 122B's decode on identical hardware (51.01 vs 21.90 tok/s at empty context) and holds 42.45 at 32k, with prefill above 1000 tok/s. clm-0023
  • 2026-08-08claimThe tool-grammar ceiling does not reproduce on llama.cpp 3653e6d. clm-0024
  • 2026-08-08claimFour models measured on identical hardware give decode from 17.51 to 55.45 tok/s, and the bandwidth model predicts the ORDER but not the magnitude — realised efficiency ranges from 34% to 62% of the theoretical ceiling. clm-0025
  • 2026-08-08claimSpeculation is close to worthless on Qwen3.6-35B-A3B with varied prompts — ngram-mod gives 1.11x with a 29% coefficient of variation, ngram-cache gives nothing, and MTP is unavailable because the model carries no NextN layers. clm-0026
  • 2026-08-08claimPrefix cache reuse is worth 9.8x on this box — an 8,000-token prefix costs 25.99 s cold and 2.65 s warm — and it is strictly PREFIX-ANCHORED: prepending three characters to an otherwise identical prompt returns it to full cold cost (26.26 s), zero reuse despite 99.9% identical content. clm-0027
  • 2026-08-08claimAcross four models on gfx1151, flash attention is worth 2.5x to 4.8x DECODE at 131k and its absence is catastrophic — a 35B loses 88% of its decode speed from empty context to 131k without it, against 44% with it. clm-0028
  • 2026-08-08claimDIAGNOSED: `llama-perplexity` produces garbage on this build — PPL 532 on deliberately repetitive English that should score 2-5, 3,654 on wikitext for a model that generates coherently at 51 tok/s and passes tool-calling guards. clm-0029
  • 2026-08-08gategpt-oss-120b: screened → benched — Throughput rows recorded. clm-0025
  • 2026-08-08gateqwen35-122b: screened → benched — Fully characterised on performance: depth to 204.8k, KV quant both ways, FA both ways, speculation curve, cache trace, grammar ceiling. clm-0022 clm-0027
  • 2026-08-08runs1 run landed on cfg-0006 (capability-guard) run-0005
  • 2026-08-08runs2 runs landed on cfg-0006, cfg-0007 (spec-ab) cfg-0006 cfg-0007
  • 2026-08-08runs84 runs landed on cfg-0008, cfg-0009, cfg-0010, cfg-0012, cfg-0015, cfg-0016, cfg-0017, cfg-0018, cfg-0019, cfg-0020, cfg-0021, cfg-0022, cfg-0023, cfg-0024 (llama-bench) cfg-0008 cfg-0009 cfg-0010 cfg-0012 cfg-0015 cfg-0016 cfg-0017 cfg-0018 cfg-0019 cfg-0020 cfg-0021 cfg-0022 cfg-0023 cfg-0024
  • 2026-08-08runs6 runs landed on cfg-0011, cfg-0013, cfg-0014 (llama-bench) cfg-0011 cfg-0013 cfg-0014
  • 2026-08-06claimLing-3.0-flash (Ant Group, 26 Jul 2026) is a 124B/5.1B-active hybrid MoE that is a near-ideal controlled comparison for our Qwen3.5-122B-A10B — same footprint, half the active parameters — but it cannot be benchmarked on a stock llama.cpp, and its MTP head ships INACTIVE, which would silently rig any decode comparison in Qwen's favour. clm-0015
  • 2026-08-06claimMeasured on Strix Halo: Qwen3.6-35B-A3B under ROCmFP4 + HIP + ngram-mod at parallel 4 sustains 121 tok/s across 500 varied IFEval prompts against a 64.8 tok/s no-speculation floor — a 1.87x production speedup with IFEval-strict at 78.6% — while prefill collapses from 1,211 tok/s cold to 136 tok/s at ~243k depth. clm-0016
  • 2026-08-06claimOn Strix Halo Vulkan, a single unmerged patch — contiguizing strided f16 KV data before the flash-attention prefill — removes a dense-model prefill collapse worth 2.5x at 32k and 6.6x at 65k, changes decode not at all, and collapses run-to-run scatter from 5-10% to 0.3%. clm-0017
  • 2026-08-06claimDeepGrove Maple-Preview (20.2B-A1.49B ternary, 5.31 GB, MIT) is a credible second-lane candidate for the Mac mini, but not a Warden candidate — its own model card concedes underperformance on agentic benchmarks. clm-0018
  • 2026-08-05claimIdentity-grounded persistent sessions with a shared memory and peer-to-peer messaging can self-organise into useful working groups without an orchestrator. clm-0008
  • 2026-08-05claimA ROCmFP4 iMatrix quant of our exact model (Qwen3.5-122B-A10B) reports 60.70 GiB and 28.505 tok/s decode with MTP OFF — against our measured 20.96 tok/s MTP-off at UD-Q4_K_M. clm-0009
  • 2026-08-05claimDraft-free ngram speculation (ngram-mod) reportedly beats MTP by a wide margin on repetitive/agentic work on Strix Halo — 71 t/s unspeculated to 216 solo on a code-edit probe, with a shared hash pool letting concurrent streams feed each other's drafts (247 pooled across 4 streams, later 302 end-to-end on HIP). clm-0010
  • 2026-08-05claimGreedy decode is NOT run-to-run deterministic on the Vulkan backend — same config, same prompt, temperature 0, batch 1, no speculation, three different outputs. clm-0011
  • 2026-08-05claim128 GB Strix Halo systems appear to be repricing upward materially — from a roughly $2,500-4,000 enthusiast tier toward $4,500-6,000 — with 128 GB SKUs scarce while 64 GB configurations remain available. clm-0012
  • 2026-08-05claimQwen3.6-35B-A3B — our chosen reflex model — is published in FastFlowLM's .q4nx NPU format (23.2 GB, plus a 1.0 GB vision encoder), making a genuinely capable MoE, not a 1-2B classifier, runnable on the XDNA2 NPU. clm-0013
  • 2026-08-05claimOn gfx1151, Vulkan measured ~22-24% faster than ROCm on the same 35B-A3B model and build — pp4096 1039 vs 840 t/s, tg128 53.1 vs 43.4 — but the Vulkan build produced GARBAGE OUTPUT for that model, so the numbers describe a broken configuration. clm-0014
  • 2026-08-03claimStatic pool utilisation did not predict the 2026-07-21 OOM. clm-0001
  • 2026-08-03claimMulti-token prediction gains MORE at heavier quantisation, not less. clm-0002
  • 2026-08-03claimQuantized KV cache is mis-implemented for Strix Halo in stock llama.cpp: the code dequantizes to full precision repeatedly during inference, which a discrete GPU hides in cache and this box cannot. clm-0006
  • 2026-08-03claimA GPU cache wall around 32-40 MB, past which read speed drops roughly 4x, is a candidate explanation for our own prefill degradation from ~354 tok/s at small context to ~144 tok/s at 128K. clm-0007

July 2026

  • 2026-07-30gatemaple-preview: listed — Ternary quantisation at a footprint that opens a second lane on hardware we already own. clm-0018
  • 2026-07-26gateling-30-flash: listed — The only candidate that functions as a controlled experiment rather than another data point.
  • 2026-07-23incidentAbrupt power loss under sustained load; board never POSTed again — RMA submitted (blocked, unresolved) inc-0005
  • 2026-07-21gatelaguna-s-21: listed — Released 2026-07-21 with a Terminal-Bench number far above our incumbent, at a footprint that fits.
  • 2026-07-21incidentUnified-memory OOM cascade to kernel panic, then a wedged boot (blocked, resolved) inc-0004
  • 2026-07-21contributionggml-org/llama.cpp fork — carrying: Disk slot save/restore silently loses all prompt reuse on hybrid/recurrent models because context checkpoints are never persisted. con-0001
  • 2026-07-20claimMTP's speedup tracks how predictable the text is: +81% on code, +21% on freeform, with the real agent workload mix landing around +29-40%. clm-0003
  • 2026-07-20runs1 run landed on cfg-0002 (mtp-decode-sweep) run-0002

June 2026

  • 2026-06-28claimgpt-oss-120b placed last on our blind conceptual set — 3/8, mean 1.38 — against Qwen3.6-35B-A3B's 7/8, mean 2.12. clm-0004
  • 2026-06-28claimNemotron 3 Super's benchmark result is NOT quant-equivalent to its peers and must not be read as a like-for-like verdict on the model. clm-0005
  • 2026-06-28runs1 run landed on cfg-0001 (conceptual-c1-c8-blind) run-0003
  • 2026-06-27gate5 candidates → acquired: gpt-oss-120b, nemotron3-super, qwen35-122b, qwen36-27b-mtp, qwen36-35b
  • 2026-06-27gate5 candidates → screened: gpt-oss-120b, nemotron3-super, qwen35-122b, qwen36-27b-mtp, qwen36-35b
  • 2026-06-27runs2 runs landed on cfg-0001 (tool-loop) cfg-0001
  • 2026-06-20gate10 candidates → listed: cascade2-30b, gemma4-26b, glm-47-flash, gpt-oss-120b, lfm2-24b, llama4-scout, nemotron3-super, qwen35-122b, qwen36-27b-mtp, qwen36-35b