Home › Log
Log
The lab notebook, generated from the record: claims as they land, candidates moving through the gates, run series, incidents and upstream contributions. Retractions and supersessions appear inline in the same timeline, badged — corrections are entries here like any other learning, and a generated filter of this stream collects them. Subscribe: RSS.
357 entries · 7 of them corrections · every entry derives from a dated record — nothing on this page is written by hand
September 2026
- 2026-09-06claimOn its proposed dense-WORKER serving config (cfg-0183: UD-Q4_K_XL, the pr27311 leak-fix build, self-speculative draft-mtp n_max 2, --parallel 4, reasoning_effort low, greedy, fronted by the slotpin proxy), Qwen3.8-27B scored 91.7% on tau2-bench airline tasks 0-25 — 22 of 24 scored passed, run end-to-end through the proxy under a real 4-slot concurrent workload, with 0 empty assistant turns over 106 minutes. clm-0129
- 2026-09-06claimThe Qwen3.8-27B dense worker earns its slot on THROUGHPUT UNDER CONCURRENCY, not single-stream speed, and only on the leak-fixed build. clm-0130
- 2026-09-06claimThe Qwen3.8-27B worker config cost 15.18 Wh per correct answer on the tau2 0-25 run (333.87 Wh whole-session across 22 correct, mean 189.6 W, 0.46 pence at 30.3 p/kWh; eng-0278) — well under the 38.29 Wh per correct of the earlier single-slot Q8_0 run (eng-0081). clm-0131
- 2026-09-06claimThe reproduced pwilkin/ilintar Strix Halo stack (IQ4_XS-imatrix + DFlash2 draft on the retained-PM4 runtime, cfg-0185) scored 83.3% on tau2 airline 0-25 — 20 of 24 scored passed, 4-slot, reasoning low, 2 cloud-user-sim infra errors excluded — versus the worker's 91.7% (run-0658) on the identical task set, harness (668d3bc), seed, reasoning and concurrency. clm-0132
- 2026-09-06claimThe pwilkin candidate wins single-stream throughput decisively but only ties the worker under concurrency. clm-0133
- 2026-09-06claimThe pwilkin candidate cost 10.26 Wh per correct answer on its tau2 0-25 run (205.2 Wh whole-session across 20 correct, mean 187 W, 0.31 pence at 30.3 p/kWh; eng-0280) — below the worker's 15.18 (eng-0278). clm-0134
- 2026-09-06claimThe retained-PM4 runtime is a real but small, lossless dispatch optimisation — NOT the source of the pwilkin speed advantage. clm-0135
- 2026-09-06gateqwen38-27b: benched → benched — WORKER-CONFIG CAMPAIGN (2026-09, aihydra) — a full HaloBench on a DIFFERENT serving profile from the 2026-08 screen's cfg-0066, aimed at the dense-worker role (capability + throughput beside a faster interactive model), not a re-measurement of the Q8_0 pick. clm-0129 clm-0130 clm-0131 run-0658 run-0660 cfg-0183 con-0018
- 2026-09-06gateqwen38-27b: benched → benched — THIRD-PARTY STACK SCREEN (pwilkin/ilintar + halo-box), gate unchanged. clm-0132 clm-0133 clm-0134 clm-0135 run-0665 run-0668 con-0019
- 2026-09-06runs1 run landed on cfg-0185 (tau2-bench-airline) run-0665
- 2026-09-06runs1 run landed on cfg-0185 (empty-output-monitor) run-0666
- 2026-09-06runs1 run landed on cfg-0185 (slotpin-mslot) run-0667
- 2026-09-06runs4 runs landed on cfg-0186 (decode-depth-served) cfg-0186
- 2026-09-06runs1 run landed on cfg-0187 (pm4-isolation) run-0672
- 2026-09-06runs1 run landed on cfg-0187 (pm4-isolation) run-0673
- 2026-09-05runs1 run landed on cfg-0183 (tau2-bench-airline) run-0658
- 2026-09-05runs1 run landed on cfg-0183 (slotpin-proxy-metrics) run-0659
- 2026-09-05runs1 run landed on cfg-0183 (empty-output-monitor) run-0660
- 2026-09-05runs4 runs landed on cfg-0184 (decode-depth-served) cfg-0184
- 2026-09-03claimQwen3.8-Flash-Next ROCmFP4-FAST-v2-ple16 (agention imatrix quant, 87.06 GiB @ 4.23 bpw, everything GPU-resident incl. the per-head n-gram/PLE table) on the agention/Laurent Vulkan fork (cfg-0182, LaurentZuijdwijk/llama.cpp branch vulkan/qwen4exp-rocmfpx commit 5e085d12, build b10809, gfx1151/RADV) — first benchmark of the ROCmFP4 quant path on this fleet. clm-0128
- 2026-09-02gateqwen38-flash-next: benched → benched — Third serving series benched, distinct from both the Unsloth UD-Q4_K_XL first-look and the KingJones full-STRIX arm: the agention ROCmFP4-FAST-v2-ple16 imatrix quant (87.06 GiB @ 4.23 bpw, per-head n-gram/PLE, fully GPU-resident) on the agention/Laurent Vulkan fork (cfg-0182, LaurentZuijdwijk/llama.cpp vulkan/qwen4exp-rocmfpx commit 5e085d12, build b10809; branch tip verified, quant byte + sha256 verified). clm-0128
- 2026-09-02runs1 run landed on cfg-0182 (tau2-airline-full26) run-0653
- 2026-09-02runs4 runs landed on cfg-0182 (decode-depth-served) cfg-0182
August 2026
- 2026-08-30claimThe KingJones Qwen3.8-Flash-Next full-STRIX ROCmFP4 series is admitted only as BOUNDED/PARTIAL on AI Hydra (cfg-0180/cfg-0181). clm-0127
- 2026-08-30gateqwen38-flash-next: acquired → benched — The separate KingJones full-STRIX ROCmFP4 campaign completed on the exact HIP/gfx1151 runtime and was independently admitted BOUNDED/PARTIAL. clm-0127 eng-0276 inc-0011 inc-0012
- 2026-08-30incidentFull-STRIX campaign retained corrected QSA evidence with bounded instrumentation gaps (degraded, resolved) inc-0010
- 2026-08-30incidentAuthorized energy join handoff was omitted before full-STRIX evidence closure (degraded, resolved) inc-0011
- 2026-08-30incidentInherited 8k output ceiling bound two full-STRIX Tau2 calls (degraded, resolved) inc-0012
- 2026-08-30runs1 run landed on cfg-0180 (tau2-airline-full26) run-0639
- 2026-08-30runs5 runs landed on cfg-0180 (qsa-cache-false-prompt-processing) cfg-0180
- 2026-08-30runs1 run landed on cfg-0180 (qsa-full-window-feasibility) run-0645
- 2026-08-30runs2 runs landed on cfg-0180 (qsa-prefix-reuse-trace) cfg-0180
- 2026-08-30runs5 runs landed on cfg-0181 (llama-bench-source-config-mmap-raw) cfg-0181
- 2026-08-29supersededOn AI Hydra's admitted HIP/gfx1151/native/Release host boundary, the KingJones R2 Qwen3.8-Flash-Next full-STRIX ROCmFP4 campaign fixed model artifact c6770d7442a06bf1d78edf28cec83e1ec93afdd34664c23ff898807b6b9349fa (121838036032 bytes, publisher revision 069dddb53bab04218d734fa9a771f8a0242ab059), runtime source 36e9acd40e10a87cd3c3ef8ec734668757dc8520, and the independently admitted post-route receipt patch 1f1c3bed910922b415c1be36c9a04c9b7fedc4162aedf504a9da007daebcb4d2. clm-0126 clm-0127
- 2026-08-29gateqwen38-flash-next: acquired → acquired — The separate KingJones R2 full-STRIX ROCmFP4 artifact/runtime campaign stopped before model execution. clm-0126
- 2026-08-29incidentQ38FN R2 receipt-only capture variant stopped at header include resolution (blocked, unresolved) inc-0009
- 2026-08-29runs1 run landed on cfg-0179 (q38fn-kingjones-r2-targeted-receipt-build-gate-v1) run-0637
- 2026-08-29runs1 run landed on cfg-0180 (tau2-airline-smoke5) run-0638
- 2026-08-28claimThe Qwen3.8-Flash-Next records from 2026-08-27 remain visible as a provisional first look: cfg-0177 and run-0632 record the reasoning-low capability arm, cfg-0178 and run-0633 through run-0636 record the non-served engine-floor depth ladder, and eng-0275 records the joined wall-meter window. clm-0125
- 2026-08-28gateqwen38-flash-next: benched → acquired — Independent review retained every first-look record but held its current leader, energy, production and role-fit interpretation. clm-0125
- 2026-08-27supersededQwen3.8-Flash-Next UD-Q4_K_XL (Qwen4-arch preview MoE: 125B MoE / ~6B active + 51B N-gram/PLE tables + vision) on the Unsloth qwen4exp Vulkan build (cfg-0177, commit 250b6144, gfx1151/RADV) — first full benchmark. clm-0124 clm-0125
- 2026-08-27gateqwen38-flash-next: acquired → screened — Screen tier on-box, Vulkan r3 build 250b6144.
- 2026-08-27gateqwen38-flash-next: screened → benched — Built the Unsloth qwen4exp fork at HEAD 250b6144 (Vulkan/RADV; the ROCm mmap-load path stalls on gfx1151, so Vulkan is the measured backend). clm-0124
- 2026-08-27contributionkingjones777/Qwen3.8-Flash-Next-ROCmFP4-STRIX-GGUF report — carrying: Immutable publisher model-page revision 069dddb53bab04218d734fa9a771f8a0242ab059 describes the full-STRIX artifact and the required kingjones30/ROCmFPX runtime. con-0017
- 2026-08-27runs1 run landed on cfg-0177 (tau2-airline-full26) run-0632
- 2026-08-27runs4 runs landed on cfg-0178 (llama-bench-decode-depth) cfg-0178
- 2026-08-26gateqwen38-flash-next: listed — Qwen4-arch preview released.
- 2026-08-26gateqwen38-flash-next: listed → acquired — UD-Q4_K_XL (the competence-validated Q4 class) staged to the Ladon library and byte/sha256-verified: 4 shards totaling 111,334,654,784 bytes (shard-1 sha256 4448186216…), matching the current HF revision exactly.
- 2026-08-26contributionggml-org/llama.cpp pr — open: PR #27311 "Scheduler UMA ring buffer (+ sanitizer and fixes)" adds an input ring buffer on UMA devices so the host cannot clobber in-flight graph inputs. con-0018
- 2026-08-26contributionROCm/rocm-systems fork — carrying: "Retained PM4 dispatch" — a custom AMD HIP/ROCm runtime (rocm-systems ilintar-experiments @ 78d1160, the CLR/hipamd graph path) that keeps a graph's low-level PM4 command buffer resident and replays it, cutting per-token launch overhead. con-0019
- 2026-08-25contributionNathanw1014/strix-halo-llamacpp report — open: Primary-source runtime-watch record: Nathanw1014/strix-halo-llamacpp v0.6.11, published 2026-08-24, fixes the fork-only image-request regression introduced by the fork's v0.6.8 DFlash M-RoPE change. con-0014
- 2026-08-24claimA Qwen3.8-27B ROCmFPX-Q4 plus DFlash2 server screening sample reported 52.2 t/s on a varied code prompt and 31.5 t/s on a varied prose prompt. clm-0120
- 2026-08-24claimA 45-turn Vulkan cache-reuse trace completed without a logged DeviceLost or lockup signature. clm-0121
- 2026-08-24claimNo Qwen3.8-27B multi-slot serving or concurrency-throughput claim is admitted from this autonomous handoff. clm-0122
- 2026-08-24claimNo Kairic Edge versus UD-Q4_K_XL head-to-head verdict is admitted from this autonomous handoff. clm-0123
- 2026-08-24gateqwen38-27b: benched → benched — COMPARABILITY ANNOTATION + v1.2 RUNTIME GATE (HK-V12-ANNOTATE, per hbreviewer verdict t_b05473c9; no new measurement). clm-0116 clm-0118 clm-0119 con-0012
- 2026-08-24gateqwen38-27b: benched → benched — AUTONOMOUS TUNING HANDOFF AUDIT — bounded screening evidence ingested without changing the measured production selection. clm-0120 clm-0121 clm-0122 clm-0123
- 2026-08-24contributionLaurentZuijdwijk/llama.cpp fork — carrying: Vulkan fork used by the Qwen3.8-27B ROCmFPX/DFlash2 autonomous tuning handoff. con-0013
- 2026-08-24contributionggml-org/llama.cpp issue — open: Raw upstream diagnostic comment by frizikk on open issue #25618, retrieved directly from GitHub on 2026-08-25. con-0015
- 2026-08-24contributionggml-org/llama.cpp pr — open: Raw upstream comment by alexpooley on open PR #25863, retrieved directly from GitHub on 2026-08-25. con-0016
- 2026-08-24runs1 run landed on cfg-0174 (qwen38-prod-tuning.server-screen-code) run-0626
- 2026-08-24runs1 run landed on cfg-0174 (tau2-airline-smoke) run-0627
- 2026-08-24runs1 run landed on cfg-0175 (qwen38-prod-tuning.cache-reuse-trace) run-0628
- 2026-08-24runs1 run landed on cfg-0176 (qwen38-prod-tuning.server-depth) run-0629
- 2026-08-24runs1 run landed on cfg-0176 (qwen38-prod-tuning.server-depth) run-0630
- 2026-08-24runs1 run landed on cfg-0174 (qwen38-prod-tuning.server-screen-prose) run-0631
- 2026-08-23claimNull-update diagnostic: the v0.5.x base-platform Stage-A sentinel empty-content defect is RESOLVED on the reviewed v0.6.10 portable Vulkan/RADV runtime (2586f6edd). clm-0109
- 2026-08-23claimOn the v0.6.10 strix-halo fork under Vulkan/RADV (build 2586f6edd), Qwen3.6-35B-A3B-MTP UD-Q4_K_M with reasoning OFF recovers MTP n2 capability that build 7077abb-upstream-on-ROCm0 destroyed. clm-0110
- 2026-08-23claimOn Qwen3.5-122B-A10B UD-Q4_K_M (build ce7689f, llama.cpp-kvfix tree, ROCm0 Radeon 8060S gfx1151 122880 MiB, q8_0/q8_0 KV baseline) the production-optimisation matrix interior conclusions hold (internal comparison, same model/quant/build): (1) interactive first-token latency (ttft) is flat at ~1420 ms across batch 512-2048 at ub=512 (all N=3, CV<1.6%), and degrades when ubatch drops below 512 — ub256 ttft ~1683-1691 ms, ub128 ttft ~2364 ms — so ub=512 is the interactive optimum and tpot stays flat ~45.8 ms/tok throughout. clm-0111
- 2026-08-23claimOn Ornith-1.0-35B-UD-Q4_K_XL (v0.6.10 fork build 2586f6ed, Vulkan/RADV, f16/f16 KV, -ngl 999 -fa 1 --parallel 1 --load-mode none -t 16, plain decode) the interactive-decode batch/ubatch tuning result is effectively FLAT at the d131072 production-depth probe: decode runs 34.342-34.430 tps across b512/ub512 (34.357, CV 0.073%), b1024/ub512 (34.430, CV 0.321%), b2048/ub1024 (34.363, CV 0.055%) and b4096/ub2048 (34.342, CV 0.089%) - a 0.26% spread with max CV 0.32%. clm-0112
- 2026-08-23claimOn the v0.6.10 strix-halo fork under Vulkan/RADV (build 2586f6edd), Qwen3.6-35B-A3B-MTP UD-Q4_K_M: raising native-MTP draft n_max from 2 to 4 (reasoning OFF) is FLAT — n4-off full-26 scored 17/26 (mean 0.653846), identical to n2-off (17/26, run-0578); Fisher exact two-sided n4-vs-n2 p=1.000000. clm-0113
- 2026-08-23claimThe llama.cpp MTP/spec-path multi-GPU fragility corroborates across a second, platform, with a validated workaround. clm-0114
- 2026-08-23claimA high fixed MTP draft-depth CEILING carries a measurable fixed cost that is INDEPENDENT of the adaptive logic: on llama.cpp PR #27210 (draft-mtp-adaptive), stew675 held the algorithm at fixed depth 3 but kept max depth 10 and found merely having max depth 10 incurred a fixed ~2.6% performance penalty on its own; the subsequent fix (2026-08-23T01:32:47Z, limit the full MTP buffer scan on truncated/short drafts) recovers ~2-3% for --spec-draft-n-max 10 --spec-draft-p-min >0.5 configs. marcusds independently confirms on a single RTX 5090 that depth CHANGING does not degrade CUDA-graph performance (GGML_CUDA_DISABLE_GRAPHS=1 vs default deltas small, baseline C0 -3.0% / -1.9% / -1.8% / -1.7%), and that adaptive 3..10 + p-min is the only config with a POSITIVE recall delta (+2.6%) — so the ~2.6% is the cost of a static high ceiling, not of adaptive switching. clm-0115
- 2026-08-23claimHK-001-FULLBENCH-REVIEW record for the affected v1.1 Kairic Edge TheRock configuration (cfg-0164): Qwen3.8-27B-IU4-Kairic-Edge.gguf on the TheRock 7.14 / AMD clang 23.0.0 ROCm runtime, c32768, --kairic-edge --no-mmap. clm-0116
- 2026-08-23claimOn Qwen3.6-35B-A3B-MTP UD-Q4_K_M (build 2586f6edd strix-halo v0.6.10 fork/Vulkan RADV, capability c32768, plain mode -rea off, no MTP, f16/f16 KV baseline, temp 0 seed 42 max_tokens 4096) the production-optimisation matrix interior conclusions hold (internal, same model/quant/build): (1) interactive first-token latency (ttft) is minimally sensitive to batch once ubatch=512 — b1024/ub512 1353.1 ms (CV 1.46%), b2048/ub512 1357.4 (CV 2.01%), b512/ub512 1364.5 (CV 1.13%), all N=5 d131072 — and degrades when ubatch drops below 512 (ub256 ttft ~1538-1545 ms, ub128 2320 ms); tpot stays flat ~28.4-28.5 ms/tok across all six arms, so b1024/ub512 is the interactive/ttft production choice (35.1 t/s decode). clm-0117
- 2026-08-23claimHK-RERUN-REASONOFF verdict on the Kairic Edge TheRock build with reasoning DISABLED (-rea off, cfg-0172) — the fix path clm-0116 prescribed for the reasoning-ON empty-assistant defect. clm-0118
- 2026-08-23claimHK-001-REASONING-MATRIX verdict on the Kairic Edge TheRock build of Qwen3.8-27B-IU4 (cfg-0173) — 5-cell tau2 smoke matrix isolating the reasoning lever on identical tasks/seeds/judge. clm-0119
- 2026-08-23gate5 candidates → benched: qwen36-35b, qwen38-27b, qwen38-27b, qwen38-27b, qwen38-27b clm-0113 clm-0110 clm-0109 run-0572 run-0573 run-0574 clm-0114 clm-0115 con-0010 con-0011 clm-0118 cfg-0172 run-0618 run-0620 clm-0119 cfg-0173 run-0621 run-0625
- 2026-08-23contributionciru-ai/ROCmFPX fork — carrying: The Kairic Edge Qwen3.8-27B-IU4 build (branch kairic-edge-qwen38-27b-v1.1, commit e1da26bb8) from ciru-ai/ROCmFPX — a vendor-authored fork that binds the ROCmFPX binary (called "TheRock" stack here because the carrier vendor certifies it against TheRock 7.14 / AMD clang 23.0.0) to a custom 4-bit IU4 sidecar layout: a Q4-ish quant GGUF (Qwen3.8-27B-IU4-Kairic-Edge.gguf) + three PROMPTFORGE_* .pfs sidecars (FFN, GDN, GDN-Output) that accelerate the FFN and GDN projector paths on gfx1151, plus a --kairc-edge server flag and a compat mode (KAIRIC_EDGE_COMPATIBILITY_MODE) that enables tool-calling where fast-greedy forbids grammar/tool calls. con-0012
- 2026-08-23runs3 runs landed on cfg-0146, cfg-0147, cfg-0148 (stageA-c32768-varied-prompts) cfg-0146 cfg-0147 cfg-0148
- 2026-08-23runs7 runs landed on cfg-0149, cfg-0150, cfg-0151, cfg-0157, cfg-0158, cfg-0159, cfg-0160 (guard-c32768-depth8000) cfg-0149 cfg-0150 cfg-0151 cfg-0157 cfg-0158 cfg-0159 cfg-0160
- 2026-08-23runs3 runs landed on cfg-0149, cfg-0150, cfg-0160 (tau2-bench-airline) cfg-0149 cfg-0150 cfg-0160
- 2026-08-23runs6 runs landed on cfg-0151, cfg-0152, cfg-0153, cfg-0154, cfg-0155, cfg-0156 (ho012-122b-prod-matrix-r2.cell1) cfg-0151 cfg-0152 cfg-0153 cfg-0154 cfg-0155 cfg-0156
- 2026-08-23runs1 run landed on cfg-0151 (ho012-122b-prod-matrix-r2.cell2-depth) run-0585
- 2026-08-23runs1 run landed on cfg-0157 (ho012-122b-prod-matrix-r2.cell3-f16) run-0586
- 2026-08-23runs1 run landed on cfg-0158 (ho012-122b-prod-matrix-r2.cell3-q4_0) run-0587
- 2026-08-23runs1 run landed on cfg-0159 (ho005-revised-ornith-q4kxl-batch-r3.control-b512-ub512) run-0592
- 2026-08-23runs1 run landed on cfg-0159 (ho005-revised-ornith-q4kxl-batch-r3.b1024-ub512) run-0593
- 2026-08-23runs1 run landed on cfg-0159 (ho005-revised-ornith-q4kxl-batch-r3.b2048-ub1024) run-0594
- 2026-08-23runs1 run landed on cfg-0159 (ho005-revised-ornith-q4kxl-batch-r3.b4096-ub2048) run-0595
- 2026-08-23runs1 run landed on cfg-0159 (ho005-revised-ornith-q4kxl-batch-r3.anchor-d262144) run-0596
- 2026-08-23runs1 run landed on cfg-0160 (eos-cliff-c32768-varied-fullprompt) run-0597
- 2026-08-23runs1 run landed on cfg-0162 (reasoning-sentinel-c32768) run-0600
- 2026-08-23runs1 run landed on cfg-0164 (hk001-fullbench-guard) run-0601
- 2026-08-23runs1 run landed on cfg-0164 (hk001-fullbench-fit) run-0602
- 2026-08-23runs2 runs landed on cfg-0164 (hk001-fullbench-depth-r3) cfg-0164
- 2026-08-23runs1 run landed on cfg-0165 (ho014a-fill-r1.guard-c32768-depth8000) run-0605
- 2026-08-23runs6 runs landed on cfg-0165, cfg-0166, cfg-0167, cfg-0168, cfg-0169, cfg-0170 (ho014a-fill-r1.cell1-latency) cfg-0165 cfg-0166 cfg-0167 cfg-0168 cfg-0169 cfg-0170
- 2026-08-23runs3 runs landed on cfg-0166 (ho014a-fill-r1.cell2-depth) cfg-0166
- 2026-08-23runs1 run landed on cfg-0166 (ho014-run-kv-overlap.guard-c32768-depth8000) run-0615
- 2026-08-23runs2 runs landed on cfg-0166, cfg-0171 (ho014-run-kv-overlap.cell3b-bench-pair) cfg-0166 cfg-0171
- 2026-08-23runs1 run landed on cfg-0172 (hk001-rerun-reasonoff.tau2-full26) run-0618
- 2026-08-23runs1 run landed on cfg-0172 (hk001-rerun-reasonoff.specsweep-spec-on) run-0619
- 2026-08-23runs1 run landed on cfg-0172 (hk001-rerun-reasonoff.specsweep-spec-off) run-0620
- 2026-08-23runs1 run landed on cfg-0173 (hk001-reasoning-matrix.R0-B4k) run-0621
- 2026-08-23runs1 run landed on cfg-0173 (hk001-reasoning-matrix.R1-B6k) run-0622
- 2026-08-23runs1 run landed on cfg-0173 (hk001-reasoning-matrix.R2-B8k) run-0623
- 2026-08-23runs1 run landed on cfg-0173 (hk001-reasoning-matrix.R3-B8k-L) run-0624
- 2026-08-23runs1 run landed on cfg-0173 (hk001-reasoning-matrix.R4-B8k-M) run-0625
- 2026-08-22claimLLaDA2.2-flash (100B-A13B block-diffusion MoE, Q4_K_S) solved 5/5 tau2-bench airline screening tasks (task_ids 0-4), mean reward 1.000, with real multi-turn tool use, running GPU-accelerated on aihydra (Strix Halo, gfx1151) via the headbouyJB/diffuse-cpp fork with the GPU-resident KV cache and full Levenshtein editing enabled. clm-0102
- 2026-08-22claimThe headbouyJB/diffuse-cpp fork makes a scaled block-diffusion LLM benchmarkable on portable AMD hardware: GPU offload of the MoE forward (gfx1151), OpenAI tool-calling, and a GPU-resident inter-step KV cache with cross-turn prompt reuse together take decode-at-long-context from ~1 tok/s (host-array cache) to ~5-12 tok/s, turning a multi-turn diffusion agent from impractical (~tens of seconds to minutes per turn) into a runnable tau2 screen (~1-3 min/task). clm-0103
- 2026-08-22claimggml_flash_attn_ext on this HIP/gfx1151 build does not fully honour an additive F16 block-causal mask: a committed prefix's per-layer K/V changed by several logits when a later masked block entered the attention window (compounding through layers), while the manual ggml_soft_max_ext path with the identical mask did not. clm-0104
- 2026-08-22superseded(operative claim) as of strix-halo-llamacpp v0.6.10. clm-0105
- 2026-08-22claimThe quantized-target divergence in the llama.cpp MTP/DSpark speculative-verify path corroborates across a NEW model, OS and GPU family: on NVIDIA Vulkan/Windows 11 (llama.cpp b10566 / bb4caa754, GTX 1660 SUPER) the official LiquidAI LFM2.5-2.6B pair diverges from vanilla on a Q8_0 target, while the SAME Q8_0 draft against an F16 target produces byte-identical greedy output to the F16 vanilla target (40/64 draft tokens accepted). clm-0106
- 2026-08-22claimThe strix-halo-llamacpp fork v0.6.9 MTP rollback-exactness tradeoff is REVERSED in v0.6.10: with the root-cause fix (server no longer re-verifies replayed draft tokens after a checkpoint restore, 9c5d899) and full-checkpoint MTP rollback re-applied (f25eefe), MTP rollback on hybrid GDN (qwen35moe) targets is token-exact again on the current stable fork, and the long-run stall is gone. clm-0107
- 2026-08-22claimHO-009-AB measured plain-vs-n2 capability A/B on Qwen3.6-35B-A3B-MTP at production agentic depth (c32768), reasoning OFF on both arms (build 7077abb, ROCm0/gfx1151, UD-Q4_K_M, f16/f16 KV). clm-0108
- 2026-08-22gatellada22-flash: listed — A genuinely novel offshoot: inclusionAI LLaDA2.2-flash, a 100B-A13B block-diffusion MoE with Levenshtein editing, is the first non-autoregressive agent model considered for the field.
- 2026-08-22gatellada22-flash: listed → screened — Screened on aihydra via the headbouyJB/diffuse-cpp fork (GPU-resident KV cache + cross-turn reuse + full Levenshtein editing). tau2 airline 5-task screen (task_ids 0-4): 5/5 solved, mean reward 1.000, real multi-turn tool use, all tasks converged in 8-10 turns (run-0553). run-0553 clm-0102
- 2026-08-22gateqwen36-35b: benched → benched — Fork v0.6.8/v0.6.9/v0.6.10 upstream context added (community report evidence, not a new HaloBench run and not a gate change). clm-0105 clm-0107 con-0007 con-0009
- 2026-08-22gateqwen38-27b: benched → benched — Upstream #25618 follow-up (RDNA4 R9700/gfx1201) added as community report evidence, not a new HaloBench run. clm-0104 con-0006
- 2026-08-22contributionAkicou/diffuse-cpp fork — carrying: Makes inclusionAI LLaDA2.2-flash (100B-A13B block-diffusion MoE) run GPU-accelerated on AMD Strix Halo (gfx1151) and benchmarkable as an agent. con-0005
- 2026-08-22contributionggml-org/llama.cpp report — open: New llama.cpp #25618 follow-up by snick525 (2026-08-22T02:11:18Z) on RDNA4 (AMD Radeon AI PRO R9700 / gfx1201, 32 GB, Vulkan/radv), target Qwen3.8-27B Q6_K_XL (MTP head intact), greedy (temperature 0, top_k 1, top_p 1, seed 42), f16 K/V. con-0006
- 2026-08-22contributionNathanw1014/strix-halo-llamacpp report — open: Nathanw1014/strix-halo-llamacpp (the Strix Halo Vulkan fork tracked in con-0002) shipped v0.6.8 and v0.6.9, two stable releases that carry a deliberate MTP rollback EXACTNESS tradeoff on hybrid GDN targets. v0.6.8 (2026-08-22T02:40:00Z) added DFlash2 speculative decoding support for mmproj/vision (community-reported vision-prompt HTTP 500 fixed by injecting draft rows at dense per-token positions and triming the draft cache in token space) and introduced a one-line change, "common: use full checkpoints for MTP rollback", that forced MTP rollback through full sequence-state checkpoints. v0.6.9 (2026-08-22T04:09:43Z) REVERTS that one line (revert commit a17e843, payload Nathanw1014/llama.cpp@a17e8432b, branch strix-halo-vulkan): on Vulkan with a hybrid GDN target (qwen35moe, e.g. con-0007
- 2026-08-22contributionggml-org/llama.cpp report — open: llama.cpp #25618 comment #16 by F-Mangini (2026-08-22T20:32:01Z) corroborates the quantized-target divergence on a NEW model/OS/GPU family: the official LiquidAI LFM2.5 DSpark pair also diverges from vanilla on a quantized target while an F16 target preserves greedy parity. con-0008
- 2026-08-22contributionNathanw1014/strix-halo-llamacpp report — open: Nathanw1014/strix-halo-llamacpp (the Strix Halo Vulkan fork tracked in con-0002) shipped STABLE v0.6.10 (2026-08-22T12:03:16Z) that REVERSES the v0.6.8/v0.6.9 MTP rollback-EXACTNESS tradeoff recorded in con-0007/clm-0105. con-0009
- 2026-08-22contributionggml-org/llama.cpp report — open: llama.cpp #27122 comment by mazinist (2026-08-22T23:51:13Z) independently confirms the MTP/CUDA multi-GPU lockup originally reported by tripletto, and validates a workaround on a completely different platform from the earlier zyxyunxin note (#issuecomment-5310571689, 2026-08-17). con-0010
- 2026-08-22runs4 runs landed on cfg-0143, cfg-0030, cfg-0144, cfg-0145 (tau2-bench-airline) cfg-0143 cfg-0030 cfg-0144 cfg-0145
- 2026-08-22runs1 run landed on cfg-0030 (ho003-guard-patched-q8_0) run-0554
- 2026-08-22runs10 runs landed on cfg-0009, cfg-0011 (llama-bench) cfg-0009 cfg-0011
- 2026-08-22runs2 runs landed on cfg-0144, cfg-0145 (guard-c32768-depth8000) cfg-0144 cfg-0145
- 2026-08-21claimHG-004 Phase-1 captured one real GLM-4.7-Flash Q4_K_M raw llama-server HTTP 200 response through the pinned installed Tau2/native-template task-1 boundary on aihydra; the normalized response was reasoning-only with empty assistant content, no tool_calls, non-empty reasoning_content, and finish_reason length, and Tau2 Phase-2 was not launched. clm-0099
- 2026-08-21claimHO-011 measured Ornith-1.5-35B Q4_K_M screening and full-26 tau2 capability on the stock ROCm0 3653e6d runtime at c32768 f16 KV. clm-0100
- 2026-08-21claimHU-004 reviewed the raw llama-server long-prompt guard for the genuine Qwen3.6-35B-A3B-MTP UD-Q4_K_M artifact on aihydra (runtime commit 7077abb, ROCm0/gfx1151, f16/f16 KV, context 204800). clm-0101
- 2026-08-21gategemma4-26b: rejected → rejected — Upstream #26750 draft-mtp defect addendum added as community report evidence, not a new HaloBench run and not a gate change (remains rejected per the HG-001 2026-08-18 verdict). clm-0102 con-0005
- 2026-08-21gateglm-47-flash: screened → screened — HG-004 Phase-1 boundary diagnostic captured one real raw llama-server HTTP 200 response through the pinned installed Tau2/native-template task-1 path. clm-0099 run-0544
- 2026-08-21gateornith-15-35b: acquired → screened — HO-011 reviewer-admitted screening at Q4_K_M, ROCm0 3653e6d, c32768, f16 KV, --parallel 1. clm-0100
- 2026-08-21gateornith-15-35b: listed — Ornith-1.5 is the direct successor to Ornith-1.0-35B on the same qwen35moe architecture, published as MIT-licensed GGUF by the vendor.
- 2026-08-21gateornith-15-35b: listed → acquired — Downloaded to aihydra ~/models/ornith-15-35b/ and sha256-verified against the sidecar: Q4_K_M (official repo) 21,713,462,848 bytes, sha256 ca6ea26329c88b78ffd90a85163be2e746c2fafd1024f56db47e499f117f9a7f.
- 2026-08-21gateqwen36-35b: benched → benched — HU-004 cleared the long-prompt guard for the same genuine MTP artifact, Qwen3.6-35B-A3B-MTP UD-Q4_K_M. clm-0101
- 2026-08-21runs2 runs landed on cfg-0140 (guard-c32768-depth8000) cfg-0140
- 2026-08-21runs3 runs landed on cfg-0140, cfg-0025 (tau2-bench-airline) cfg-0140 cfg-0025
- 2026-08-21runs2 runs landed on cfg-0140 (llama-bench) cfg-0140
- 2026-08-21runs2 runs landed on cfg-0141, cfg-0142 (guard-hu004-longprompt-8kto40k) cfg-0141 cfg-0142
- 2026-08-20claimNemotron-3-Super-120B-A12B now has a single record-backed Vulkan d204800 throughput cell under the existing cfg-0108 fingerprint: 130.8127 t/s prefill (run-0492) and 15.8832 t/s decode (run-0493), rc=0, contention=false, with the same stock llama.cpp 3653e6d6d Vulkan / UD-Q4_K_M / f16 KV / --load-mode none boundary as the prior d131072 Vulkan rows. clm-0091
- 2026-08-20claimOrnith-1.0-35B UD-Q4_K_XL has a bounded HO-005 r3 preflight record on the stock 3653e6d6d Vulkan/f16-KV boundary: the Q4 artifact identity was sidecar/local-sha matched, FIT/header/load at c32768 was healthy, the schema-visible guard passed 4/4 (run-0494), and the 5-task tau2 airline smoke passed 5/5 with mean_reward 1.000 (run-0495). clm-0092
- 2026-08-20claimOrnith-1.0-35B UD-Q4_K_XL has its own measured HO-005 full-26 tau2 airline capability record on the stock 3653e6d6d Vulkan/f16-KV boundary: after a fresh schema-visible guard passed 4/4 (run-0498), the full task-id 0..25 tau2 airline run completed rc=0 at 22/26 with mean_reward 0.8461538461538461 (run-0499), 185 tool-call messages, 0 empty-argument tool calls, 0 empty assistant turns, 0 infrastructure errors, and 0 max-step cuts. clm-0093
- 2026-08-20claimThe llama.cpp #25618 F16-V result is a prompt-sensitive exactness result, not a general proof that MTP preserves the target trajectory on Qwen3.8-27B. clm-0094
- 2026-08-20claimHO-005 measured a reviewer-admitted Ornith-1.0-35B UD-Q4_K_XL production-depth performance frontier for three exact guarded fingerprints: Vulkan f16/f16 KV, ROCm f16/f16 KV, and Vulkan q8_0/q8_0 KV. clm-0095
- 2026-08-20claimHO-001 v0.6.6 admitted fork-specific DeepSeek-V4-Flash sparse-path and performance evidence for Nathanw1014 portable Vulkan payload source 7b6c6133 build 10569, with matching f16/f16, q8_0/q8_0, and q4_0/q4_0 K/V cache guard plus throughput cells through d262144 and sparse-path proof for q8_0 and q4_0. clm-0096
- 2026-08-20claimHO-009 measured reviewer-admitted MTP activation/safety/performance evidence for the genuine Qwen3.6-35B-A3B-MTP UD-Q4_K_M artifact on aihydra. clm-0097
- 2026-08-20claimHO-004 measured a negative/speculation-path diagnostic for Qwen3.8-27B Q8_0 on the reviewed v0.6.5 portable Vulkan runtime: stock MTP n_max=3 and DFlash2 Q8_0 width 4 both showed log-backed draft activation, but plain, stock-MTP, and DFlash2 arms all failed Stage A on code/toolish sentinel-empty-content checks, so no guard, Stage B performance, throughput, or energy-efficiency cell is admitted. clm-0098
- 2026-08-20gate7 candidates → benched: deepseek-v4-flash, ornith-35b, ornith-35b, ornith-35b, qwen36-35b, qwen38-27b, qwen38-27b clm-0096 clm-0092 run-0494 run-0495 run-0496 run-0497 clm-0093 run-0498 run-0499 eng-0202 eng-0203 clm-0095 run-0500 run-0501 run-0502 run-0503 run-0504 run-0505 run-0506 run-0507 run-0508 run-0509 run-0510 run-0511 run-0512 run-0513 run-0514 run-0515 run-0516 run-0517 run-0518 run-0519 run-0520 run-0521 run-0522 clm-0097 clm-0094 con-0004 clm-0098 run-0541 run-0542 run-0543 eng-0222 eng-0223 eng-0224
- 2026-08-20contributionggml-org/llama.cpp report — open: Independent Strix Halo Vulkan follow-up on llama.cpp #25618. con-0004
- 2026-08-20runs2 runs landed on cfg-0129 (ho005-ornith-q4kxl-vulkan-preflight-r3) cfg-0129
- 2026-08-20runs28 runs landed on cfg-0129, cfg-0130, cfg-0131, cfg-0132, cfg-0134 (llama-bench) cfg-0129 cfg-0130 cfg-0131 cfg-0132 cfg-0134
- 2026-08-20runs6 runs landed on cfg-0129, cfg-0130, cfg-0131, cfg-0132, cfg-0134 (guard-c32768-depth8000) cfg-0129 cfg-0130 cfg-0131 cfg-0132 cfg-0134
- 2026-08-20runs1 run landed on cfg-0129 (tau2-bench-airline) run-0499
- 2026-08-20runs1 run landed on cfg-0133 (mtp-context-proof) run-0523
- 2026-08-20runs4 runs landed on cfg-0132, cfg-0133, cfg-0134, cfg-0135 (stageA-c32768) cfg-0132 cfg-0133 cfg-0134 cfg-0135
- 2026-08-20runs5 runs landed on cfg-0134 (llama-server-completion) cfg-0134
- 2026-08-20runs3 runs landed on cfg-0136, cfg-0137, cfg-0138 (stageA-c32768-varied-prompts) cfg-0136 cfg-0137 cfg-0138
- 2026-08-20runs1 run landed on cfg-0139 (hg004-phase1-boundary-diagnostic) run-0544
- 2026-08-19claimLing-3.0-flash's shipped NextN head is genuinely activatable through llama.cpp 7077abb's draft-MTP path on ROCm: each bounded n_max=1, 2 and 3 arm created an MTP draft context, emitted nonzero draft accounting, passed all 15 varied output-sanity generations and passed the house guard 4/4. clm-0086
- 2026-08-19claimNVIDIA Nemotron 3.5 Lightning 30B-A3B's native NextN head is genuinely active through llama.cpp 7077abb's draft-MTP path on ROCm. clm-0087
- 2026-08-19claimHG-002 Phase C's corrected matched plain control completed the five-task tau2 airline smoke at 4/5, mean 0.80, with six valid nonempty tool calls, zero empty turns, zero infrastructure errors and zero max-step cuts. clm-0088
- 2026-08-19claimUnder HG-003's corrected tokenizer-measured protocol, LFM2-24B-A2B Q4_K_M failed retrieval at the minimum tested depth of 2048 tokens with the needle planted early at token position 512. clm-0089
- 2026-08-19claimllama.cpp PR #27342's DFlash2 community evidence does not yet justify replacing Qwen3.8-27B's measured serving recommendation, but it does justify a future matched lab check once the implementation is stable: on one Strix Halo Vulkan report DFlash2 Q8_0 slightly beat MTP on prose within noise (18.96 vs 18.36 tok/s) and more clearly on code (22.68 vs 20.15 tok/s), while adjacent reports show hardware- and concurrency-specific hazards including Intel B70 multi-agent collapse, a V100 vision/M-RoPE failure in the draft context, V100 n_max=7 regression, RTX 3090 near-parity/modest gain, and Blackwell better scaling at higher parallelism. clm-0090
- 2026-08-19gate3 candidates → screened: lfm2-24b, nemotron35-lightning-30b, nemotron35-lightning-30b clm-0089 run-0473 eng-0197 clm-0087 clm-0088
- 2026-08-19gate3 candidates → benched: ling-30-flash, nemotron3-super, qwen38-27b run-0455 run-0457 run-0459 clm-0086 run-0492 run-0493 clm-0091 clm-0090
- 2026-08-19incidentHG-002 smoke parser rejected a valid task with no required tool call (data-integrity, resolved) inc-0006
- 2026-08-19incidentHG-003 retrieval runner stopped on an unbound backend variable before model work (data-integrity, resolved) inc-0007
- 2026-08-19incidentHG-003 nominal depth prompt exceeded its server context before placement (data-integrity, resolved) inc-0008
- 2026-08-19runs8 runs landed on cfg-0115, cfg-0119, cfg-0120, cfg-0121, cfg-0122, cfg-0123, cfg-0124, cfg-0125 (house-capability-guard) cfg-0115 cfg-0119 cfg-0120 cfg-0121 cfg-0122 cfg-0123 cfg-0124 cfg-0125
- 2026-08-19runs4 runs landed on cfg-0115, cfg-0119, cfg-0120, cfg-0121 (mtp-engagement-probe) cfg-0115 cfg-0119 cfg-0120 cfg-0121
- 2026-08-19runs4 runs landed on cfg-0122, cfg-0123, cfg-0124, cfg-0125 (mtp-engagement-probe) cfg-0122 cfg-0123 cfg-0124 cfg-0125
- 2026-08-19runs3 runs landed on cfg-0122, cfg-0123 (tau2-bench-airline) cfg-0122 cfg-0123
- 2026-08-19runs3 runs landed on cfg-0126 (HG-003-lfm2-retrieval-promotion) cfg-0126
- 2026-08-19runs18 runs landed on cfg-0127, cfg-0128 (pr25494-vulkan-q8kv-stock-r7) cfg-0127 cfg-0128
- 2026-08-19runs2 runs landed on cfg-0108 (llama-bench) cfg-0108
- 2026-08-18claimLing-3.0-flash's corrected post-PR-26608 Q4_K_M GGUF completed a valid full 26-task tau2 airline run on aihydra at 0.500 mean reward (13/26), with 170 tool-call messages, 0 empty assistant messages, 0 infrastructure errors and 0 max-step cuts. clm-0083
- 2026-08-18claimOn gfx1151 with the same DeepSeek-V4-Flash-0731 UD-IQ3_XXS artifact and f16 KV, the carried Strix Halo Vulkan fork at baf6360be passed the house 4-item guard and completed every bounded llama-bench cell through d262144 without a device loss. clm-0084
- 2026-08-18claimGemma 4 26B-A4B UD-Q4_K_M did not earn a reflex-tier promotion on HG-001: the corrected ROCm stock run passed coherence, native tool-call arguments and the parallel-1 isolation skip, but failed the house needle retrieval at the first declared depth, 2048 tokens, for a 3/4 guard result. clm-0085
- 2026-08-18gatedeepseek-v4-flash: benched → benched — Follow-up on the carried Strix Halo Vulkan fork at baf6360be changes the backend-stability finding without changing the production recommendation yet. clm-0084
- 2026-08-18gateling-30-flash: screened → benched — Full publishable bench completed on the corrected GGUF under the conservative ROCm/plain-decode path. run-0439 clm-0083 eng-0173
- 2026-08-18gategemma4-26b: screened → rejected — HG-001 bounded promotion retried the shallower retrieval-envelope check and stopped at the first declared depth: house guard 3/4, with coherence and native tool-call arguments passing but needle retrieval lost at 2048 tokens. clm-0085 run-0451
- 2026-08-18gateling-30-flash: screened → screened — Corrected post-PR-26608 GGUF acquired and screened: bloomer010/Ling-3.0-flash-GGUF `Ling-3.0-flash-Q4_K_M.gguf`, 78,285,287,904 bytes, sha256 bcce6e32799749e8e52a52c161127b989db52430e00fbebc77f4d61ef94e754d, staged first on ladon then copied to aihydra local NVMe and re-verified.
- 2026-08-18contributionggml-org/llama.cpp fork — carrying: Nathanw1014's Strix Halo Vulkan fork, staged first as the v0.6.4 portable payload at build baf6360be and then as the reviewed v0.6.6 portable payload at build 7b6c6133/source 7b6c61330edf370659f531932e0b91aca67ba055. con-0002
- 2026-08-18contributionggml-org/llama.cpp pr — open: Open llama.cpp DFlash2 support PR for Qwen3.8-27B-style drafters. con-0003
- 2026-08-18runs3 runs landed on cfg-0115, cfg-0117, cfg-0118 (house-capability-guard) cfg-0115 cfg-0117 cfg-0118
- 2026-08-18runs4 runs landed on cfg-0114 (llama-bench) cfg-0114
- 2026-08-18runs1 run landed on cfg-0115 (tau2-bench-airline) run-0439
- 2026-08-18runs10 runs landed on cfg-0116 (llama-bench) cfg-0116
- 2026-08-17claimqwen36-27b-mtp is the WORST full-bench candidate on this board by energy efficiency, by a wide margin, despite passing the cheap house guard twice (4/4, no cliff) and a directionally-consistent SMOKE screen (0.80 mean, 2026-08-13). clm-0076
- 2026-08-17claimDeep-Thought-Posttrain's KV cache costs a measured 40.00 KiB/token (f16 KV, ROCm and Vulkan identical — both default to the same type_k/type_v) — reproduced exactly from first principles against the GGUF's own attention metadata (2 x 32 layers x 5 KV heads x 64 head_dim x 2 bytes) and cross-checked against a real two-point GTT probe (c=2048 vs c=8192, /sys/class/drm/card0/device/mem_info_gtt_used readings before and after each load): (1,258,139,648 - 1,006,481,408) bytes / (8192-2048) tokens = 40,960 bytes/token exactly. clm-0077
- 2026-08-17claimDeep-Thought-Posttrain's full 26-task tau2 airline run (run-0372) recorded tool_call_messages = 0 across all 26 simulations — the same SMOKE validity gate (protocol.json comparison_rules, added 2026-08-14) that flagged npu-lfm2's vacuous 1.00 mean reward as invalid on this site's first NPU screen. clm-0078
- 2026-08-17claimQwen3-Coder-Next 80B (qwen3next hybrid, 512 experts/10 active + 1 shared, ~3B active) Q8_0's own throughput matrix has Vulkan leading decode at EVERY depth measured (44.13/37.0/26.4/21.73/19.15 t/s at d0/d32768/d131072/d204800/d262144 vs ROCm's 37.66/31.06/21.29/17.51/15.46) and prefill at the two shallowest cells (614.0/450.4 vs 476.6/346.8 t/s at d0/d32768) -- but the pattern INVERTS at the two deepest cells: ROCm overtakes prefill at d204800 (124.6 vs 98.75 t/s, Vulkan -20.7%) and by a wider margin at d262144, this candidate's model-max (104.56 vs 68.77 t/s, Vulkan -34.2%). clm-0079
- 2026-08-17claimQwen3-Coder-Next 80B's KV cache costs 24 KiB/token (f16 KV), computed from the GGUF's own architecture metadata rather than a live GTT probe: qwen3next. full_attention_interval=4 over 48 blocks means 12 full-attention layers (GQA, attention.head_count_kv=2, attention.key_length=attention.value_length=256), each costing 2 kv_heads x 256 head_dim x 2(K+V) x 2 bytes(f16) = 2,048 bytes/token, and the other 36 blocks are gated-DeltaNet-class SSM/linear- attention with NO growing KV cache at all. clm-0080
- 2026-08-17claimQwen3-Coder-Next 80B ships NO speculative-decode or MTP head in its official Q8_0 GGUF. clm-0081
- 2026-08-17claimQwen3-Coder-Next 80B Q8_0 completed the FULL standard 26-task tau2 AIRLINE set cleanly (rc=0, not a wall-bound cut) at 0.5385 mean reward (14/26), 192 tool-call messages, 390 assistant messages, 0 empty assistant turns, 0 infrastructure errors -- VALID under the SMOKE gate, served on Vulkan (this bench's own throughput-matrix pick, clm-0079) at plain decode (no MTP/ speculation exists in this GGUF, clm-0081). clm-0082
- 2026-08-17gatedeep-thought-posttrain: listed — Listed on operator direction as a deliberate negative-control candidate for the house benchmark suite, to be run with the same rigour as every capability candidate rather than as a joke entry — the humour, if any, is expected to come from the deadpan measurement, not from the write-up winking at the reader.
- 2026-08-17gateqwen3-coder-next-80b: listed — Direct friend-request, PRIORITY per operator.
- 2026-08-17gatedeep-thought-posttrain: listed → acquired — always42-universal.gguf (726 MB, F16, sha256 fa0a43b12b90ed9e8aa0e3c346b8f0e4e627d971f171a0b5b549fad31ce2d5df) downloaded to aihydra ~/models/deep-thought-posttrain/ and mirrored to the NAS gguf-library (/volume1/Models/gguf-library/deep-thought-posttrain/, remote size verified byte-identical to local).
- 2026-08-17gateqwen3-coder-next-80b: listed → acquired — Downloaded to aihydra ~/models/qwen3-coder-next-80b/ (official Qwen/Qwen3-Coder-Next-GGUF, Q8_0, 4 shards, 84,812,055,968 bytes total) and sha256-verified exact against HF's own X-Linked-ETag on all 4 shards (shard1 30b7554fc0c846a5dc3ecf585884c77471f73e3da698a8ba4fabd8e7868c6533, shard2 3f96379de5a5c4655cb378710ea571d5e9cc96f260120a44a6477198efcdc27d, shard3 5dd1ce07eaae95ee430331dc9c6f3120ff88e4211ad3a0cceeaa963f25328504, shard4 76730702c630bf76305139165cb85421858604030851dfd64fe96a5e67cda99d).
- 2026-08-17gate3 candidates → screened: deep-thought-posttrain, ling-30-flash, qwen3-coder-next-80b clm-0077
- 2026-08-17gate9 candidates → benched: deep-thought-posttrain, deepseek-v4-flash, laguna-s-21, nemotron3-super, ornith-35b, ornith-35b, qwen3-coder-next-80b, qwen36-27b-mtp, qwen38-27b run-0363 run-0372 eng-0146 run-0391 run-0392 eng-0156 run-0385 run-0386 run-0387 run-0388 run-0389 run-0390 eng-0153 eng-0154 eng-0155 run-0377 run-0378 run-0379 run-0380 run-0381 run-0382 run-0383 run-0384 eng-0152 clm-0079 clm-0080 clm-0081 clm-0082 run-0361 run-0362 run-0373 run-0374 run-0375 run-0376
- 2026-08-17contributionggml-org/llama.cpp pr — open: llama.cpp PR #27210 by stew675 "spec : add adaptive MTP draft depth (draft-mtp-adaptive)" — adds a new --spec-type draft-mtp-adaptive with a counting state machine (climb counter + weighted drop-pressure accumulator) so draft depth adjusts per segment instead of staying at a fixed n_max. con-0011
- 2026-08-17runs1 run landed on cfg-0099 (qwen36-27b-mtp-fullbench) run-0359
- 2026-08-17runs3 runs landed on cfg-0099, cfg-0110, cfg-0111 (tau2-bench-airline) cfg-0099 cfg-0110 cfg-0111
- 2026-08-17runs66 runs landed on cfg-0100, cfg-0101, cfg-0102, cfg-0103, cfg-0104, cfg-0105, cfg-0106, cfg-0107, cfg-0108, cfg-0109, cfg-0112, cfg-0113 (llama-bench) cfg-0100 cfg-0101 cfg-0102 cfg-0103 cfg-0104 cfg-0105 cfg-0106 cfg-0107 cfg-0108 cfg-0109 cfg-0112 cfg-0113
- 2026-08-17runs2 runs landed on cfg-0101 (deep-thought-posttrain-fullbench) cfg-0101
- 2026-08-17runs2 runs landed on cfg-0110, cfg-0111 (house-capability-guard) cfg-0110 cfg-0111
- 2026-08-16claimDeepSeek-V4-Flash-0731 UD-IQ3_XXS completed the FULL standard 26-task tau2 airline set cleanly (rc=0, not a wall-bound cut) at 0.846 mean reward (22/26), 272 tool-call messages, 0 empty assistant turns, 0 infrastructure errors -- VALID under the SMOKE gate. clm-0061
- 2026-08-16claimNemotron-3-Super-120B-A12B's KV cache costs 8.00 KiB per token at UD-Q4_K_M on llama.cpp — measured byte-exact from a two-point GTT delta, ROCm, f16 KV, --load-mode none, --parallel 1: 224 MiB between a c=4096 and a c=32768 load probe (78,754 vs 78,978 MiB gtt_used), over 28,672 additional tokens (229,376 KiB / 28,672 tok = 8.00 KiB/token exactly). clm-0062
- 2026-08-16claimOn gfx1151 at UD-Q4_K_M / f16 KV, Nemotron-3-Super-120B-A12B is the FIRST model in this programme where Vulkan decode measurably beats ROCm: +8.97% at d0 (18.2043 vs 16.7065 t/s) and +8.19% at d32768 (17.7386 vs 16.3962 t/s), a stable signature by depth. clm-0063
- 2026-08-16claimNemotron-3-Super-120B-A12B UD-Q4_K_M completed the FULL standard 26-task tau2 airline set cleanly (rc=0, not a wall-bound cut) at 0.7692 mean reward (20/26), 263 tool-call messages, 429 assistant messages, 0 empty assistant turns, 0 infrastructure errors -- VALID under the SMOKE gate. clm-0064
- 2026-08-16claimOn gfx1151, Laguna-S-2.1 Q4_K_M is a sixth model measured under the per-model, per-phase backend rule (clm-0050), and unlike nemotron3-super and deepseek-v4-flash (both Vulkan-favoured on this box), ROCm wins outright here: ROCm swept the full throughput matrix clean at every depth tried (d0, d32768, d131072 -- 300.7/23.68 t/s down to 116.43/7.19 t/s pp/tg), while Vulkan completed only d0 and d32768 (256.9/13.29 and 51.6/12.27 t/s pp/tg, ROCm ahead on every figure at every depth both backends completed) and lost the GPU device on both allowed attempts at d131072 (rc=134, vk::DeviceLostError; kernel evidence: amdgpu ring timeout, ring reset FAILED, a MODE2 GPU reset, "device wedged, but recovered through reset" -- the same device-loss class as clm-0054/clm-0059's precedent, now a fourth confirmed occurrence on this silicon). clm-0065
- 2026-08-16claimLaguna-S-2.1's KV cache costs approximately 48.0 KiB per token on llama.cpp at f16 -- reportedly measured from a two-point GTT load probe (c=4096 -> 92,131.38 MiB gtt_used, c=131,072 -> 98,083.38 MiB gtt_used; delta 5,952.00 MiB over 126,976 tokens). clm-0066
- 2026-08-16claimLaguna-S-2.1 Q4_K_M completed the FULL standard 26-task tau2 airline set cleanly (rc=0, not a wall-bound cut) at 0.6923 mean reward (18/26), 187 tool-call messages, 0 empty assistant turns, 0 infrastructure errors -- VALID under the SMOKE gate. clm-0067
- 2026-08-16claimOrnith-1.0-35B (ornith-ai/deepreinforce-ai, arch qwen35moe, A3B-class MoE, MIT) passed the house standard screen on release-adjacent acquisition, stock 3653e6d ROCm, Q8_0, -np 1, c=32768: FIT loaded in 12s at 35,607 MiB GTT (comfortable against the 122,880 MiB boot window), GUARD 4/4 (coherence, tool call, needle at 8000 tokens, isolation skipped under -np 1), SMOKE 5/5 with mean_reward 1.000 on the 5-task tau2 airline smoke (21 tool-call messages, 0 empty assistant turns, 0 infra errors -- VALID per the smoke gate; run-0330, run-0331). clm-0068
- 2026-08-16claimgithub.com/julianmb/q38rocm (r/StrixHalo 1vpiwz0) is a genuine, substantial llama.cpp fork -- charlie12345/ROCmFPX, based on official llama.cpp b9438 (commit 22cadc194), pinned at e87d53e for this artifact -- carrying real custom ROCmFP4/ROCmFP4_FAST GGUF block-quant tensor formats and a real --spec-mtp-strict-qwen exact-verification mode for qwen35/qwen35moe MTP (implemented in tools/server/server-context.cpp + common/arg.cpp, gated on -np 1 and sufficient recurrent-rollback depth, dynamically caps draft length to stay inside one 256-token dense-attention KV block to avoid a ROCm floating-point rounding divergence they found at block boundaries). clm-0069
- 2026-08-16claimOn gfx1151, Ornith-1.0-35B Q8_0 is measured under the per-model, per-phase backend rule (clm-0050), and Vulkan wins outright here -- unlike laguna-s-21 (ROCm-favoured, clm-0065) but matching nemotron3-super and deepseek-v4-flash's Vulkan-favoured pattern: Vulkan leads decode at every measured depth (55.70 vs 47.66 t/s at d0, +17%; 46.22 vs 39.75 at d32768, +16%; 32.37 vs 27.48 at d131072, +18%) and prefill at d0/ d32768 (1063.7 vs 865.3, +23%; 679.2 vs 525.2, +29%), with ROCm only marginally ahead on prefill at the deepest cell (244.8 vs 238.5, +2.6%). clm-0070
- 2026-08-16claimOrnith-1.0-35B's KV cache costs 20.0 KiB per token on llama.cpp at f16 (default KV, ROCm), measured from a two-point GTT load probe (c=4096 -> 36,750,315,520 bytes gtt_used, c=131,072 -> 39,350,784,000 bytes gtt_used; delta 2,600,468,480 bytes over 126,976 tokens). clm-0071
- 2026-08-16claimOrnith-1.0-35B Q8_0 completed the FULL standard 26-task tau2 airline set cleanly (rc=0, not a wall-bound cut) at 0.8846 mean reward (23/26), 202 tool-call messages, 0 empty assistant turns, 0 infrastructure errors -- VALID under the SMOKE gate. clm-0072
- 2026-08-16claimThe staged qwen36-27b-mtp GGUF (unsloth/Qwen3.6-27B-GGUF, Q4_K_M) is NOT the uniform-dense-attention, MTP-capable model this candidate's own record assumed. clm-0073
- 2026-08-16claimOn qwen36-27b-mtp Q4_K_M, ROCm wins prefill decisively at every tested depth over Vulkan (357.8 vs 302.7 t/s at d0, +18%; 210.5 vs 95.7 t/s at d32768, +120%) and is the ONLY backend that completed the deepest matrix cell (d131072) without a device-loss event — Vulkan device-lost on BOTH allowed attempts at that depth (vk::DeviceLostError, GPU wedged, recovered via kernel amdgpu ring reset each time; ROCm completed cleanly on its first attempt at 95.4 pp / 8.44 tg t/s). clm-0074
- 2026-08-16claimqwen36-27b-mtp's KV cache costs 64.00 KiB per token on llama.cpp at f16 (explicit KV, ROCm), measured from a two-point GTT load probe (c=4096 -> 16,821,088,256 bytes gtt_used, c=131,072 -> 25,142,587,392 bytes gtt_used; delta 8,321,499,136 bytes over 126,976 tokens = exactly 64.00 KiB/token). clm-0075
- 2026-08-16gate6 candidates → benched: deepseek-v4-flash, laguna-s-21, nemotron3-super, ornith-35b, qwen36-27b-mtp, qwen38-27b clm-0059 clm-0060 clm-0061 clm-0065 clm-0066 clm-0067 clm-0062 clm-0063 clm-0064 clm-0070 clm-0071 clm-0072 run-0332 run-0333 run-0334 run-0335 run-0336 run-0337 run-0338 run-0339 run-0340 run-0341 run-0342 run-0343 run-0344 run-0345 clm-0069
- 2026-08-16gateornith-35b: listed — Listed on the architecture-match rationale above (qwen35moe, same family as production 122B and qwen38-27b) plus the model's own MIT-licensed, vendor-published agentic-coding benchmark table (Terminal-Bench 2.1 64.2, SWE-bench Verified 75.6, both ahead of Qwen3.6-35B on the vendor's own numbers) -- vendor-claim provenance, not evidence, per house policy, but a real published measurement all the same.
- 2026-08-16gateornith-35b: listed → acquired — Downloaded to aihydra ~/models/ornith-35b/ and sha256-verified exact against the HF LFS oid on all three files: Q8_0 (primary screening quant, OFFICIAL repo) 36,903,138,880 bytes, sha256 cbc992bca07901c1a51f33e65e6fc5d687de179c852a772dfd15e4c3261dbf5c; UD-Q4_K_XL (comparability arm, unsloth -- matches the quant class protocol.json quants.expected already carries) 22,324,804,000 bytes, sha256 67081ae4a1a291bd6c72834094ea056332cb3cb5fa15e88536ec7f233a475b71; mmproj-F16 (vision file, unsloth -- the official repo has none) 899,283,680 bytes, sha256 217ee3ae58ef7b1f743341a0037c9da50aa6ae06e051ccb49c517132b0ac2bf6.
- 2026-08-16gateornith-35b: acquired → screened — House standard screen, stock 3653e6d ROCm, Q8_0, -np 1, c=32768. clm-0068 run-0330 run-0331
- 2026-08-16runs2 runs landed on cfg-0085, cfg-0086 (eos-cliff-isolation) cfg-0085 cfg-0086
- 2026-08-16runs3 runs landed on cfg-0066, cfg-0089, cfg-0092 (house-capability-guard) cfg-0066 cfg-0089 cfg-0092
- 2026-08-16runs42 runs landed on cfg-0087, cfg-0088, cfg-0090, cfg-0091, cfg-0094, cfg-0096, cfg-0097, cfg-0098 (llama-bench) cfg-0087 cfg-0088 cfg-0090 cfg-0091 cfg-0094 cfg-0096 cfg-0097 cfg-0098
- 2026-08-16runs4 runs landed on cfg-0089, cfg-0092, cfg-0095, cfg-0099 (tau2-bench-airline) cfg-0089 cfg-0092 cfg-0095 cfg-0099
- 2026-08-16runs2 runs landed on cfg-0093 (ornith-35b-screen) cfg-0093
- 2026-08-16runs1 run landed on cfg-0095 (ornith-35b-fullbench) run-0338
- 2026-08-16runs1 run landed on cfg-0099 (qwen36-27b-mtp-fullbench) run-0357
- 2026-08-15claimOn gfx1151 at f16 KV, Qwen3.8-27B (dense) is a fourth model measured under the per-model, per-phase backend rule (clm-0050), and it lands emphatically on the prefill-favours-ROCm side: DECODE is backend-independent on this model (every matched cell within 5%), but Vulkan PREFILL collapses with depth — 0.55x ROCm at d32768 on Q8_0 (87.82 vs 159.98 t/s) and 0.48x on UD-Q4_K_XL (99.80 vs 207.61) — and at d131072 stock Vulkan cannot complete the cell at all: 2 of 2 reps in BOTH quants aborted with vk::DeviceLostError, the kernel logging an amdgpu ring timeout and recovering the device by ring reset, while ROCm completed every cell it was offered (83.44 pp / 6.11 tg Q8_0, 95.59 pp / 8.14 tg UD-Q4_K_XL at d131072). clm-0054
- 2026-08-15claimdraft-mtp speculation on Qwen3.8-27B (stock 3653e6d, gfx1151) peaks at spec-draft-n-max=3 over CLEAN cells — Vulkan Q8_0 17.75 t/s (2.26x its 7.86 no-speculation floor, acceptance 0.626), Vulkan UD-Q4_K_XL 26.68 t/s (2.23x, 0.623), ROCm Q8_0 18.25 t/s (2.33x, 0.6235) — and at n_max >= 4 the feature is BROKEN on this model: after accumulated generation volume in a live session (sequential VARIED prompts at full length; not fresh servers, not one repeated prompt, not short generations), generations start terminating at 1 token with <|im_end|> (id 248046). clm-0055
- 2026-08-15claimQwen3.8-27B's KV cache on llama.cpp costs exactly 64.00 KiB per token — the hybrid-attention allocation working as designed, measured byte-exact on this box: only 16 of the 64 layers carry a conventional KV cache (the 3:1 linear-to-full-attention layout), and 16 layers x 4 KV heads x 256 head_dim x 2 (K+V) x 2 bytes (f16) = 65,536 bytes/token. clm-0056
- 2026-08-15claimQwen3.8-27B met its registered agentic prediction: on the standing 5-task tau2 airline smoke subset it scored 1.000 (5/5, 21 tool-call messages, 0 empty assistant turns, VALID under the smoke gate) against the >= 0.80 bar recorded in the candidate record on 2026-08-14, BEFORE any measurement — and against the incumbent qwen36-27b-mtp's 0.80 on the identical tasks under the identical pinned-simulator protocol, where the incumbent failed task 2 and this model passed it. clm-0057
- 2026-08-15claimclm-0035's retracted quality-drop story does not reappear at n=14 on the PATCHED (ce7689f) build: task-matched against an f16 control, quantised KV on Qwen3.6-35B-A3B-UD-Q4_K_XL scores q8_0 0.857 against f16 0.786 — q8_0 AHEAD, not behind. 13 of the 14 matched tasks scored identically in both arms; the single exception is the one f16 failed and q8_0 passed. clm-0058
- 2026-08-15supersededOn gfx1151 at f16 KV, DeepSeek-V4-Flash-0731 UD-IQ3_XXS is a fifth model measured under the per-model, per-phase backend rule (clm-0050), and it lands on the same side as Qwen3.8-27B (clm-0054): ROCm swept the full throughput matrix clean (d0 through d262144) while Vulkan lost the GPU device on both allowed attempts at every depth >=32768 (rc=134/SIGABRT, vk::DeviceLostError; kernel evidence: amdgpu ring timeout, ring reset, "device wedged, but recovered through reset" — 12 reset cycles total across the four failed cells). clm-0059 clm-0084
- 2026-08-15claimDeepSeek-V4-Flash-0731's KV cache costs approximately 7.13 KiB per token on llama.cpp — about a ninth of Qwen3.8-27B's 64.00 KiB/token hybrid-attention cache (clm-0056) — measured from a two-point GTT delta at --parallel 1, f16 KV: 884 MiB between a c=4096 and a c=131072 load probe (99,464 vs 100,348 MiB gtt_used), over 126,976 additional tokens. clm-0060
- 2026-08-15gate3 candidates → screened: deepseek-v4-flash, muse-glimmer-30b, qwen38-27b clm-0055 clm-0056 clm-0057 run-0267
- 2026-08-15gatellada-22-flash: listed — Listed on release-week signal (r/AIDeveloperNews thread 1vk21p9).
- 2026-08-15gatenemotron3-super: benched → benched — Q4_K_M re-run DONE - the fair-quant screen this record has flagged outstanding since 2026-08-09. run-0274
- 2026-08-15gateqwen38-27b: screened → benched — DAY-ONE SCREEN COMPLETE, all four phases, overnight on release day. clm-0054 clm-0055 clm-0056 clm-0057 run-0271 run-0267 run-0268
- 2026-08-15runs1 run landed on cfg-0066 (eos-cliffguard) run-0267
- 2026-08-15runs1 run landed on cfg-0067 (mtp-probe) run-0268
- 2026-08-15runs2 runs landed on cfg-0066, cfg-0084 (tau2-bench-airline) cfg-0066 cfg-0084
- 2026-08-15runs2 runs landed on cfg-0070, cfg-0071 (kv-quality-patched) cfg-0070 cfg-0071
- 2026-08-15runs1 run landed on cfg-0072 (nemotron3-super-q4km-screen) run-0274
- 2026-08-15runs2 runs landed on cfg-0073, cfg-0074 (eos-cliff-q8kv) cfg-0073 cfg-0074
- 2026-08-15runs10 runs landed on cfg-0075, cfg-0076, cfg-0077, cfg-0078, cfg-0081, cfg-0079, cfg-0080 (eos-cliff-isolation) cfg-0075 cfg-0076 cfg-0077 cfg-0078 cfg-0081 cfg-0079 cfg-0080
- 2026-08-15runs16 runs landed on cfg-0082, cfg-0083 (llama-bench) cfg-0082 cfg-0083
- 2026-08-15runs1 run landed on cfg-0084 (house-capability-guard) run-0393
- 2026-08-14claimAMENDED 2026-08-14 — the rule below is PER-MODEL AND PER-PHASE, not fleet-wide; see the correction history and the amendment note. clm-0050
- 2026-08-14claimOn gfx1151, stock ROCm 3653e6d's own BF16 KV decode is SLOWER than its own F16 KV decode: 29.26 t/s vs 42.35 t/s at 32,768 depth (Qwen3.6-35B-A3B, fa on) — a 31% penalty for switching to the KV type the community recommends for its quality, with prefill roughly unchanged (587.4 vs 595.0). stew675's rdna-boosts branch (commit ed89854) fixes it with a native BF16 flash-attention tile kernel: 47.31 t/s BF16 decode, +61.7% over stock's own BF16 and stock's best config on this chip. clm-0052
- 2026-08-14claimA benchmarking-harness defect, not a hardware limitation, produced this lab's earlier "Nemotron-3-Super cannot allocate on ROCm at any depth" verdict (clm-0050's first amendment, corrected 2026-08-14). clm-0053
- 2026-08-14gate4 candidates → screened: npu-embeddinggemma, npu-lfm2, npu-qwen3-4b-thinking, npu-whisper
- 2026-08-14gateqwen38-27b: listed → listed — RELEASED and CONFIRMED runnable.
- 2026-08-14runs8 runs landed on cfg-0050, cfg-0051 (llama-bench) cfg-0050 cfg-0051
- 2026-08-14runs44 runs landed on cfg-0052, cfg-0053, cfg-0056, cfg-0059, cfg-0060, cfg-0032, cfg-0062, cfg-0063, cfg-0064, cfg-0065 (llama-bench) cfg-0052 cfg-0053 cfg-0056 cfg-0059 cfg-0060 cfg-0032 cfg-0062 cfg-0063 cfg-0064 cfg-0065
- 2026-08-14runs8 runs landed on cfg-0054, cfg-0055 (llama-bench) cfg-0054 cfg-0055
- 2026-08-14runs2 runs landed on cfg-0057, cfg-0058 (perf-matrix-gtt120) cfg-0057 cfg-0058
- 2026-08-14runs1 run landed on cfg-0061 (memgate-test) run-0242
- 2026-08-14runs2 runs landed on cfg-0068, cfg-0069 (mtp-probe) cfg-0068 cfg-0069
- 2026-08-13claimQuantised KV on gfx1151 splits three ways by build. clm-0051
- 2026-08-13gate7 candidates → screened: gemma4-12b, glm-45-air, ling-30-flash, llama4-scout, maple-preview, nemotron35-lightning-30b, qwen36-27b-mtp
- 2026-08-13gatenemotron35-lightning-30b: listed → acquired — Official ggml-org Q4_K_M downloaded to aihydra (25,430,738,944 bytes, exactly the listed size) and sha256-verified against the HF LFS oid (6110e2e2e6cd324e6ee69ddced5a6b34fad6c94ca9827222a1e420fb92e3c90b).
- 2026-08-13runs44 runs landed on cfg-0038, cfg-0039, cfg-0040, cfg-0041, cfg-0042, cfg-0043, cfg-0044, cfg-0045, cfg-0046, cfg-0047, cfg-0048, cfg-0049 (llama-bench) cfg-0038 cfg-0039 cfg-0040 cfg-0041 cfg-0042 cfg-0043 cfg-0044 cfg-0045 cfg-0046 cfg-0047 cfg-0048 cfg-0049
- 2026-08-12gatedeepseek-v4-flash: listed → listed — Identity now concrete via the r/LocalLLaMA Strix Halo guide thread: DeepSeek V4 Flash 0731, deepseek4 arch, 256 experts/6 active + 1 shared, MIT licence, 1M native context. clm-0031
- 2026-08-12gatenemotron35-lightning-30b: listed — Surfaced via an operator-shared r/AIDeveloperNews link ("NVIDIA has launched Nemotron 3.5 Lightning") plus a follow-up r/StrixHalo post on a community ROCmFP4 requant with hardware-matched Strix Halo numbers.
- 2026-08-12runs32 runs landed on cfg-0032, cfg-0033, cfg-0034, cfg-0035 (llama-bench) cfg-0032 cfg-0033 cfg-0034 cfg-0035
- 2026-08-12runs16 runs landed on cfg-0036, cfg-0037 (llama-bench) cfg-0036 cfg-0037
- 2026-08-11claimA second independent Strix Halo source (llama.cpp PR #26856 + its Reddit write-up) reports Vulkan ahead of ROCm on decode at depth by ~10.5% on a clean same-binary comparison — same direction as clm-0031's +55% but a fifth the magnitude, confirming that figure was mostly build-gap and private patches. clm-0044
- 2026-08-11claimThe KV dequant patch removes most of quantised KV's agentic cost, not just its speed cost: on identical seeded tasks, patched q8_0 takes +9.3% more turns than patched f16 (234 vs 214 over 11 paired tasks) where the stock build cost +39% (clm-0038). clm-0045
- 2026-08-11claimThe community BF16 flash-attention predictions (clm-0044) reproduce on this hardware under single-binary methodology: at 32k depth, stock Vulkan leads ROCm +17.4% prefill / +17.1% decode at f16 KV, and the PR-26856 patch inverts prefill to ROCm +46.9% with bf16 KV. clm-0046
- 2026-08-11claimNemotron's quant confound resolves cleanly: on 12 identical seeded tasks, UD-Q4_K_M and UD-IQ4_XS produce IDENTICAL reward on every task (0.583 both), but IQ4_XS takes +11% more total turns (630 vs 567). clm-0047
- 2026-08-11claimSimulator identity changes what tau2 measures. clm-0048
- 2026-08-11claimThe 122B has a reproducible rule-precedence bug in policy application: on tau2 airline task 9 it cancels a partially-flown reservation in 7 of 8 trials, every failure with the identical signature — cancel_reservation called without ever checking flight status. clm-0049
- 2026-08-11gatemuse-glimmer-30b: blocked → screened — Screened on min-62bf73d (its minimum build, anchor-calibrated: +2.8%/+0.9% vs fleet baseline, identical tau2 capability).
- 2026-08-11runs9 runs landed on cfg-0029, cfg-0030, cfg-0025, cfg-0031, cfg-0028, cfg-0027 (tau2-bench-airline) cfg-0029 cfg-0030 cfg-0025 cfg-0031 cfg-0028 cfg-0027
- 2026-08-10claimThe 122B's real τ²-bench airline score is 0.545 +/-0.208, not the 1.00 reported by 5-task runs, which sampled only the easiest tasks in the domain. clm-0037
- 2026-08-10claimPaired on identical tasks, q8_0 KV costs TURN EFFICIENCY: 228 turns against f16's 164 over the same 9 tasks, +39%, taking more turns on 6 of 9 and fewer on 1. clm-0038
- 2026-08-10claimOn reward the 122B and Nemotron are indistinguishable, but reward is the wrong headline: on tasks both get RIGHT, the 122B needs 19 turns and 2.0 minutes against Nemotron's 26 and 3.7 — 37% fewer loops and 85% less wall-clock to the same correct answer. clm-0039
- 2026-08-10supersededby clm-0042's per-task measurement, which found the true energy cost roughly 10x lower — this run's figures are whole-arm totals padded by model loading and non-scoring tasks, not the model's actual energy per answer, and must not be used to rank models. clm-0040 clm-0042
- 2026-08-10claimThe KV dequant patch cuts energy 42% at 200k context — 146.1 Wh unpatched against 85.3 Wh patched for the same throughput benchmark — and patched q8_0 (85.3 Wh) beats f16 (89.6 Wh). clm-0041
- 2026-08-10claimPer-task energy windows, matched to a common task set across arms, put a correct τ² answer at 6.81 Wh on the 122B with f16 KV, 9.48 Wh with q8_0, and 12.75 Wh on Nemotron. clm-0042
- 2026-08-10claimEvery τ² arm was run with --user-llm set to the same model as --agent-llm, so cross-model comparisons changed the agent AND the user simulator together — exactly what the runbook forbids ("hold both --user-llm and the judge fixed across comparisons, or results re-baseline silently"). clm-0043
- 2026-08-10gate7 candidates → acquired: cascade2-30b, gemma4-26b, glm-47-flash, laguna-s-21, lfm2-24b, muse-glimmer-30b, qwen36-27b-mtp
- 2026-08-10gate5 candidates → screened: cascade2-30b, gemma4-26b, glm-47-flash, laguna-s-21, lfm2-24b
- 2026-08-10gate16 candidates → listed: deepseek-v4-flash, gemma4-12b, glm-45-air, ling-30-flash, llama4-scout, maple-preview, muse-glimmer-30b, npu-embeddinggemma, npu-embeddinggemma, npu-lfm2, npu-qwen3-4b-thinking, npu-qwen3-4b-thinking, npu-whisper, npu-whisper, qwen38-27b, qwen38-27b clm-0031 clm-0034 clm-0015 clm-0018
- 2026-08-10gate4 candidates → blocked: laguna-s-21, laguna-s-21, muse-glimmer-30b, muse-glimmer-30b
- 2026-08-10runs2 runs landed on cfg-0026, cfg-0027 (tau2-bench-airline) cfg-0026 cfg-0027
- 2026-08-09supersededthe Pass^1 = 1.000 reported by this run came from a 3-task subsample biased toward the domain's easiest tasks — the other two of the original five never terminated and were excluded as infrastructure errors. clm-0030 clm-0037
- 2026-08-09claimA community DeepSeek-V4-Flash report independently confirms our mmap/GTT double-residency finding, demonstrates a 120 GiB GTT ceiling in production use, and — most consequentially — reports Vulkan BEATING ROCm on 3 of 4 cells including 55% faster decode at depth. clm-0031
- 2026-08-09claimThe tau2 "runaway" tasks are a failure to escalate — and clm-0033 later established the failure is CAUSED BY THINKING, which this claim wrongly ruled out. clm-0032
- 2026-08-09claimOn tau2-bench airline, thinking is a NET NEGATIVE for this model: thinking OFF solves 5/5 tasks at reward 1.0 in ~10 minutes, while thinking ON solves 3/5 and deadlocks indefinitely on the other two (4h15m and 2h10m in unbounded runs). clm-0033
- 2026-08-09claimdomdoss/Warden (unrelated project, coincidental name) implements the multi-model architecture we have been designing toward — a small local orchestrator routing to named specialists with per-agent model selection. clm-0034
- 2026-08-09retractionthe per-model reward rankings and quantised-KV cost reported by this run do not hold — they came from 5-task tau2 arms whose ~0.40 run-to-run noise and 40-step cap bias were only characterised afterward (clm-0036), so the reward numbers below are not usable. clm-0035 clm-0037 clm-0039
- 2026-08-09claimτ²-bench at 5 tasks cannot resolve the differences drawn from it in this project's capability matrix. clm-0036
- 2026-08-09gatedeepseek-v4-flash: listed — Community report worth testing directly. clm-0031
- 2026-08-09gategemma4-12b: listed — Evidence from a comparable project that a 12B suffices for the orchestrator role. clm-0034
- 2026-08-09gatenemotron3-super: screened → benched — Ran the tau2 arms. clm-0035 clm-0036
- 2026-08-09gateqwen36-35b: screened → benched — Throughput, FA, ngram speculation and tau2 arms run. clm-0026
- 2026-08-09runs1 run landed on cfg-0006 (tau2-bench-airline) run-0098
- 2026-08-09runs1 run landed on cfg-0006 (tau2-bench-airline-nothink) run-0099
- 2026-08-09runs1 run landed on cfg-0025 (tau2-bench-airline) run-0100
- 2026-08-08claimOn aihydra (gfx1151, ROCm 7.1, llama.cpp 3653e6d), running llama-server with `--parallel 4` destroys long-context needle retrieval — 0/8 across controlled trials — while `--parallel 1` on the same build, model and prompt succeeds 8/8. clm-0019
- 2026-08-08claimOn ROCm/gfx1151 with a stock llama.cpp, flash attention is unambiguously BETTER at depth — at 32k it is worth 1.22x prefill and 1.84x decode on the 122B MoE — which is the opposite of the Vulkan cliff reported in clm-0017. clm-0020
- 2026-08-08claimThe dense-model flash-attention prefill cliff reported in clm-0017 does NOT exist on ROCm. clm-0021
- 2026-08-08claimThe community KV-dequantisation fix is real, large, and scales monotonically with depth: one cherry-picked commit recovers +18.3% at 32k, +55.6% at 131k and **+70.3% at 204,800 — production's own context** — while leaving f16 unchanged at every depth. clm-0022
- 2026-08-08claimQwen3.6-35B-A3B is 2.3x the 122B's decode on identical hardware (51.01 vs 21.90 tok/s at empty context) and holds 42.45 at 32k, with prefill above 1000 tok/s. clm-0023
- 2026-08-08claimThe tool-grammar ceiling does not reproduce on llama.cpp 3653e6d. clm-0024
- 2026-08-08claimFour models measured on identical hardware give decode from 17.51 to 55.45 tok/s, and the bandwidth model predicts the ORDER but not the magnitude — realised efficiency ranges from 34% to 62% of the theoretical ceiling. clm-0025
- 2026-08-08claimSpeculation is close to worthless on Qwen3.6-35B-A3B with varied prompts — ngram-mod gives 1.11x with a 29% coefficient of variation, ngram-cache gives nothing, and MTP is unavailable because the model carries no NextN layers. clm-0026
- 2026-08-08claimPrefix cache reuse is worth 9.8x on this box — an 8,000-token prefix costs 25.99 s cold and 2.65 s warm — and it is strictly PREFIX-ANCHORED: prepending three characters to an otherwise identical prompt returns it to full cold cost (26.26 s), zero reuse despite 99.9% identical content. clm-0027
- 2026-08-08claimAcross four models on gfx1151, flash attention is worth 2.5x to 4.8x DECODE at 131k and its absence is catastrophic — a 35B loses 88% of its decode speed from empty context to 131k without it, against 44% with it. clm-0028
- 2026-08-08claimDIAGNOSED: `llama-perplexity` produces garbage on this build — PPL 532 on deliberately repetitive English that should score 2-5, 3,654 on wikitext for a model that generates coherently at 51 tok/s and passes tool-calling guards. clm-0029
- 2026-08-08gategpt-oss-120b: screened → benched — Throughput rows recorded. clm-0025
- 2026-08-08gateqwen35-122b: screened → benched — Fully characterised on performance: depth to 204.8k, KV quant both ways, FA both ways, speculation curve, cache trace, grammar ceiling. clm-0022 clm-0027
- 2026-08-08runs1 run landed on cfg-0006 (capability-guard) run-0005
- 2026-08-08runs2 runs landed on cfg-0006, cfg-0007 (spec-ab) cfg-0006 cfg-0007
- 2026-08-08runs84 runs landed on cfg-0008, cfg-0009, cfg-0010, cfg-0012, cfg-0015, cfg-0016, cfg-0017, cfg-0018, cfg-0019, cfg-0020, cfg-0021, cfg-0022, cfg-0023, cfg-0024 (llama-bench) cfg-0008 cfg-0009 cfg-0010 cfg-0012 cfg-0015 cfg-0016 cfg-0017 cfg-0018 cfg-0019 cfg-0020 cfg-0021 cfg-0022 cfg-0023 cfg-0024
- 2026-08-08runs6 runs landed on cfg-0011, cfg-0013, cfg-0014 (llama-bench) cfg-0011 cfg-0013 cfg-0014
- 2026-08-06claimLing-3.0-flash (Ant Group, 26 Jul 2026) is a 124B/5.1B-active hybrid MoE that is a near-ideal controlled comparison for our Qwen3.5-122B-A10B — same footprint, half the active parameters — but it cannot be benchmarked on a stock llama.cpp, and its MTP head ships INACTIVE, which would silently rig any decode comparison in Qwen's favour. clm-0015
- 2026-08-06claimMeasured on Strix Halo: Qwen3.6-35B-A3B under ROCmFP4 + HIP + ngram-mod at parallel 4 sustains 121 tok/s across 500 varied IFEval prompts against a 64.8 tok/s no-speculation floor — a 1.87x production speedup with IFEval-strict at 78.6% — while prefill collapses from 1,211 tok/s cold to 136 tok/s at ~243k depth. clm-0016
- 2026-08-06claimOn Strix Halo Vulkan, a single unmerged patch — contiguizing strided f16 KV data before the flash-attention prefill — removes a dense-model prefill collapse worth 2.5x at 32k and 6.6x at 65k, changes decode not at all, and collapses run-to-run scatter from 5-10% to 0.3%. clm-0017
- 2026-08-06claimDeepGrove Maple-Preview (20.2B-A1.49B ternary, 5.31 GB, MIT) is a credible second-lane candidate for the Mac mini, but not a Warden candidate — its own model card concedes underperformance on agentic benchmarks. clm-0018
- 2026-08-05claimIdentity-grounded persistent sessions with a shared memory and peer-to-peer messaging can self-organise into useful working groups without an orchestrator. clm-0008
- 2026-08-05claimA ROCmFP4 iMatrix quant of our exact model (Qwen3.5-122B-A10B) reports 60.70 GiB and 28.505 tok/s decode with MTP OFF — against our measured 20.96 tok/s MTP-off at UD-Q4_K_M. clm-0009
- 2026-08-05claimDraft-free ngram speculation (ngram-mod) reportedly beats MTP by a wide margin on repetitive/agentic work on Strix Halo — 71 t/s unspeculated to 216 solo on a code-edit probe, with a shared hash pool letting concurrent streams feed each other's drafts (247 pooled across 4 streams, later 302 end-to-end on HIP). clm-0010
- 2026-08-05claimGreedy decode is NOT run-to-run deterministic on the Vulkan backend — same config, same prompt, temperature 0, batch 1, no speculation, three different outputs. clm-0011
- 2026-08-05claim128 GB Strix Halo systems appear to be repricing upward materially — from a roughly $2,500-4,000 enthusiast tier toward $4,500-6,000 — with 128 GB SKUs scarce while 64 GB configurations remain available. clm-0012
- 2026-08-05claimQwen3.6-35B-A3B — our chosen reflex model — is published in FastFlowLM's .q4nx NPU format (23.2 GB, plus a 1.0 GB vision encoder), making a genuinely capable MoE, not a 1-2B classifier, runnable on the XDNA2 NPU. clm-0013
- 2026-08-05claimOn gfx1151, Vulkan measured ~22-24% faster than ROCm on the same 35B-A3B model and build — pp4096 1039 vs 840 t/s, tg128 53.1 vs 43.4 — but the Vulkan build produced GARBAGE OUTPUT for that model, so the numbers describe a broken configuration. clm-0014
- 2026-08-03claimStatic pool utilisation did not predict the 2026-07-21 OOM. clm-0001
- 2026-08-03claimMulti-token prediction gains MORE at heavier quantisation, not less. clm-0002
- 2026-08-03claimQuantized KV cache is mis-implemented for Strix Halo in stock llama.cpp: the code dequantizes to full precision repeatedly during inference, which a discrete GPU hides in cache and this box cannot. clm-0006
- 2026-08-03claimA GPU cache wall around 32-40 MB, past which read speed drops roughly 4x, is a candidate explanation for our own prefill degradation from ~354 tok/s at small context to ~144 tok/s at 128K. clm-0007
July 2026
- 2026-07-30gatemaple-preview: listed — Ternary quantisation at a footprint that opens a second lane on hardware we already own. clm-0018
- 2026-07-26gateling-30-flash: listed — The only candidate that functions as a controlled experiment rather than another data point.
- 2026-07-23incidentAbrupt power loss under sustained load; board never POSTed again — RMA closed (full refund) (blocked, resolved) inc-0005
- 2026-07-21gatelaguna-s-21: listed — Released 2026-07-21 with a Terminal-Bench number far above our incumbent, at a footprint that fits.
- 2026-07-21incidentUnified-memory OOM cascade to kernel panic, then a wedged boot (blocked, resolved) inc-0004
- 2026-07-21contributionggml-org/llama.cpp fork — carrying: Disk slot save/restore silently loses all prompt reuse on hybrid/recurrent models because context checkpoints are never persisted. con-0001
- 2026-07-20claimMTP's speedup tracks how predictable the text is: +81% on code, +21% on freeform, with the real agent workload mix landing around +29-40%. clm-0003
- 2026-07-20runs1 run landed on cfg-0002 (mtp-decode-sweep) run-0002
June 2026
- 2026-06-28claimgpt-oss-120b placed last on our blind conceptual set — 3/8, mean 1.38 — against Qwen3.6-35B-A3B's 7/8, mean 2.12. clm-0004
- 2026-06-28claimNemotron 3 Super's benchmark result is NOT quant-equivalent to its peers and must not be read as a like-for-like verdict on the model. clm-0005
- 2026-06-28runs1 run landed on cfg-0001 (conceptual-c1-c8-blind) run-0003
- 2026-06-27gate5 candidates → acquired: gpt-oss-120b, nemotron3-super, qwen35-122b, qwen36-27b-mtp, qwen36-35b
- 2026-06-27gate5 candidates → screened: gpt-oss-120b, nemotron3-super, qwen35-122b, qwen36-27b-mtp, qwen36-35b
- 2026-06-27runs2 runs landed on cfg-0001 (tool-loop) cfg-0001
- 2026-06-20gate10 candidates → listed: cascade2-30b, gemma4-26b, glm-47-flash, gpt-oss-120b, lfm2-24b, llama4-scout, nemotron3-super, qwen35-122b, qwen36-27b-mtp, qwen36-35b