Corrections
Every claim this lab has withdrawn, kept at its citable id with its full original text — struck through nowhere, hidden never. Each entry names what replaced it; the worked diagnosis of what went wrong lives on the record itself. These entries also appear inline in the Log, badged, in the same timeline as every other learning: this page is a generated filter of that stream, not a shrine.
9 corrections on record · a headline result that gets retracted is the process working — catching it before it became policy is the retraction's whole value
On AI Hydra's admitted HIP/gfx1151/native/Release host boundary, the KingJones R2 Qwen3.8-Flash-Next full-STRIX ROCmFP4 campaign fixed model artifact c6770d7442a06bf1d78edf28cec83e1ec93afdd34664c23ff898807b6b9349fa (121838036032 bytes, publisher revision 069dddb53bab04218d734fa9a771f8a0242ab059), runtime source 36e9acd40e10a87cd3c3ef8ec734668757dc8520, and the independently admitted post-route receipt patch 1f1c3bed910922b415c1be36c9a04c9b7fedc4162aedf504a9da007daebcb4d2. The exact tracked 1,017-byte rotate-bits header was present at blob 75c4881fc322f2e6a6ee9d809e696852531abb8c and the patch clean-applied at +109/-2. The card-authorized targeted llama-server capture build, not an all-target build, then failed because sha256.c could not resolve rotate-bits/rotate-bits.h. Therefore no receipt-backed guard, Tau2 capability, llama-bench floor, served-path, cache, depth or energy result exists for this campaign. This is not a model quality, fit, throughput, capability or source-runtime performance claim. The earlier Unsloth UD-Q4_K_XL/Vulkan records are a separate artifact, runtime and backend series and are not merged with this full-STRIX result.
corrected by: clm-0127 — the withdrawn record keeps its URL and full text; the successor carries the number to cite
- 2026-08-30
This receipt-only R2 capture-build stop remains a valid record of that instrumented variant, but it no longer defines the current KingJones full-STRIX boundary. A later uninstrumented campaign produced bounded Tau2, served QSA, cache, memory and energy evidence in clm-0127.
Qwen3.8-Flash-Next UD-Q4_K_XL (Qwen4-arch preview MoE: 125B MoE / ~6B active + 51B N-gram/PLE tables + vision) on the Unsloth qwen4exp Vulkan build (cfg-0177, commit 250b6144, gfx1151/RADV) — first full benchmark. ADMITTED CLAIM 1 (capability leader): on the standard tau2 airline full-26 suite (seed 42, claude-haiku-4.5 simulator, deterministic reward) it scored mean_reward 0.9231, 24/26 tasks at reward 1.0 (run-0632). This LEADS the tau2 airline board — the prior best was Ornith-1.0-35B UD-Q4_K_XL at 22/26 (0.846). Tool-calling was clean: 184 tool calls, 0 empty-argument calls, 1 tool-error message, 24.7 messages/task. Failures = tasks 7 and 20. ADMITTED CLAIM 2 (energy leader per correct answer): wall-metered join (eng-0275, HA counter-difference) = 256.08 Wh total over 6605 s, 237.55 Wh active above the 10.1 W idle floor, = 9.90 Wh per correct answer (0.300 p @ 30.3 p/kWh). This is ~4x cheaper per correct answer than the ho003 stock-f16 tau2 control (42.06 Wh/correct), because it solves more than twice as many tasks in less wall time. ADMITTED CLAIM 3 (footprint): GTT-resident 76.7 GiB with ~43 GiB free for KV. The 26.8 GiB n-gram/PLE table (per_layer_token_embd, iq4_nl) is placed on CPU automatically by the Vulkan backend regardless of -ot, so it never occupies GTT; KV scales ~48 MiB per 1000 tokens (hybrid QSA attention). Performance: prefill 277-313 tok/s at 5-8K, decode ~22 tok/s short falling to 8.5 at 131K; TTFT ~2.9 s. BLOCKED CLAIMS (explicitly NOT made): single trial (N=1), no variance bound. reasoning_effort=low, not the model's default xhigh — a higher-reasoning result is untested and may differ (both tasks 7 and 20 may be reasoning-recoverable). The runtime is a draft/WIP PR (#27742), not a released/immutable build; the GGUF's advertised LICENSE artifact 404s and conversion-base provenance is unpinned, so this is a LAB result with NO weight redistribution or public/production recommendation. MTP speculative decoding is unavailable (the GGUF lacks MTP head layers) and n-gram speculation did not help this quant. 262K context hits a Vulkan workgroup-count assertion (cap -c <= 262140). No UD-Q2 control or cross-host pair run yet.
corrected by: clm-0125 — the withdrawn record keeps its URL and full text; the successor carries the number to cite
- 2026-08-28
The first-look observations remain published under their original record ids, but this claim no longer carries current capability-leader, energy-per-correct, production-throughput, or role-fit authority. The run lacked the protocol screen, formal guard, repetitions, retained exact full-run serving transport, and an unambiguous admitted energy denominator. clm-0125 records the interim evidence boundary while a protocol-valid intended-config supersession is prepared.
SUPERSEDED (operative claim) as of strix-halo-llamacpp v0.6.10. As written against v0.6.9 (2026-08-22): on hybrid GDN (qwen35moe) targets the then-stable Strix Halo Vulkan fork did NOT guarantee token-exact MTP rollback after a rejected draft. v0.6.8 had introduced a one-line change forcing MTP rollback through full sequence-state checkpoints; v0.6.9 (2026-08-22T04:09:43Z) reverted that line because the full-checkpoint save/restore path deadlocked deterministically on Vulkan with a hybrid target (Qwen3.6-35B-A3B stalled a few hundred tokens into a long response, all threads parked in futex_do_wait, GPU idle, GTT flat). With the revert, rollback returned to the v0.6.4-v0.6.7 fast snapshot-plane restore, which the vendor stated "can diverge slightly from a no-draft run after a rejected draft" - a subtle distribution drift after rejection, not garbled output. That was a deliberate availability / exactness tradeoff by the fork, pending a state-save-path fix. v0.6.10 (clm-0107) re-lands token-exact full-checkpoint rollback; the operative tradeoff claim above no longer holds on the current stable fork.
corrected by: clm-0107 — the withdrawn record keeps its URL and full text; the successor carries the number to cite
draft-mtp speculation on Qwen3.8-27B (stock 3653e6d, gfx1151) peaks at spec-draft-n-max=3 over CLEAN cells — Vulkan Q8_0 17.75 t/s (2.26x its 7.86 no-speculation floor, acceptance 0.626), Vulkan UD-Q4_K_XL 26.68 t/s (2.23x, 0.623), ROCm Q8_0 18.25 t/s (2.33x, 0.6235) — and at n_max >= 4 the feature is BROKEN on this model: after accumulated generation volume in a live session (sequential VARIED prompts at full length; not fresh servers, not one repeated prompt, not short generations), generations start terminating at 1 token with <|im_end|> (id 248046). The decisive measurement is that the TARGET model's own pre-sampling distribution puts im_end at logprob -0.084 (~92%) on a prompt the same server answers normally with speculation off — the speculative path is corrupting the target's forward pass, not merely mis-accepting drafts — and ignore_eos:true restores both correct text AND draft accounting. The hazard compounds into a measurement artifact: llama.cpp reports a 1-token generation as 1,000,000 tokens/s, so unfiltered throughput averages FLATTER exactly the broken cells; a community sweep on this silicon reporting a peak at n_max=5 sits on what is measured here as one of the two worst cells (11/15 degenerate on Q8_0), under a workload shape (single repeated codegen prompt) that this trigger analysis shows cannot reproduce the failure. Recommendation from measurement: draft-mtp ON, n_max hard-capped at 3 until the mechanism is understood upstream.
- 2026-08-15
The headline text's "at n_max >= 4 the feature is BROKEN on this model" carries no backend qualifier and could be read as a fleet-wide (any backend) claim. A same-day cause-isolation matrix (see the second 2026-08-15 amendment below) falsifies that broad reading: flipping ONLY the backend from the failing Vulkan/Q8_0/f16-KV baseline to ROCm, with quant/KV/n_max held at n_max=5 (one of the two historically worst cells), produced ZERO anomalies of any kind across 15 requests AND an extended 45-request confirmation (run-0278, run-0281) -- where the Vulkan baseline reproduced at 11/15 degenerate in the same session shape (run-0277). "BROKEN on this model" should be read as "BROKEN on this model on Vulkan, at this session depth" -- the ROCm result is real evidence, not an absence of evidence, and the standing n_max<=3 hard-cap recommendation is unchanged (it was never contingent on backend), but a reader taking the original sentence as backend-general would be wrong. This also retires the immediately-prior amendment's open question ("a matched ROCm + f16-KV + n_max=7 control... has not been run") -- run-0278/run-0281 is that control, run at n_max=5 instead of 7, which is the more diagnostic choice since it is a known-worse cell on Vulkan.
On gfx1151 at f16 KV, DeepSeek-V4-Flash-0731 UD-IQ3_XXS is a fifth model measured under the per-model, per-phase backend rule (clm-0050), and it lands on the same side as Qwen3.8-27B (clm-0054): ROCm swept the full throughput matrix clean (d0 through d262144) while Vulkan lost the GPU device on both allowed attempts at every depth >=32768 (rc=134/SIGABRT, vk::DeviceLostError; kernel evidence: amdgpu ring timeout, ring reset, "device wedged, but recovered through reset" — 12 reset cycles total across the four failed cells). Even at the one depth Vulkan completed, d0, ROCm decode was already ahead (15.30 vs 12.44 t/s, +23%) while prefill was close (140.80 vs 134.73, ROCm +4.5%); ROCm's prefill lead widens with depth on every model measured so far on this chip. Serving this model at any useful context on this box requires ROCm — Vulkan cannot be trusted to survive a conversation that grows past 32k.
corrected by: clm-0084 — the withdrawn record keeps its URL and full text; the successor carries the number to cite
- 2026-08-18
The stock 3653e6d Vulkan measurements remain valid for that build, but the general conclusion that Vulkan cannot survive past d0 is superseded by the clean baf6360be carried-fork series through d262144 (clm-0084).
AMENDED 2026-08-14 — the rule below is PER-MODEL AND PER-PHASE, not fleet-wide; see the correction history and the amendment note. On gfx1151 at f16 KV, stock Vulkan beat stock ROCm in EVERY cell of a matched matrix measured on Qwen3.6-35B-A3B (one binary commit 3653e6d, depths 0 to 131,072): decode +19-21% at every depth, prefill +4% to +20% growing with depth to 65k. That finding was correct as measured and remains the reference case for this model. It does NOT generalise: the 2026-08-14 perf-matrix sweep found gpt-oss-120b's PREFILL inverts hard in ROCm's favour at depth — Vulkan −52% at d65536, −84% at d131072 (23.1 t/s vs ROCm's 144.3, tight across 3 reps) — while gpt-oss-120b's own DECODE still favours Vulkan (+10 to +16%), same direction as Qwen3.6-35B. Nemotron-3-Super is now a THIRD supporting data point rather than an absence: the OOM first recorded against it (twice, including at an expanded 122 GiB GTT ceiling) was a benchmarking-harness artifact of running stock ROCm under mmap on a model within ~40 GiB of the box's full memory, corrected 2026-08-14 — see clm-0053. On the identical binary with `--load-mode none`, Nemotron-3-Super UD-Q4_K_M at d32768 measures 244.05 pp / 16.43 tg (median of 3 post-reboot reps; an earlier single-rep check the same day landed within 2% at 248.79 pp / 16.39 tg) against Vulkan's own 139.8 pp / 17.56 tg at the same cell: ROCm +74.6% prefill, Vulkan +6.4% decode — the same prefill-favours-ROCm, decode-favours-Vulkan split gpt-oss-120b shows, now measured on a third architecture. The r/LocalLLaMA claim that ROCm leads Vulkan 3.5x at 65k depth — measured on an RDNA2 V620 — still inverts on gfx1151 for Qwen3.6-35B decode/prefill and for gpt-oss-120b decode, but NOT for gpt-oss-120b's own prefill, which is the live counter-example. The community fork (v0.6.1) over stock Vulkan remains a prefill-only win at f16 KV on Qwen3.6-35B: +13/+13/+1/+2.5/+16% by depth, decode unchanged (±1%), and no stride bug (41/41 layers on GPU, CPU utilisation identical to stock) — untested on the other two models.
corrected by: clm-0053 — the withdrawn record keeps its URL and full text; the successor carries the number to cite
- 2026-08-14
Original text ("stock Vulkan beats stock ROCm in EVERY cell") was true of its one-model matrix but read as a fleet-wide backend rule. The 2026-08-14 perf-matrix sweep (gpt-oss-120b, Nemotron-3-Super) falsified the fleet-wide reading: gpt-oss-120b prefill favours ROCm by a wide and growing margin at depth (opposite direction from Qwen3.6-35B and from gpt-oss-120b's own decode). Rule restated as per-model, per-phase. The original Qwen3.6-35B numbers are unchanged and correct; only the scope of the claim was wrong.
- 2026-08-14
The Nemotron-3-Super OOM behind this claim's first amendment ("favours Vulkan only in the degenerate sense") was itself a benchmarking-harness artifact, not a hardware limitation: the perf-matrix and perf-matrix-gtt120 llama-bench queues that produced run-0234/run-0235 never carried --load-mode none, so stock ROCm was benchmarked under mmap on a model within ~40 GiB of the box's full 122 GiB — full diagnosis in clm-0053. Re-run with --load-mode none on the identical binary, Nemotron-3-Super runs cleanly on ROCm (run-0236, run-0237: 244.05 pp / 16.43 tg, median of 3 post-reboot reps at d32768; an earlier single-rep check the same day, run-0243/run-0244, landed within 2% at 248.79 pp / 16.39 tg) and becomes a THIRD supporting data point for the per-model/per-phase rule rather than an absence: ROCm +74.6% prefill / Vulkan +6.4% decode against Vulkan's own run-0190/run-0191 (139.8 pp / 17.56 tg) — the same prefill-favours-ROCm, decode-favours-Vulkan split gpt-oss-120b shows. The per-model, per-phase rule itself is unchanged; only Nemotron's status in it moves from absent to confirming.
SUPERSEDED by clm-0042's per-task measurement, which found the true energy cost roughly 10x lower — this run's figures are whole-arm totals padded by model loading and non-scoring tasks, not the model's actual energy per answer, and must not be used to rank models. As measured here: a correct τ² answer cost 78.0 Wh on the 122B with f16 KV, 100.1 Wh on Nemotron, and 115.3 Wh on the 122B with q8_0 KV, but only 8-27% of each arm's wall time fell inside a scored task. aihydra's power envelope stands on its own: idle 10.1 W, 150-168 W under inference, peaking at 218 W.
corrected by: clm-0042 — the withdrawn record keeps its URL and full text; the successor carries the number to cite
SUPERSEDED: the Pass^1 = 1.000 reported by this run came from a 3-task subsample biased toward the domain's easiest tasks — the other two of the original five never terminated and were excluded as infrastructure errors. The sustained score across a realistic sample is clm-0037's 0.545 (n=22), which is the number to cite for this model on tau2 airline. This run was nonetheless the project's first genuine capability measurement rather than a throughput number, and it proved the harness and scoring path work.
corrected by: clm-0037 — the withdrawn record keeps its URL and full text; the successor carries the number to cite
RETRACTED: the per-model reward rankings and quantised-KV cost reported by this run do not hold — they came from 5-task tau2 arms whose ~0.40 run-to-run noise and 40-step cap bias were only characterised afterward (clm-0036), so the reward numbers below are not usable. The corrected 122B score is clm-0037's; the corrected cross-model comparison is clm-0039's. What survives is categorical, not scored: the 122B fails to terminate some tasks with thinking on, and gpt-oss fails the domain outright.
corrected by: clm-0037 clm-0039 — the withdrawn record keeps its URL and full text; the successor carries the number to cite