Method
Why any of this should be believed. Every published number is wall-measured on the lab's own hardware, carries its provenance and confidence, joins to the run and configuration that produced it, and is corrected in public when it falls. The pages below are the machinery: the protocol that constrains how measurements are made, the standing rules extracted from measurement failures, and the corrections the process has already produced.
The protocol
Two measurements, not one: capability is what a model can do — slow, expensive, run rarely; performance is how fast it delivers that — cheap, run often, and only meaningful behind a guard, the cheap cliff-detector run at every performance configuration. What makes that split legal is a classification of every configuration lever by how it can break capability, and a machine-readable constraint file the harness checks before every run — prose is read once and remembered badly; checks fire every time.
- The enforced protocol — bench/protocol.json rendered: pins, units, comparison rules, known builds
- The full protocol document — the prose version, with the reasoning behind every rule
fleet default build 3653e6d ·11 known builds, each pinned with an anchor pair · user simulator pinned to openrouter/anthropic/claude-haiku-4.5 ·6 comparison rules enforced
Corrections — the record correcting itself
Corrections live in the Log timeline with every other entry, badged inline — visible, not enshrined. The registry below them is a generated filter of that same stream: every retracted or superseded claim, what was wrong, and what fixed it.
- The corrections registry — 7 withdrawn claims, kept citable with their replacements
On AI Hydra's admitted HIP/gfx1151/native/Release host boundary, the KingJones R2 Qwen3.8-Flash-Next full-STRIX ROCmFP4 campaign fixed model artifact c6770d7442a06bf1d78edf28cec83e1ec93afdd34664c23ff898807b6b9349fa (121838036032 bytes, publisher revision 069dddb53bab04218d734fa9a771f8a0242ab059), runtime source 36e9acd40e10a87cd3c3ef8ec734668757dc8520, and the independently admitted post-route receipt patch 1f1c3bed910922b415c1be36c9a04c9b7fedc4162aedf504a9da007daebcb4d2. The exact tracked 1,017-byte rotate-bits header was present at blob 75c4881fc322f2e6a6ee9d809e696852531abb8c and the patch clean-applied at +109/-2. The card-authorized targeted llama-server capture build, not an all-target build, then failed because sha256.c could not resolve rotate-bits/rotate-bits.h. Therefore no receipt-backed guard, Tau2 capability, llama-bench floor, served-path, cache, depth or energy result exists for this campaign. This is not a model quality, fit, throughput, capability or source-runtime performance claim. The earlier Unsloth UD-Q4_K_XL/Vulkan records are a separate artifact, runtime and backend series and are not merged with this full-STRIX result.
Qwen3.8-Flash-Next UD-Q4_K_XL (Qwen4-arch preview MoE: 125B MoE / ~6B active + 51B N-gram/PLE tables + vision) on the Unsloth qwen4exp Vulkan build (cfg-0177, commit 250b6144, gfx1151/RADV) — first full benchmark. ADMITTED CLAIM 1 (capability leader): on the standard tau2 airline full-26 suite (seed 42, claude-haiku-4.5 simulator, deterministic reward) it scored mean_reward 0.9231, 24/26 tasks at reward 1.0 (run-0632). This LEADS the tau2 airline board — the prior best was Ornith-1.0-35B UD-Q4_K_XL at 22/26 (0.846). Tool-calling was clean: 184 tool calls, 0 empty-argument calls, 1 tool-error message, 24.7 messages/task. Failures = tasks 7 and 20. ADMITTED CLAIM 2 (energy leader per correct answer): wall-metered join (eng-0275, HA counter-difference) = 256.08 Wh total over 6605 s, 237.55 Wh active above the 10.1 W idle floor, = 9.90 Wh per correct answer (0.300 p @ 30.3 p/kWh). This is ~4x cheaper per correct answer than the ho003 stock-f16 tau2 control (42.06 Wh/correct), because it solves more than twice as many tasks in less wall time. ADMITTED CLAIM 3 (footprint): GTT-resident 76.7 GiB with ~43 GiB free for KV. The 26.8 GiB n-gram/PLE table (per_layer_token_embd, iq4_nl) is placed on CPU automatically by the Vulkan backend regardless of -ot, so it never occupies GTT; KV scales ~48 MiB per 1000 tokens (hybrid QSA attention). Performance: prefill 277-313 tok/s at 5-8K, decode ~22 tok/s short falling to 8.5 at 131K; TTFT ~2.9 s. BLOCKED CLAIMS (explicitly NOT made): single trial (N=1), no variance bound. reasoning_effort=low, not the model's default xhigh — a higher-reasoning result is untested and may differ (both tasks 7 and 20 may be reasoning-recoverable). The runtime is a draft/WIP PR (#27742), not a released/immutable build; the GGUF's advertised LICENSE artifact 404s and conversion-base provenance is unpinned, so this is a LAB result with NO weight redistribution or public/production recommendation. MTP speculative decoding is unavailable (the GGUF lacks MTP head layers) and n-gram speculation did not help this quant. 262K context hits a Vulkan workgroup-count assertion (cap -c <= 262140). No UD-Q2 control or cross-host pair run yet.
SUPERSEDED (operative claim) as of strix-halo-llamacpp v0.6.10. As written against v0.6.9 (2026-08-22): on hybrid GDN (qwen35moe) targets the then-stable Strix Halo Vulkan fork did NOT guarantee token-exact MTP rollback after a rejected draft. v0.6.8 had introduced a one-line change forcing MTP rollback through full sequence-state checkpoints; v0.6.9 (2026-08-22T04:09:43Z) reverted that line because the full-checkpoint save/restore path deadlocked deterministically on Vulkan with a hybrid target (Qwen3.6-35B-A3B stalled a few hundred tokens into a long response, all threads parked in futex_do_wait, GPU idle, GTT flat). With the revert, rollback returned to the v0.6.4-v0.6.7 fast snapshot-plane restore, which the vendor stated "can diverge slightly from a no-draft run after a rejected draft" - a subtle distribution drift after rejection, not garbled output. That was a deliberate availability / exactness tradeoff by the fork, pending a state-save-path fix. v0.6.10 (clm-0107) re-lands token-exact full-checkpoint rollback; the operative tradeoff claim above no longer holds on the current stable fork.
Lessons — standing rules with their scars
Rules extracted from measurement failures, not a diary of them. Each cites the claim it was drawn from, so the worked diagnosis is one click away.
- 1. A small-n benchmark can silently become a coin flip
- 2. Sanity-check an instrument against known bounds before trusting its output
- 3. Pin the user simulator, or cross-model comparisons are confounded
- 4. A bound chosen for convenience is not a property of the system
- 5. Match the energy denominator to the question
- 6. A benchmark harness that drifts from the serving config produces capability
- 7. Watch the kernel signal, not the userspace symptom, for a detection-lag defect
Essays — the long-form record
Instruments
The bench scripts are the instrument set — each enforces part of the protocol rather than trusting a reader to remember it. Energy is read from a wall meter's cumulative kWh counter, differenced across each run's window; never modelled, never joules.
scripts: cache-trace-probe.py cache-trace.sh capability-probe.sh env-lock.sh export-tool-schemas.sh grammar-ceiling.sh guard-probe.py guard.sh ha-energy-join.py ho004-derive.py ho004-dump-evidence.py ho005-revised-energy-join.py ho005-revised-ornith-batch-energy-join.py ho009-v0610-energy-join.py ho012-122b-energy-join.py ho014-energy-join.py kv-quality.sh preflight.sh protocol-check.sh run.py runmeta.sh spec-ab.sh sweep.sh tool-calling.sh
meters: TP-Link smart plug via Home Assistant AI Hydra energy/power sensors via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy; peak/mean drawn from sensor.hardware_ai_hydra_power) aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy) aihydra energy sensor via Home Assistant (5-minute statistics on sensor.hardware_ai_hydra_energy, linearly interpolated to the run window) aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, 5-minute statistics interpolated to the run-meta window) aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s history samples interpolated to the run-meta window) aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the run-meta window) aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to each run-meta window; two windows summed, not a single counter-diff over the outer span) aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the session window) aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the transcript-derived window) aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the recorded window) aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, high-confidence raw ~10s-resolution samples linearly interpolated to the recorded window) aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s samples linearly interpolated to the recorded window) aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the recorded window; power from sensor.hardware_ai_hydra_power where available) aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples linearly interpolated to the recorded window; power from sensor.hardware_ai_hydra_power where available) aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy; raw ~10s-resolution history samples queried from hborchestrator using retained HA history) aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the run window) aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s samples queried by hborchestrator via Nabu Casa API, idle baseline 10.1 W) aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-history samples queried by hborchestrator via Nabu Casa API and linearly interpolated to the recorded window; power from sensor.hardware_ai_hydra_power) aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-history samples queried by hborchestrator via Nabu Casa API and linearly interpolated to the window; power from sensor.hardware_ai_hydra_power) aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-history samples queried by hborchestrator via Nabu Casa and interpolated to the window; power from sensor.hardware_ai_hydra_power) aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, sampled and interpolated to the window; power from sensor.hardware_ai_hydra_power) aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-history samples queried by hborchestrator via the Nabu Casa API and linearly interpolated to the window; power from sensor.hardware_ai_hydra_power) aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-history samples queried by hborchestrator via the HA API and linearly interpolated to the window; power from sensor.hardware_ai_hydra_power) aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-history samples queried by hborchestrator via the HA API and linearly interpolated to the band edges; power from sensor.hardware_ai_hydra_power) aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw history samples queried by hborchestrator via the HA API and linearly interpolated to the window edges; power from sensor.hardware_ai_hydra_power) aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw history queried via ha_get_history and linearly interpolated to the window edges; power cross-checked from sensor.hardware_ai_hydra_power 5-minute statistics) AI Hydra Home Assistant cumulative kWh counter sensor.hardware_ai_hydra_energy, queried by the authorized hborchestrator reader and linearly interpolated to retained exact UTC edges. aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw history queried via ha_get_history at the window edges and a midpoint; power cross-checked from sensor.hardware_ai_hydra_power hourly statistics) · corpus pinned in bench/CORPUS.md
Contributions — upstream work, fork-carrying disclosed
- con-0017 · kingjones777/Qwen3.8-Flash-Next-ROCmFP4-STRIX-GGUF report — carrying Immutable publisher model-page revision 069dddb53bab04218d734fa9a771f8a0242ab059 describes the full-STRIX artifact and the required kingjones30/ROCmFPX runtime. Its long-context table was measured on gfx1151 with ROCm 7.2.4, different target prompt counts and context allocations, and no HaloBench 512-checkpoint/ 2048-token-cadence proof. The page reports full-STRIX rows through c131072, while its c262144 rows are explicitly STRIX_LEAN. HaloBench used the full-STRIX artifact on HIP 7.1.52801-9999, exact cache-false prompt_n targets, c262144, and the corrected checkpoint flags. The 32k-131k rates are close in magnitude, but neither source transfers: runtime version, prompt bytes/counts, allocation, repetition method and checkpoint boundary differ. The publisher page is a comparison source and attribution caveat, not evidence for a HaloBench number. · upstream
- con-0018 · ggml-org/llama.cpp pr — open PR #27311 "Scheduler UMA ring buffer (+ sanitizer and fixes)" adds an input ring buffer on UMA devices so the host cannot clobber in-flight graph inputs. OPEN / unmerged as of the 2026-08-26 head commit. Carried on aihydra as branch pr27311 (commit c530ea79c) and is the load-bearing dependency of the Qwen3.8-27B worker serving config (cfg-0183): 0 EMPTY across a 106-minute 4-slot run vs the master build's ~4-minute failure (run-0660), and it is why the worker can serve --parallel 4 at all where the earlier Q8_0 config had to pin --parallel 1. A reproducibility hazard until merged: any rebuild of this worker must fetch the PR branch, not stock master. See cfg-0183, clm-0130. · upstream
- con-0019 · ROCm/rocm-systems fork — carrying "Retained PM4 dispatch" — a custom AMD HIP/ROCm runtime (rocm-systems ilintar-experiments @ 78d1160, the CLR/hipamd graph path) that keeps a graph's low-level PM4 command buffer resident and replays it, cutting per-token launch overhead. Activated by ENABLE_RETAINED_PM4=1 -> DEBUG_HIP_GRAPH_PM4=1 with HIP graphs on. This is a RUNTIME change (libamdhip64), not a llama.cpp change. Carried on aihydra as the custom rocr+clr runtime behind the pwilkin 27B build (cfg-0185/cfg-0186). It is the ONE piece of the pwilkin stack with NO upstream PR: the author states "all of this is in upstream PRs except for the PM4 stuff which I don't expect to be accepted (but I might retire for proper HRX support when it's there)." A novel-fork reproducibility hazard: any rebuild must build this custom runtime from source; there is no merge path. Measured effect on our worker model is small (~+4%, clm-0135) — it is not the source of the pwilkin speed advantage. The pwilkin llama.cpp changes, by contrast, ARE upstream PRs (e.g. #27311, con-0018). · upstream
- con-0014 · Nathanw1014/strix-halo-llamacpp report — open Primary-source runtime-watch record: Nathanw1014/strix-halo-llamacpp v0.6.11, published 2026-08-24, fixes the fork-only image-request regression introduced by the fork's v0.6.8 DFlash M-RoPE change. For a position-sharing draft-mtp context, the old token-space trim left the draft behind the target after an image span and led to inconsistent M-RoPE positions, llama_decode failure, and HTTP 500. Commit 0eb528051a56f34567312ce63ab4e14a3fc71d89 records whether a draft cache is dense-row indexed: DFlash/DSpark retain token-space trimming, while draft-mtp and eagle3 use target position space. The release's gfx1151 Vulkan/RADV smoke on Qwen3.8-27B-UD-Q6_K_XL plus mmproj-F16 reports that draft-mtp n_max=3 image requests now return HTTP 200 with speculation remaining active; its DFlash image and speculation-off controls are unchanged. The release asset digest is sha256:a4306edefdefb2eff925cbc43cbf32cd42cf07c00e06011e906957217bf1ddad and MANIFEST.txt identifies the portable payload source as 0eb528051. This is a fork-vendor functional-correctness report, not a HaloBench benchmark: the release explicitly says its smoke checks are not BENCHMARKS.md throughput and makes no throughput claim. It has no text-only agent outcome, guard result, capability score, energy join, or served-path/floor pair. The regression never applied to upstream llama.cpp or to fork configurations with speculation off, so it cannot rebaseline existing text-only Qwen3.8 throughput/capability records or replace their recorded v0.6.10/stock runtime provenance. A future vision/MTP prep card is justified only as a separately scoped, on-box validation of this exact payload with multimodal request controls and the house guard; it is not a reason to dispatch a benchmark or change the production pick now. · upstream
- con-0013 · LaurentZuijdwijk/llama.cpp fork — carrying Vulkan fork used by the Qwen3.8-27B ROCmFPX/DFlash2 autonomous tuning handoff. The handoff records commit 16f0799a6 and Mesa RADV 26.0.3. Its production-performance evidence remains screening-only until the required floor, guard, and artifact-identity records are captured.
- con-0015 · ggml-org/llama.cpp issue — open Raw upstream diagnostic comment by frizikk on open issue #25618, retrieved directly from GitHub on 2026-08-25. Surface: pinned llama.cpp sources, Vulkan, Qwen3.8-27B Q6_K target plus Q8_0 native MTP, flash attention on, f16 K/V, greedy sampling, and --spec-draft-n-max 1. Its first confirmed target-verify boundary is the natural generation-1 layer-0 Q8_0 x F32 MUL_MAT: sequential N=1 uses F32 MUL_MAT_VEC, while speculative N=2 stages activations as Q8_1 for MMVQ; the direct predecessor is exact but every one of the 5120 output values changes. Replaying the N=2 fixture through the non-MMVQ F32 path made row 0 bit-exact, but later full generation logits still differ, so the narrow diagnostic gate is not a complete fix. The same comment identifies a second N-dependent ADD/RMS-reduction boundary and explicitly confines its flash-attention observation to one captured fixture. This is upstream operator-level evidence, not a HaloBench benchmark or a production fix. It corroborates the standing varied-prompt/raw-token MTP-invariance requirement, but does not identify any local Qwen3.8 result as affected: the closest recorded local Vulkan arms use Q8_0 targets and n_max=3 (cfg-0068/run-0269 and cfg-0147/run-0573), already retain a negative varied-prompt invariance result and are not promoted as baseline/MTP equivalent. It has no direct transfer to the local ROCm Qwen3.8 records. · upstream
- con-0016 · ggml-org/llama.cpp pr — open Raw upstream comment by alexpooley on open PR #25863, retrieved directly from GitHub on 2026-08-25. The report is HIP/ROCm on the integrated gfx1151 surface, using Qwen3-4B-Instruct-2507 Q8_0, llama-perplexity, c4096, b4096, 12 chunks, and variable ubatch. Its master-versus-proposed-fix PPL table is stable at ub4096 (9.1542 versus 9.1542), but degrades on master as prompt chunking increases: ub2048 11.8957 versus 9.1527, ub1024 2259.7396 versus 9.1557, and ub512 49142.0371 versus 9.1594. The proposed explanation is that the scheduler does not protect caller-writable input during chunking; the author disclosed AI assistance for identifying and implementing the patch, while stating that the final change was personally tested and reviewed. This is a source-bound upstream report, not a HaloBench result or a validated fix. It shares HIP/gfx1151 and low-ubatch shape with the local Qwen3.5-122B production-optimisation matrix (cfg-0151, b2048/ub512), but differs in model, weight quant, tool, corpus, batch, context and runtime. The local four-item guard (run-0588) passed but is not a chunked-input forward correctness test; it cannot prove this source does or does not apply. Therefore no recorded Qwen3.5 number is reinterpreted or invalidated, but future or modified HIP/gfx1151 low-ubatch performance work needs the new exact-fingerprint chunked-input control before a claim is admitted. · upstream
- con-0012 · ciru-ai/ROCmFPX fork — carrying The Kairic Edge Qwen3.8-27B-IU4 build (branch kairic-edge-qwen38-27b-v1.1, commit e1da26bb8) from ciru-ai/ROCmFPX — a vendor-authored fork that binds the ROCmFPX binary (called "TheRock" stack here because the carrier vendor certifies it against TheRock 7.14 / AMD clang 23.0.0) to a custom 4-bit IU4 sidecar layout: a Q4-ish quant GGUF (Qwen3.8-27B-IU4-Kairic-Edge.gguf) + three PROMPTFORGE_* .pfs sidecars (FFN, GDN, GDN-Output) that accelerate the FFN and GDN projector paths on gfx1151, plus a --kairc-edge server flag and a compat mode (KAIRIC_EDGE_COMPATIBILITY_MODE) that enables tool-calling where fast-greedy forbids grammar/tool calls. HK-001 on aihydra rebuilt this at commit e1da26bb8 (binary e121a4d3, model sha 360caf73, sidecar hashes adcbb90a/82f93129/3b07e7b1, gfx1151 / AMD clang 23.0.0). It is a carried fork — results produced with it are fork-specific and must not be attributed to stock llama.cpp, and not to any one bundled patch. Backwards the fixed GDN output path is NOT accelerated (gdn_output_fallback 48) on this aihydra build and the toolchain differs from the vendor-certified one (ROCm 7.1 vs 7.14, LLVM 21 vs AMD clang 23), so the vendor's throughput claims do not transfer; every number must be treated as measured-here on this exact build. [2026-08-24 v1.2 RUNTIME MANDATE — HK-V12-ANNOTATE, per t_b05473c9] The vendor released runtime tag kairic-edge-qwen38-27b-v1.2 @ 205a3e5f (HF repo updated 2026-08-23T14:28:32Z): the v1.1 native IU4 M65 verifier could flip a greedy target token on a low-margin reproduced case; v1.2 defaults to the strict-compact authoritative verify path (unsafe native only via KAIRIC_UNSAFE_NATIVE_M65_VERIFY=1). GGUF + all three .pfs sidecars are byte-unchanged (gguf sha 360caf73…), so throughput records stand. Vendor validation: 6/6 exact-match to frozen no-spec target, draft acceptance 99.73%. ALL FUTURE Kairic Edge box runs MUST build from the v1.2 tag; every HK run recorded to date used v1.1 and carries a comparability annotation on its claim (clm-0116, clm-0118, clm-0119). · upstream
- con-0005 · Akicou/diffuse-cpp fork — carrying Makes inclusionAI LLaDA2.2-flash (100B-A13B block-diffusion MoE) run GPU-accelerated on AMD Strix Halo (gfx1151) and benchmarkable as an agent. Over the Akicou base (b799157): GPU offload of the MoE forward to the HIP/ROCm backend; OpenAI tool-calling in diffuse-server; a manual soft_max attention path that works around a gfx1151 flash_attn_ext mask leak [clm-0104]; a resurrected + GPU-resident inter-step KV cache with in-place flash decode and cross-turn prompt reuse (the key speedup, ~1 -> ~5-12 tok/s at long context) [clm-0103]; and a faithful port of the model's Levenshtein editing (M2T + T2T + DELETE/SPLIT + anti-loop + post-steps) restoring its self-correction. Fork to be published at headbouyJB/diffuse-cpp; patch currently held on aihydra. Base runtime AND the diffuse.* GGUF are a matched pair, so this carries the Akicou base rather than rebasing. · upstream
- con-0006 · ggml-org/llama.cpp report — open New llama.cpp #25618 follow-up by snick525 (2026-08-22T02:11:18Z) on RDNA4 (AMD Radeon AI PRO R9700 / gfx1201, 32 GB, Vulkan/radv), target Qwen3.8-27B Q6_K_XL (MTP head intact), greedy (temperature 0, top_k 1, top_p 1, seed 42), f16 K/V. The reporter additionally built open PR #27342 (DFlash2, an external 2B draft model) so the built-in MTP head is no longer the only speculation type in the comparison. Three findings against a --spec-type none baseline on the same binary: 1. DIVERGENCE IS DRAFTER-INDEPENDENT. At -c 65536 / f16 KV the code prompt diverges at the same first-difference byte 204 for the built-in MTP head AND for DFlash2 at three drafter quants (Q4_K_M, Q8_0, BF16); all four configs are lossless on a reasoning prompt. Same first-diff byte holds across two context sizes (16384 and 65536) and two builds. Whatever drifts is on the target/verify side; the drafter is not a variable in it. 2. DRAFTER QUANTIZATION HAS NO EFFECT. DFlash2 Q4_K_M vs Q8_0 vs BF16 produce byte-identical outputs on both prompts and both context sizes, and identical acceptance counters (draft 347, accepted 283). Consistent with the verify path emitting only the target's own token: a drafter can only change output by proposing different tokens, and here even a Q4 drafter proposes identically, so drafter precision can be ruled out when bisecting. 3. KV-CACHE QUANTIZATION IS ITSELF OUTPUT-MOVING AND COMPOUNDS. Swapping f16 for q8_0 KV (everything else fixed) turns the reasoning prompt from LOSSLESS under f16 to DIVERGES at byte 155 under q8_0 for both MTP and DFlash2, and moves the code prompt's first-diff from byte 204 to byte 277. Separately, KV quantization alone moves greedy output with no speculation on either side: --spec-type none with f16 KV versus q8_0 KV differs on the code prompt at byte 561. The reporter notes q8_0 KV is accordingly never output-preserving and compounds with whatever this issue turns out to be. Throughput context (mentioned only because it came up in #27342, and reported with an active caveat): -c 65536 / f16 KV on Qwen3.8-27B Q6_K_XL on a 32 GB card gave 23.8 tok/s no-spec, 40.4 tok/s draft-mtp n_max 1, 61.9 tok/s DFlash2 Q4_K_M, and 29.4 tok/s for DFlash2 Q8_0 (the Q8 and BF16 drafter rows are NOT drafter-cost evidence: at 31.2-31.6 GB they cross the card's usable ceiling and silently fall back to host memory per #26432, with acceptance unchanged). Below the ceiling every drafter quant performs identically, so the Q8 spill is a memory-footprint artefact, not a drafter-precision finding. Upstream diagnostic record only. Surface is RDNA4 R9700/gfx1201 Vulkan, not our ROCm/Vulkan Strix Halo gfx1151 aihydra box, so no HaloBench benchmark claim is made from these figures. · upstream
- con-0007 · Nathanw1014/strix-halo-llamacpp report — open Nathanw1014/strix-halo-llamacpp (the Strix Halo Vulkan fork tracked in con-0002) shipped v0.6.8 and v0.6.9, two stable releases that carry a deliberate MTP rollback EXACTNESS tradeoff on hybrid GDN targets. v0.6.8 (2026-08-22T02:40:00Z) added DFlash2 speculative decoding support for mmproj/vision (community-reported vision-prompt HTTP 500 fixed by injecting draft rows at dense per-token positions and triming the draft cache in token space) and introduced a one-line change, "common: use full checkpoints for MTP rollback", that forced MTP rollback through full sequence-state checkpoints. v0.6.9 (2026-08-22T04:09:43Z) REVERTS that one line (revert commit a17e843, payload Nathanw1014/llama.cpp@a17e8432b, branch strix-halo-vulkan): on Vulkan with a hybrid GDN target (qwen35moe, e.g. Qwen3.6-35B-A3B) the full-checkpoint save/restore path deadlocked a few hundred tokens into a long response - every server thread parked in futex_do_wait with the GPU idle and GTT flat, reproduced deterministically and bisected on hardware. v0.6.7, CPU, and DFlash2 (incl. the vision path) were unaffected. The vendor states the tradeoff plainly: with the revert, MTP rollback returns to v0.6.4-v0.6.7 behaviour - fast snapshot-plane restore that CAN DIVERGE SLIGHTLY from a no-draft run after a rejected draft (called "a subtle distribution drift after rejected drafts, not garbled output"). Re-land of the full-checkpoint restore is deferred until the state-save path is fixed on Vulkan. Speed figures in the notes are single runs, explicitly NOT the BENCHMARKS.md protocol. Treat as an upstream/report record only: single-payload fork-vendor verification on gfx1151 Vulkan, not a HaloBench benchmark result, and not inherited into any HaloBench number. SUPERSEDED as an operative statement by fork v0.6.10 (see con-0009/clm-0107, 2026-08-22T12:03:16Z), which restored token-exact MTP rollback (root cause: the server re-verifying replayed draft tokens after a checkpoint restore) — this record stands as the historical v0.6.8/v0.6.9 report. · upstream
- con-0008 · ggml-org/llama.cpp report — open llama.cpp #25618 comment #16 by F-Mangini (2026-08-22T20:32:01Z) corroborates the quantized-target divergence on a NEW model/OS/GPU family: the official LiquidAI LFM2.5 DSpark pair also diverges from vanilla on a quantized target while an F16 target preserves greedy parity. Environment: llama.cpp b10566 (bb4caa754), Vulkan build, Windows 11, NVIDIA GTX 1660 SUPER 6 GB (driver 551.76). Target: LiquidAI/LFM2.5-2.6B-GGUF Q8_0 (SHA-256 36587fdf27bdfc69caf2637273679a0870ec155162161bde6fd16e8c70bdb757); draft: LiquidAI/LFM2.5-2.6B-DSpark-GGUF Q8_0 (SHA-256 85a98fafd9af1328b6876fd1360d7ed69e74c6cefc14dd07fb6306e1940386c87). Minimal raw /completion reproduction without a chat template, server `-c 4096 -np 1 -fa on -ngl 99 -ctk f16 -ctv f16`, requests temperature=0 seed=42 n_predict=256, on the LRU-cache Python prompt. Quantized-target finding: vanilla Q8_0 target outputs SHA-256 2eeb8174..., same target + DSpark --spec-draft-n-max 1 outputs a8d38c69... (draft accepted 107/147) -- DIVERGES. Differential controls: (1) Q8_0 target + DSpark loaded but --spec-draft-p-min 1 (zero draft tokens) exactly matches vanilla Q8_0; (2) the active Q8_0 DSpark run is deterministic but consistently differs from vanilla; (3) the SAME Q8_0 draft against an F16 target produces byte-identical greedy output to the F16 vanilla target while accepting 40/64 proposed tokens over 64 generated tokens -- i.e. F16 target preserves greedy parity. The Q8_0 mismatch also persists with --spec-draft-n-max 1 and target KV in F16, so it is not caused by a larger speculative block or by quantized target KV. Notably the n_max=1 divergence differs from the Qwen3 boundary reported earlier in the same issue, where n_max=1 was said to remain lossless. Reporter concludes this is consistent with path-dependent target numerics between sequential single-token decoding and the speculative verification path for quantized weights, and confirms the issue affects LFM2 DSpark, NVIDIA Vulkan, Windows and Q8_0 -- not only the Q4/MTP configs in the original report. Upstream diagnostic record only: surface is NVIDIA Vulkan/Windows, not our ROCm/Vulkan Strix Halo gfx1151 aihydra box, so no HaloBench benchmark claim is made from these figures. · upstream
- con-0009 · Nathanw1014/strix-halo-llamacpp report — open Nathanw1014/strix-halo-llamacpp (the Strix Halo Vulkan fork tracked in con-0002) shipped STABLE v0.6.10 (2026-08-22T12:03:16Z) that REVERSES the v0.6.8/v0.6.9 MTP rollback-EXACTNESS tradeoff recorded in con-0007/clm-0105. Token-exact MTP rollback returns on hybrid GDN (qwen35moe) targets (e.g. Qwen3.6-35B-A3B). Root cause of the v0.6.9 deadlock was NOT the state-save path: the server was RE-VERIFYING replayed draft tokens after a checkpoint restore, which livelocked the slot. Fix lands in two commits — "server: do not re-verify replayed draft tokens after a checkpoint restore" (9c5d899) and re-apply "common: use full checkpoints for MTP rollback" (f25eefe). With the re-verify eliminated, MTP rollback through full sequence-state checkpoints is token-exact again and the long-run stall is gone (smoke: Qwen3.6-35B-A3B MTP Q6_K 800-token repro 13.9 s no stall; 3000-token run 46.4 s / ~65 tok/s past the former stall horizon). v0.6.10 ALSO adds DSpark speculative-decode support for bailingmoe3 (Ling 3.0), cherry-picked from upstream llama.cpp PR #27508 (merged 2026-08-22T09:19:49Z, btw616). Runtime side is complete; the reworked bailingmoe3 forward pass was A/B verified byte-identical to the prior tip on Ling-3.0-tiny (greedy, 96 tokens), but the DSpark path is NOT exercisable end-to-end until Ling-3.0-flash draft GGUFs are published (inclusionAI/Ling-3.0-flash-dspark has model.safetensors up, GGUFs not out as of 2026-08-22T16:19Z). Also gates the RADV coopmat LDS pad of 2 to RADV >= 25.3 (older drivers violated VUID-08986). Payload built from Nathanw1014/llama.cpp@2586f6ed (branch strix-halo-vulkan). Speed figures are single runs, explicitly NOT the BENCHMARKS.md protocol; a pre-existing CPU-only divergence in one 800-token exactness matrix cell (token 776) predates v0.6.8 and is unrelated to the rollback path. Treat as an upstream/fork release-note record only: single-payload fork-vendor verification on gfx1151 Vulkan, not a HaloBench benchmark result, and not inherited into any HaloBench number. · upstream
- con-0010 · ggml-org/llama.cpp report — open llama.cpp #27122 comment by mazinist (2026-08-22T23:51:13Z) independently confirms the MTP/CUDA multi-GPU lockup originally reported by tripletto, and validates a workaround on a completely different platform from the earlier zyxyunxin note (#issuecomment-5310571689, 2026-08-17). Hardware: Threadripper Pro 3945WX (WRX80, Gigabyte MC62-G40), 4x RTX A4000 16GB (PCIe 4.0 x16, NO NVLink / no P2P — nvidia-smi topo shows NODE), Ubuntu 24.04, kernel 6.14, closed driver 595.84, llama.cpp build 749f688fc (Aug 21, CUDA 13.2). Repro: Qwen3.8-27B-UD-Q6_K, --split-mode tensor -ts 25,25,25,25, --spec-type draft-mtp --spec-draft-n-max 3, 131072-token deep-context prefill via llama-benchy. 6/6 hard crashes, dying 2-4 min into the prefill; symptom on this platform is harsher than a lockup — the entire machine hard-powers-off (BMC logs "Power off/down", no Xid/AER/panic). KEY DISCRIMINATOR: the same workload on 2 GPUs (which uses the internal 2-device AllReduce path, not the meta-backend) is completely stable with MTP enabled, so the defect is specific to the multi-GPU tensor-split MTP path. WORKAROUND CONFIRMED: LLAMA_GRAPH_REUSE_DISABLE=1 lets the exact crash config survive the full 131072 prefill (610.82 t/s prefill, 27.53 t/s decode @131K, 38.43 t/s @4K), consistent with PR #24549's mechanism (graph reuse leaves dangling per-device tensor references when MTP/target contexts share memory under SPLIT_MODE_TENSOR). Also matches the reporter's notes: no-MTP tensor is stable, layer split never affected, and lockup frequency scales with --spec-draft-n-max. Upstream diagnostic record only: surface is CUDA multi-GPU, not our single-GPU ROCm/Vulkan Strix Halo gfx1151 aihydra box, so no HaloBench benchmark claim is made from these figures. · upstream
- con-0004 · ggml-org/llama.cpp report — open Independent Strix Halo Vulkan follow-up on llama.cpp #25618. The report reproduced the earlier F16-V byte-exact PASS on the original prose prompt using matching Qwen3.8-27B Q6_K target and Q8_0 MTP draft artifacts, llama.cpp 9d57ce456c94d241dde672b2db9cf18879766568, f16 K/V cache and the reporter's greedy request contract. Under the same runtime, server flags, template path and completion path, five additional prompts had 0/5 exact-token parity, and a separate fixed-sampler replay of those same prompts also had 0/5 exact parity with unchanged first mismatch locations. Treat this as an upstream diagnostic report, not a HaloBench measured-here run: it establishes prompt sensitivity of the exactness check for that target/draft/runtime combination and argues that a single exact MTP trajectory is not evidence of general baseline/MTP invariance. · upstream
- con-0002 · ggml-org/llama.cpp fork — carrying Nathanw1014's Strix Halo Vulkan fork, staged first as the v0.6.4 portable payload at build baf6360be and then as the reviewed v0.6.6 portable payload at build 7b6c6133/source 7b6c61330edf370659f531932e0b91aca67ba055. The payloads carry RADV wave32/cooperative-matrix tuning, dense and sparse prefill work, DeepSeek-V4 Lightning Indexer support and upstream merges. Results produced with them are fork-specific and must not be attributed to stock llama.cpp or to any one bundled patch. · upstream
- con-0003 · ggml-org/llama.cpp pr — open Open llama.cpp DFlash2 support PR for Qwen3.8-27B-style drafters. Community comments currently report hardware-dependent results: modest single-slot Strix Halo Vulkan gains over MTP on code but near-parity on prose, a severe Intel B70 multi-agent collapse, a V100 multimodal/M-RoPE draft-context failure with a proposed fix, RTX 3090 near-parity to modest gains, and Blackwell scaling better at higher parallelism. This is upstream/community context only; no HaloBench measured-here result depends on it yet. · upstream
- con-0011 · ggml-org/llama.cpp pr — open llama.cpp PR #27210 by stew675 "spec : add adaptive MTP draft depth (draft-mtp-adaptive)" — adds a new --spec-type draft-mtp-adaptive with a counting state machine (climb counter + weighted drop-pressure accumulator) so draft depth adjusts per segment instead of staying at a fixed n_max. Suggested config --spec-draft-n-max 12; floor/cold-start default 3. Author's own table (Qwen3.8-27B Q8_0, 2x Radeon AI PRO R9700/gfx1201 ROCm, temp 0.6, ctx 8192): on coding, adaptive (n_max 12, floor 2) reaches 86.4 t/s vs fixed depth 3 at 78.8 t/s, driven by long mean draft length (8.0) with 58.4% accept; on prose and hard prose the adaptive rows sit roughly AT or slightly BELOW fixed depth 3 (54.4 vs 56.4 and 51.0 vs 52.6), which is the small reasoning/prose penalty trace. KEY FINDING (comment 5382662863, 2026-08-22T21:21:33Z): stew675 modified the algorithm to stay at FIXED depth 3 but kept max depth 10, and found that merely HAVING max depth 10 incurred a fixed ~2.6% performance penalty completely independent of the adaptive logic — i.e. the penalty is the cost of a high fixed n_max CEILING, not the adaptive switching itself. NEW FIX pushed 2026-08-23T01:32:47Z (comment 5383599943): limits the full MTP buffer scan on truncated/short drafts, giving a ~2-3% speed boost for --spec-draft-n-max 10 --spec-draft-p-min >0.5 configs. INDEPENDENT CORROBORATION (comments 5375448677 at 2026-08-21T21:09:53Z and 5377578075 at 2026-08-22T03:18:20Z, marcusds on a single RTX 5090 32GB, CUDA 13.3, same e1a5754 commit and script): GGML_CUDA_DISABLE_GRAPHS=1 vs default deltas are small (baseline C0 -3.0% / -1.9% / -1.8% / -1.7%), i.e. depth CHANGING does not slow CUDA graph perf; the adaptive 3..10 + p-min config (C3) was the only row with a POSITIVE recall delta (+2.6%) — corroborating that a fixed high ceiling, not adaptive switching, is the cost. Upstream PR record only: surface is CUDA/gfx1201, not our single-GPU ROCm/Vulkan Strix Halo gfx1151 aihydra box, so no HaloBench benchmark claim is made from these figures. · upstream
- con-0001 · ggml-org/llama.cpp fork — carrying Disk slot save/restore silently loses all prompt reuse on hybrid/recurrent models because context checkpoints are never persisted. Fixed with a .ckpt sidecar and pushed to headbouyJB/llama.cpp@fix-25913. Independently confirmed working by two community testers. A third-party PR (#26004) fixes the same bug by appending checkpoints inside the save file instead; production carries our fork until one of the two lands upstream. · upstream
carrying marks a fork in production use but not merged upstream — a reproducibility hazard disclosed on every configuration that depends on it, not a badge of honour