con-0006
open
citable URL: https://halobench.com/records/con-0006/ — this address never moves; the anchor /records/#con-0006 keeps resolving
kind report · upstream ggml-org/llama.cpp
opened 2026-08-22 · https://github.com/ggml-org/llama.cpp/issues/25618#issuecomment-5377293972
New llama.cpp #25618 follow-up by snick525 (2026-08-22T02:11:18Z) on RDNA4 (AMD Radeon AI PRO R9700 / gfx1201, 32 GB, Vulkan/radv), target Qwen3.8-27B Q6_K_XL (MTP head intact), greedy (temperature 0, top_k 1, top_p 1, seed 42), f16 K/V. The reporter additionally built open PR #27342 (DFlash2, an external 2B draft model) so the built-in MTP head is no longer the only speculation type in the comparison. Three findings against a --spec-type none baseline on the same binary:
1. DIVERGENCE IS DRAFTER-INDEPENDENT. At -c 65536 / f16 KV the code prompt
diverges at the same first-difference byte 204 for the built-in MTP head
AND for DFlash2 at three drafter quants (Q4_K_M, Q8_0, BF16); all four
configs are lossless on a reasoning prompt. Same first-diff byte holds
across two context sizes (16384 and 65536) and two builds. Whatever drifts
is on the target/verify side; the drafter is not a variable in it.
2. DRAFTER QUANTIZATION HAS NO EFFECT. DFlash2 Q4_K_M vs Q8_0 vs BF16 produce
byte-identical outputs on both prompts and both context sizes, and
identical acceptance counters (draft 347, accepted 283). Consistent with
the verify path emitting only the target's own token: a drafter can only
change output by proposing different tokens, and here even a Q4 drafter
proposes identically, so drafter precision can be ruled out when bisecting.
3. KV-CACHE QUANTIZATION IS ITSELF OUTPUT-MOVING AND COMPOUNDS. Swapping
f16 for q8_0 KV (everything else fixed) turns the reasoning prompt from
LOSSLESS under f16 to DIVERGES at byte 155 under q8_0 for both MTP and
DFlash2, and moves the code prompt's first-diff from byte 204 to byte 277.
Separately, KV quantization alone moves greedy output with no speculation
on either side: --spec-type none with f16 KV versus q8_0 KV differs on the
code prompt at byte 561. The reporter notes q8_0 KV is accordingly never
output-preserving and compounds with whatever this issue turns out to be.
Throughput context (mentioned only because it came up in #27342, and reported with an active caveat): -c 65536 / f16 KV on Qwen3.8-27B Q6_K_XL on a 32 GB card gave 23.8 tok/s no-spec, 40.4 tok/s draft-mtp n_max 1, 61.9 tok/s DFlash2 Q4_K_M, and 29.4 tok/s for DFlash2 Q8_0 (the Q8 and BF16 drafter rows are NOT drafter-cost evidence: at 31.2-31.6 GB they cross the card's usable ceiling and silently fall back to host memory per #26432, with acceptance unchanged). Below the ceiling every drafter quant performs identically, so the Q8 spill is a memory-footprint artefact, not a drafter-precision finding.
Upstream diagnostic record only. Surface is RDNA4 R9700/gfx1201 Vulkan, not our ROCm/Vulkan Strix Halo gfx1151 aihydra box, so no HaloBench benchmark claim is made from these figures.
Cited by — computed at build time, never stored
candidate gate history qwen38-27b