Home › Evidence › Records › con-0015
⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.

con-0015

open
citable URL: https://halobench.com/records/con-0015/ — this address never moves; the anchor /records/#con-0015 keeps resolving
kind issue · upstream ggml-org/llama.cpp
opened 2026-08-24 · https://github.com/ggml-org/llama.cpp/issues/25618#issuecomment-5392171333
Raw upstream diagnostic comment by frizikk on open issue #25618, retrieved directly from GitHub on 2026-08-25. Surface: pinned llama.cpp sources, Vulkan, Qwen3.8-27B Q6_K target plus Q8_0 native MTP, flash attention on, f16 K/V, greedy sampling, and --spec-draft-n-max 1. Its first confirmed target-verify boundary is the natural generation-1 layer-0 Q8_0 x F32 MUL_MAT: sequential N=1 uses F32 MUL_MAT_VEC, while speculative N=2 stages activations as Q8_1 for MMVQ; the direct predecessor is exact but every one of the 5120 output values changes. Replaying the N=2 fixture through the non-MMVQ F32 path made row 0 bit-exact, but later full generation logits still differ, so the narrow diagnostic gate is not a complete fix. The same comment identifies a second N-dependent ADD/RMS-reduction boundary and explicitly confines its flash-attention observation to one captured fixture. This is upstream operator-level evidence, not a HaloBench benchmark or a production fix. It corroborates the standing varied-prompt/raw-token MTP-invariance requirement, but does not identify any local Qwen3.8 result as affected: the closest recorded local Vulkan arms use Q8_0 targets and n_max=3 (cfg-0068/run-0269 and cfg-0147/run-0573), already retain a negative varied-prompt invariance result and are not promoted as baseline/MTP equivalent. It has no direct transfer to the local ROCm Qwen3.8 records.