con-0008
open
citable URL: https://halobench.com/records/con-0008/ — this address never moves; the anchor /records/#con-0008 keeps resolving
kind report · upstream ggml-org/llama.cpp
opened 2026-08-22 · https://github.com/ggml-org/llama.cpp/issues/25618#issuecomment-5382455878
llama.cpp #25618 comment #16 by F-Mangini (2026-08-22T20:32:01Z) corroborates the quantized-target divergence on a NEW model/OS/GPU family: the official LiquidAI LFM2.5 DSpark pair also diverges from vanilla on a quantized target while an F16 target preserves greedy parity. Environment: llama.cpp b10566 (bb4caa754), Vulkan build, Windows 11, NVIDIA GTX 1660 SUPER 6 GB (driver 551.76). Target: LiquidAI/LFM2.5-2.6B-GGUF Q8_0 (SHA-256 36587fdf27bdfc69caf2637273679a0870ec155162161bde6fd16e8c70bdb757); draft: LiquidAI/LFM2.5-2.6B-DSpark-GGUF Q8_0 (SHA-256 85a98fafd9af1328b6876fd1360d7ed69e74c6cefc14dd07fb6306e1940386c87). Minimal raw /completion reproduction without a chat template, server `-c 4096 -np 1 -fa on -ngl 99 -ctk f16 -ctv f16`, requests temperature=0 seed=42 n_predict=256, on the LRU-cache Python prompt. Quantized-target finding: vanilla Q8_0 target outputs SHA-256 2eeb8174..., same target + DSpark --spec-draft-n-max 1 outputs a8d38c69... (draft accepted 107/147) -- DIVERGES. Differential controls: (1) Q8_0 target + DSpark loaded but --spec-draft-p-min 1 (zero draft tokens) exactly matches vanilla Q8_0; (2) the active Q8_0 DSpark run is deterministic but consistently differs from vanilla; (3) the SAME Q8_0 draft against an F16 target produces byte-identical greedy output to the F16 vanilla target while accepting 40/64 proposed tokens over 64 generated tokens -- i.e. F16 target preserves greedy parity. The Q8_0 mismatch also persists with --spec-draft-n-max 1 and target KV in F16, so it is not caused by a larger speculative block or by quantized target KV. Notably the n_max=1 divergence differs from the Qwen3 boundary reported earlier in the same issue, where n_max=1 was said to remain lossless. Reporter concludes this is consistent with path-dependent target numerics between sequential single-token decoding and the speculative verification path for quantized weights, and confirms the issue affects LFM2 DSpark, NVIDIA Vulkan, Windows and Q8_0 -- not only the Q4/MTP configs in the original report. Upstream diagnostic record only: surface is NVIDIA Vulkan/Windows, not our ROCm/Vulkan Strix Halo gfx1151 aihydra box, so no HaloBench benchmark claim is made from these figures.
Cited by — computed at build time, never stored
claims clm-0106