Home › Evidence › Records › clm-0106
⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.

clm-0106

communitymed ●●○
citable URL: https://halobench.com/records/clm-0106/ — this address never moves; the anchor /records/#clm-0106 keeps resolving

The quantized-target divergence in the llama.cpp MTP/DSpark speculative-verify path corroborates across a NEW model, OS and GPU family: on NVIDIA Vulkan/Windows 11 (llama.cpp b10566 / bb4caa754, GTX 1660 SUPER) the official LiquidAI LFM2.5-2.6B pair diverges from vanilla on a Q8_0 target, while the SAME Q8_0 draft against an F16 target produces byte-identical greedy output to the F16 vanilla target (40/64 draft tokens accepted). The divergence is a property of the quantized-target x speculative-verify interaction, not the drafter: with the Q8_0 target the DSpark output consistently and deterministically differs from vanilla, but loading the draft model with --spec-draft-p-min 1 (zero draft tokens) restores byte-exact match to vanilla Q8_0, and the Q8_0 mismatch persists with --spec-draft-n-max 1 and target KV held in F16 (so it is neither a larger speculative block nor quantized target KV). It also diverges at n_max=1, unlike the Qwen3 boundary earlier in #25618 where n_max=1 was reported lossless.

verified 2026-08-22 · volatility high
evidence con-0008

Note — the record's own working

Upstream report only, the fifth arm of the #25618 MTP-correctness dossier (with con-0004/clm-0094, con-0005/clm-0102, con-0006/clm-0104, con-0007/clm-0105). This is INDEPENDENT corroboration of the quantized-target-divergence axis, deliberately on a non-Qwen3.8, non-ROCm/AMD family: LiquidAI LFM2.5-2.6B, NVIDIA GTX 1660 SUPER, Windows 11, Vulkan, Q8_0 target (the original report was Qwen3.8 Q6_K/mixed Q4 and ROCm/RDNA4/Vulkan/Metal). The strong differential control (same Q8_0 drafter, only the target quant differs, F16 target byte-identical to vanilla) isolates target quantization as the causal variable and broadens the house warning that the drift is on the target/verify side, not the drafter. Reported surface is NVIDIA Vulkan/Windows, NOT our ROCm/Vulkan Strix Halo gfx1151 aihydra box, so none of these figures are a HaloBench benchmark gain. It does NOT change the house varied-prompt/raw-token MTP-invariance requirement -- it STRENGTHENS it: quantized-target drift is now evidenced on multiple unrelated families, so a local varied-prompt / raw-token control remains required before any baseline/MTP invariance claim on our surface, and the target quantization axis must be treated as a live loss source regardless of drafter or model family. The n_max=1 divergence on LFM2 (vs lossless on Qwen3.8 at n_max=1) also warns that the narrow n_max=1 boundary is not a general safety property across targets.

Cited by — computed at build time, never stored