Home › Evidence › Records › clm-0055
⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.

clm-0055

measured-heremed ●●○
citable URL: https://halobench.com/records/clm-0055/ — this address never moves; the anchor /records/#clm-0055 keeps resolving

draft-mtp speculation on Qwen3.8-27B (stock 3653e6d, gfx1151) peaks at spec-draft-n-max=3 over CLEAN cells — Vulkan Q8_0 17.75 t/s (2.26x its 7.86 no-speculation floor, acceptance 0.626), Vulkan UD-Q4_K_XL 26.68 t/s (2.23x, 0.623), ROCm Q8_0 18.25 t/s (2.33x, 0.6235) — and at n_max >= 4 the feature is BROKEN on this model: after accumulated generation volume in a live session (sequential VARIED prompts at full length; not fresh servers, not one repeated prompt, not short generations), generations start terminating at 1 token with <|im_end|> (id 248046). The decisive measurement is that the TARGET model's own pre-sampling distribution puts im_end at logprob -0.084 (~92%) on a prompt the same server answers normally with speculation off — the speculative path is corrupting the target's forward pass, not merely mis-accepting drafts — and ignore_eos:true restores both correct text AND draft accounting. The hazard compounds into a measurement artifact: llama.cpp reports a 1-token generation as 1,000,000 tokens/s, so unfiltered throughput averages FLATTER exactly the broken cells; a community sweep on this silicon reporting a peak at n_max=5 sits on what is measured here as one of the two worst cells (11/15 degenerate on Q8_0), under a workload shape (single repeated codegen prompt) that this trigger analysis shows cannot reproduce the failure. Recommendation from measurement: draft-mtp ON, n_max hard-capped at 3 until the mechanism is understood upstream.

verified 2026-08-15 · volatility medium
evidence run-0267 run-0268 run-0269 run-0270 run-0275 run-0276 run-0277 run-0278 run-0279 run-0280 run-0281 run-0282 run-0283 run-0284 run-0285 run-0286
corrections 2026-08-15: The headline text's "at n_max >= 4 the feature is BROKEN on this model" carries no backend qualifier and could be read as a fleet-wide (any backend) claim. A same-day cause-isolation matrix (see the second 2026-08-15 amendment below) falsifies that broad reading: flipping ONLY the backend from the failing Vulkan/Q8_0/f16-KV baseline to ROCm, with quant/KV/n_max held at n_max=5 (one of the two historically worst cells), produced ZERO anomalies of any kind across 15 requests AND an extended 45-request confirmation (run-0278, run-0281) -- where the Vulkan baseline reproduced at 11/15 degenerate in the same session shape (run-0277). "BROKEN on this model" should be read as "BROKEN on this model on Vulkan, at this session depth" -- the ROCm result is real evidence, not an absence of evidence, and the standing n_max<=3 hard-cap recommendation is unchanged (it was never contingent on backend), but a reader taking the original sentence as backend-general would be wrong. This also retires the immediately-prior amendment's open question ("a matched ROCm + f16-KV + n_max=7 control... has not been run") -- run-0278/run-0281 is that control, run at n_max=5 instead of 7, which is the more diagnostic choice since it is a known-worse cell on Vulkan.

Note — the record's own working

METHOD (sweep) — llama-server harness (llama-bench has no speculative support; spec params are force-set to 0 there), 5 varied prompts x 3 reps = 15 generations per cell, greedy (temperature 0, seed 42), n_predict 500, cache_prompt off, -c 32768, --load-mode none, stock 3653e6d, IOMMU on. The floor (spec off) travels in-session with every cell; the RATIO is the deliverable, absolutes are not comparable to llama-bench numbers. "degenerate" = predicted_n <= 1, and the per-cell degenerate fraction is part of the record: Vulkan sweep, decode t/s median of clean samples (degenerate/attempted): | n_max | Q8_0 | degen | UD-Q4_K_XL | degen | |---|---|---|---|---| | off | 7.86 | 0/15 | 11.98 | 0/15 | | 2 | 16.50 | 0/15 | 23.74 | 0/15 | | 3 | 17.75 | 0/15 | 26.68 | 0/15 | | 4 | (artifact) | 12/15 | 25.55 | 5/15 | | 5 | (artifact) | 11/15 | 24.73 | 7/15 | | 6 | 34.41 | 3/15 | 43.72 | 2/15 | The Q8_0 n_max=4/5 "medians" computed naively are 1,000,000 t/s — the 1-token sentinel value, not a throughput; that is the artifact hazard in one line. The n_max=6 cells are faster on their clean samples and markedly LESS degenerate than 4/5 — unexplained, and left as an open question rather than a recommendation. The clean-cell peak is n_max=3 on every arm measured. The n_max >= 4 cells are deliberately NOT recorded as run records; the clean cells behind this claim are (run-0268 ROCm Q8_0, run-0269 Vulkan Q8_0, run-0270 Vulkan UD-Q4_K_XL). THE EOS CLIFF, characterised (full evidence bundle: ~/bench-results/qwen38-screen/eos-cliff-evidence.md on the bench box, prepared for upstream review): - Trigger needs accumulated generation volume in one server session: fresh server per request 0/6 degenerate; one repeated prompt sequentially 0/10; varied prompts at n_predict 48 (~480 tokens total) 0/10; varied prompts at n_predict 500 hit the cliff by request 2-3 (~1,400 tokens) — 11/15 degenerate at n_max=5. Recovery is possible mid-session (a later request produced a clean 500), so the state is not latched. - Affected requests show tokens_cached == tokens_evaluated (the KV slot WAS reset — not stale prompt cache) and draft_n/draft_n_accepted ABSENT entirely, vs 666-847 drafted on healthy requests in the same session. - Not a sampler artifact: reproduces under greedy (seeds 42 and 1234, bit-identical sequences) and under the vendor thinking sampler (temp 1.0, top_p 0.95, top_k 20, min_p 0) at seeds 42 and 7 with seed-varying rates. Bit-identical greedy failure sequences across four independent sessions, across server restarts and drop_caches. - Worse on Q8_0 (12/15 at n_max=4) than UD-Q4_K_XL (5/15) — the OPPOSITE of what a quantisation-noise story predicts. - ignore_eos:true control: same server, same n_max, same prompt — correct text returns AND draft accounting reappears (draft_n 341, accepted 129). The corruption expresses entirely through the EOG token. - CHECKED against the nearest upstream issues; none matches this signature: 23302/23335 are draft-mtp token DIVERGENCE with complete generations (Metal); 25618 is quant-dependent divergence (and the Q8-worse observation here cuts against it); 26750 is a CUDA acceptance collapse — HIP acceptance here (0.6235) is statistically indistinguishable from Vulkan (0.6261), so it does not generalise to this box. CONFIDENCE NOTES — what is and is not established. Established by measurement: the trigger conditions, the n_max threshold, the pre-sampling distribution shift onto im_end, the ignore_eos control, backend-indifferent acceptance at n_max=3, and determinism. NOT established: the root cause (the MTP head's internal state is the natural suspect but was not instrumented), why a single repeated prompt escapes, and why n_max=6 partially recovers. ROCm at n_max >= 4 was not probed for the symptom (the clean ROCm n_max=3 cell is not evidence either way) — see the 2026-08-15 amendment below, which is the first ROCm probe at n_max >= 4. UPSTREAM: not yet filed — the evidence bundle is written for review first; this claim should be re-verified against whatever the eventual issue thread establishes. Confidence is medium on the characterisation and the n_max<=3 operating rule; the mechanism language above is deliberately interpretation-free. COMMUNITY CONFLICT, stated fairly: the r/StrixHalo sweep reporting draft-mtp peaking at n_max=5 for this model family ran a different build lineage (the strix-halo-vulkan fork, b10283/b10397 era), amd_iommu=off, and a single repeated 500-token codegen prompt — the exact workload shape measured here as unable to trigger the cliff — and llama-bench-style unfiltered averages cannot distinguish a fast cell from a broken one once 1-token generations enter the mean. Their numbers may be correct for their harness; the recommendation that follows from them is what this claim disputes for live serving. Guard evidence for the recommended operating point: 25 sequential full-length vendor-sampler generations at n_max=3 on the tau2 serving config, 0/25 degenerate, ~11,000 generated tokens in one session (run-0267) — deeper session volume than any failing cell needed. ENERGY (retroactive join, computed 2026-08-15 against the same nmaxoff/nmax3 sub-windows behind the sweep table above; eng-0077/0078/0079/0080). MTP's decode speedup does NOT translate 1:1 into energy savings, because mean power draw rises under speculation rather than holding constant — the extra draft/verify compute costs watts, not just wall-clock: | arm | floor Wh/1000tok | n_max=3 Wh/1000tok | energy improvement | decode speedup | mean W floor -> n3 | |---|---|---|---|---|---| | Vulkan Q8_0 | 5.03 | 2.42 | 2.08x | 2.26x | 142.4 -> 154.7 (+8.6%) | | Vulkan UD-Q4_K_XL | 3.41 | 1.88 | 1.82x | 2.23x | 147.2 -> 180.5 (+22.6%) | | ROCm Q8_0 | 5.21 | 2.70 | 1.93x | 2.33x | 146.8 -> 177.6 (+21.0%) | Every arm still burns meaningfully less energy per output token with speculation on — this is not a wash — but the win undershoots the raw decode ratio by 8-23 percentage points depending on backend/quant, entirely because power draw is not flat. UD-Q4_K_XL shows the largest power increase and therefore the smallest energy win despite a comparable decode speedup, which is the opposite of what "MTP is free, only wall-clock changes" would predict. Wh/1000tok = wh_total / (decode_tps * window_seconds) — an approximation that assumes the window is generation-dominated (verified true: prompts are ~20 tokens, prefill negligible against 500-token generations). Recommendation unchanged (n_max=3, hard-capped) — the energy case reinforces rather than revises it — but "2.3x faster" should not be read as "2.3x less energy"; it is closer to 1.8-2.1x less energy, per this measurement. AMENDMENT 2026-08-15 — EOS-CLIFF x q8_0 KV, and the first ROCm probe at n_max >= 4. Prompted by a r/Qwen_AI thread (1vorjo7) recommending n_max=7 + q8_0 KV for this model, and by upstream #25618 (a q8_0 V-cache + MTP divergence interaction on a DIFFERENT model) — two open questions this claim had not addressed: does the cliff reproduce under quantised KV, and is our own recommended combo (n_max=3, q8_0 KV — clm-0057/clm-0058's serving config) free of that interaction. Same 5-prompt-sequence x3-reps harness as the sweep above, greedy seed 42, n_predict 500, cache_prompt off, ROCm, UD-Q4_K_XL weight quant, `-ctk q8_0 -ctv q8_0`: | n_max | KV | degenerate | evidence | |---|---|---|---| | 3 | q8_0 | 0/15 | run-0275 | | 7 | q8_0 | 0/15 | run-0276 | Q1 ANSWERED FOR OUR SERVING COMBO: n_max=3 stays clean under q8_0 KV (run-0275), same as it already was under f16 KV (run-0267's 0/25 at the tau2 serving config). The recommended Warden combination is not exposed to this interaction by this test. Q2 — DOES NOT REPRODUCE, BUT THE TEST IS CONFOUNDED. n_max=7 under q8_0 KV on ROCm produced zero degenerate generations (run-0276) — no im_end-as-first-token signature anywhere in the 15-request sequence, at an n_max deep inside the zone that broke Vulkan/f16 KV (12/15 and 11/15 at n_max=4/5). But this cell changes TWO variables from the original repro at once: backend (ROCm, not Vulkan) AND KV type (q8_0, not f16). It therefore does NOT cleanly show that q8_0 KV suppresses the cliff — it is equally consistent with ROCm simply not exhibiting this symptom at any KV type, which is exactly the gap this claim already flagged ("ROCm at n_max >= 4 was not probed for the symptom"). This is the first data point on that gap, and it is clean, but it does not close it: a matched ROCm + f16-KV + n_max=7 control is the cell that would decompose backend from KV-quant as the protective factor, and it has not been run. PRACTICAL READING: this does not change the standing recommendation (n_max=3, hard-capped) — that recommendation was never contingent on n_max=7 being safe or unsafe, and the clean n_max=3 result under q8_0 KV is confirmatory, not novel. What it does do is narrow the open backend question without closing it, and it gives a direct, sourced answer to the r/Qwen_AI thread's n_max=7+q8_0 suggestion for THIS box's ROCm build: no cliff observed in 15 requests, but the test that would show whether that is the KV type or the backend doing the protecting has not been run. AMENDMENT 2026-08-15 (SECOND) — CAUSE-ISOLATION MATRIX: backend is the dominant single-variable gate, session accumulation is a real secondary factor, closes the "matched ROCm + f16-KV control has not been run" gap from the amendment above. Full method, raw matrix, and reasoning: eos-cliff-evidence.md section 11 (bench box + local project copy) and ~/bench-results/eos-cliff-isolation-20260815/matrix-summary.md on the box. One variable changed per cell from the failing baseline (Vulkan, Q8_0, f16 KV, n_max=5), 15 requests each (same 5-prompt x3-rep greedy harness as the sweep above), plus 45-request extended confirmation of the two clean results, plus four avoidance probes on the still-failing config: | cell | backend | quant | KV | n_max | n | degenerate | short-stop | |---|---|---|---|---|---|---|---| | 0 baseline (positive control) | Vulkan | Q8_0 | f16 | 5 | 15 | 11/15 | 1/15 | | B backend-only flip | ROCm | Q8_0 | f16 | 5 | 15 | 0/15 | 0/15 | | C KV-only flip | Vulkan | Q8_0 | q8_0 | 5 | 15 | 0/15 | 1/15 | | E new-evidence combo at n5 | ROCm | UD-Q4_K_XL | q8_0 | 5 | 15 | 0/15 | 3/15 | | B-ext (45 req) | ROCm | Q8_0 | f16 | 5 | 45 | 0/45 | 0/45 | | N7-ext (45 req) | ROCm | UD-Q4_K_XL | q8_0 | 7 | 45 | 0/45 | 9/45 | | avoid-1 n_max=3 | Vulkan | Q8_0 | f16 | 3 | 15 | 0/15 | 0/15 | | avoid-2 p-min 0.9 | Vulkan | Q8_0 | f16 | 5 | 15 | 1/15 | 0/15 | | avoid-3 cache_prompt on | Vulkan | Q8_0 | f16 | 5 | 15 | 12/15 | 1/15 | | avoid-4 restart every 3 | Vulkan | Q8_0 | f16 | 5 | 15 (5 sessions) | 1/15 | 3/15 | "degenerate"=predicted_n<=1 (this claim's established definition); "short-stop" = any other premature stop, the same "intermediate case" already present in the original characterisation's minimal repro (394/500). Evidence: run-0277 (cell 0), run-0278 (cell B), run-0279 (cell C), run-0280 (cell E), run-0281 (B-ext), run-0282 (N7-ext), run-0283 (avoid-1), run-0284 (avoid-2), run-0285 (avoid-3), run-0286 (avoid-4). LOCALISATION: backend (ROCm vs Vulkan) is the dominant single-variable gate at this session depth. Flipping it alone took 11/15 degenerate to ZERO anomalies of any kind at 3x the request volume (B-ext). KV-only and quant-only flips both leave the model still perturbed (KV-only: 1 short-stop; quant-only, already in this claim's own sweep table: 7/15 degenerate). The "new evidence" triple-flip combination, re-tested at the actually-diagnostic n_max=5 instead of the untested-in-the-failure-zone n_max 3/7, shows a HIGHER short-stop rate (3/15) than the KV-only flip — combining the backend flip with quant+KV changes does not add protection over the backend flip alone, and may reintroduce a partial echo of the mechanism. This is a mitigation gradient, not a clean/broken binary: even N7-ext (fully clean of true collapse) shows a small, perfectly periodic short-stop tied to one prompt position, recurring every cycle, never escalating with volume. SESSION ACCUMULATION IS REAL BUT PARTIAL. Restarting the server every 3 requests on the FAILING Vulkan config cut the degenerate rate 11/15 -> 1/15 — the single largest mitigation measured — but one fresh 3-request sub-session still produced a full collapse on its 3rd request; a targeted pre-sampling replay of that exact case (post_sampling_probs:false, this claim's decisive- measurement method) shows im_end at logprob -0.148 (~86.2%) vs the #2 candidate at -2.115 (~12.1%) — as strong a shift as the original 92% characterisation. Raising --spec-draft-p-min to 0.9 shows the same profile (11/15 -> 1/15, not zero); when it fires, im_end sits at -0.391 (~67.6%) vs -1.175 (~30.9%) — weaker but the same direction. cache_prompt:true is NOT a mitigation — it made the baseline WORSE (12/15 vs 11/15). REPETITION-LOOP CHECK (motivated by llama.cpp #26425 "MTP retains inter-request state... model degradation" and #23577 "Qwen3.6-27B MTP outputs repeated //// after long session" — near-miss reports sharing the accumulated-session-state trigger class but a different reported symptom). All 26 full-raw JSON captures from this matrix were grepped for repeated- character (8+) and repeated-short-token (7+) runs. Zero genuine hits — every match was a legitimate markdown separator inside a correct, complete answer. This box's failure signature stays exclusively in the EOS-collapse family; it never drifts into a repetition-attractor mode under any configuration tested. AVOIDANCE ENVELOPE: --spec-draft-n-max<=3 remains safe regardless of backend (avoid-1 reconfirms 0/15 on the original failing Vulkan combo) and is still the only zero-anomaly result achieved on Vulkan without changing backend, quant, or KV. ROCm held clean at 3x this matrix's depth (45 req / ~22,500 generated tokens) at n_max=5 — the strongest result in the matrix — but this scope is 15-45 requests, NOT production-scale session depths; whether it holds at hundreds-to-thousands of requests is not established and should not be assumed. On Vulkan at n_max>=4-5, no cheap mitigation measured here reaches zero: restart-every-3 and p-min=0.9 both cut the rate ~10x without eliminating it; cache_prompt:true is counterproductive. WHAT THIS DOES NOT ESTABLISH: the mechanism itself is still uninstrumented — this matrix localises WHICH variable gates the symptom, not WHY. "ROCm doesn't show it here" is not the same claim as "ROCm is immune," and this claim should not be read as making the stronger one. Standing recommendation UNCHANGED: n_max=3, hard-capped, on any backend — that recommendation was never contingent on this result and remains the only zero-anomaly, any-backend, any-depth-tested operating point measured to date.

Cited by — computed at build time, never stored

model pages qwen38-27b
candidate gate history qwen38-27b