Home › Evidence › Records › con-0011
⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.

con-0011

open
citable URL: https://halobench.com/records/con-0011/ — this address never moves; the anchor /records/#con-0011 keeps resolving
kind pr · upstream ggml-org/llama.cpp
opened 2026-08-17 · https://github.com/ggml-org/llama.cpp/pull/27210
llama.cpp PR #27210 by stew675 "spec : add adaptive MTP draft depth (draft-mtp-adaptive)" — adds a new --spec-type draft-mtp-adaptive with a counting state machine (climb counter + weighted drop-pressure accumulator) so draft depth adjusts per segment instead of staying at a fixed n_max. Suggested config --spec-draft-n-max 12; floor/cold-start default 3. Author's own table (Qwen3.8-27B Q8_0, 2x Radeon AI PRO R9700/gfx1201 ROCm, temp 0.6, ctx 8192): on coding, adaptive (n_max 12, floor 2) reaches 86.4 t/s vs fixed depth 3 at 78.8 t/s, driven by long mean draft length (8.0) with 58.4% accept; on prose and hard prose the adaptive rows sit roughly AT or slightly BELOW fixed depth 3 (54.4 vs 56.4 and 51.0 vs 52.6), which is the small reasoning/prose penalty trace. KEY FINDING (comment 5382662863, 2026-08-22T21:21:33Z): stew675 modified the algorithm to stay at FIXED depth 3 but kept max depth 10, and found that merely HAVING max depth 10 incurred a fixed ~2.6% performance penalty completely independent of the adaptive logic — i.e. the penalty is the cost of a high fixed n_max CEILING, not the adaptive switching itself. NEW FIX pushed 2026-08-23T01:32:47Z (comment 5383599943): limits the full MTP buffer scan on truncated/short drafts, giving a ~2-3% speed boost for --spec-draft-n-max 10 --spec-draft-p-min >0.5 configs. INDEPENDENT CORROBORATION (comments 5375448677 at 2026-08-21T21:09:53Z and 5377578075 at 2026-08-22T03:18:20Z, marcusds on a single RTX 5090 32GB, CUDA 13.3, same e1a5754 commit and script): GGML_CUDA_DISABLE_GRAPHS=1 vs default deltas are small (baseline C0 -3.0% / -1.9% / -1.8% / -1.7%), i.e. depth CHANGING does not slow CUDA graph perf; the adaptive 3..10 + p-min config (C3) was the only row with a POSITIVE recall delta (+2.6%) — corroborating that a fixed high ceiling, not adaptive switching, is the cost. Upstream PR record only: surface is CUDA/gfx1201, not our single-GPU ROCm/Vulkan Strix Halo gfx1151 aihydra box, so no HaloBench benchmark claim is made from these figures.

Cited by — computed at build time, never stored

claims clm-0115
candidate gate history qwen38-27b
docs benchmark-protocol