⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.
clm-0115
communitymed ●●○
citable URL: https://halobench.com/records/clm-0115/ — this address never moves; the anchor /records/#clm-0115 keeps resolving
A high fixed MTP draft-depth CEILING carries a measurable fixed cost that is INDEPENDENT of the adaptive logic: on llama.cpp PR #27210 (draft-mtp-adaptive), stew675 held the algorithm at fixed depth 3 but kept max depth 10 and found merely having max depth 10 incurred a fixed ~2.6% performance penalty on its own; the subsequent fix (2026-08-23T01:32:47Z, limit the full MTP buffer scan on truncated/short drafts) recovers ~2-3% for --spec-draft-n-max 10 --spec-draft-p-min >0.5 configs. marcusds independently confirms on a single RTX 5090 that depth CHANGING does not degrade CUDA-graph performance (GGML_CUDA_DISABLE_GRAPHS=1 vs default deltas small, baseline C0 -3.0% / -1.9% / -1.8% / -1.7%), and that adaptive 3..10 + p-min is the only config with a POSITIVE recall delta (+2.6%) — so the ~2.6% is the cost of a static high ceiling, not of adaptive switching. The house reading: a deep static or deep blind-sweep n_max is not a safe/default target; a shallow-plus-adaptive draft depth (low floor, bounded ceiling, adaptive climb/drop) is the target.
Draft-depth-cost arm of the MTP-correctness dossier. Reinforces the standing shallow/adaptive draft-depth guidance (protocol §1a/§0 and the house n_max<=3 caution on qwen38-27b, clm-0055) and the "no deep blind sweep" position: a high fixed n_max ceiling carries a fixed cost even when the model never uses it. Reported surface is CUDA/gfx1201 upstream PR context, not our single-GPU ROCm/Vulkan Strix Halo aihydra box, so no HaloBench benchmark claim or n_max guidance change is made from these figures; the shallow/adaptive + varied-prompt requirement stands intact.