⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.
clm-0060
measured-heremed ●●○
citable URL: https://halobench.com/records/clm-0060/ — this address never moves; the anchor /records/#clm-0060 keeps resolving
DeepSeek-V4-Flash-0731's KV cache costs approximately 7.13 KiB per token on llama.cpp — about a ninth of Qwen3.8-27B's 64.00 KiB/token hybrid-attention cache (clm-0056) — measured from a two-point GTT delta at --parallel 1, f16 KV: 884 MiB between a c=4096 and a c=131072 load probe (99,464 vs 100,348 MiB gtt_used), over 126,976 additional tokens. The fixed (weights + compute-graph) component this implies, 97.10 GiB, matches the GGUF's own reported model_size (104,202,502,492 bytes = 97.06 GiB) almost exactly, which cross-checks the KV-delta method. THE FIT CONSEQUENCE IS LARGE: at this model's full declared 1,048,576-token native context, total GTT usage projects to ~104.1 GiB (fixed 97.1 GiB + ~7.1 GiB KV) — comfortably inside the 120 GiB (122,880 MiB) GTT boot window with ~16 GiB to spare. This model is nowhere near memory-bound on this hardware at ANY context it natively supports; the wall-clock cost of prefilling that far (clm-0059's decode curve implies hours) is the real ceiling, not fit.
verified 2026-08-15 · volatility low
Note — the record's own working
METHOD — two llama-server load-only probes on aihydra, stock ROCm 3653e6d, UD-IQ3_XXS, f16 KV, -ngl 999, -fa on, --load-mode none, --parallel 1 (explicit — the house hazard-class flag; NOT how the 2026-08-15 screen's own c=32768 fit probe was run, which omitted --parallel and defaulted to n_slots=4/kv_unified — so that number, gtt_used=100,422 MiB at c=32768, is NOT directly comparable to this pair and is not used in the arithmetic here). gtt_used read from /sys/class/drm/card0/device/mem_info_gtt_used before/after each load, server torn down between probes. Raw values: c=4096 -> 99,464 MiB (delta 99,446 MiB from an ~18 MiB idle baseline); c=131,072 -> 100,348 MiB (delta 100,330 MiB). Not captured as formal run/config records — a load probe produces no throughput or capability metric in this schema's sense, the same treatment clm-0056's phase-A footprint probes received.
ARITHMETIC — (100,348 - 99,464) MiB / (131,072 - 4,096) tokens = 884 MiB / 126,976 tokens = 905,216 KiB / 126,976 tokens = 7.1296 KiB/token. Fixed component at c=4096: 99,464 MiB - (4,096 * 7.1296 KiB) = 99,464 - 28.51 MiB = 99,435.5 MiB = 97.10 GiB. Cross-check against the GGUF's declared model_size (llama-bench JSON): 104,202,502,492 bytes = 97.06 GiB — agrees to within 0.04 GiB, i.e. the "fixed" component IS essentially just the resident weights; this quant's compute-graph/KV overhead at low context is near zero, consistent with an MLA-style (multi-head latent attention) compressed cache rather than a conventional per-head KV allocation.
FIT PROJECTION — GTT total 128,849,018,880 bytes = 122,880.9 MiB. At the model's declared 1,048,576-token native context: 99,435.5 MiB (fixed) + 1,048,576 tokens * 7.1296 KiB/token / 1024 = 99,435.5 + 7,301.4 = 106,736.9 MiB = 104.24 GiB, leaving ~16.05 GiB headroom against the 122,880.9 MiB boot budget. Solving the same equation for the ceiling depth the GTT budget alone would permit (ignoring a safety margin) gives ~3.4M tokens — more than 3x the model's own native context — so the GTT window was never going to be the binding constraint for this model on this hardware; native context is. This directly informed the matrix depth selection (clm-0059): depths were chosen for wall-clock usefulness, not because anything deeper risked OOM.