Home › Evidence › Records › clm-0075
⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.

clm-0075

measured-herehigh ●●●
citable URL: https://halobench.com/records/clm-0075/ — this address never moves; the anchor /records/#clm-0075 keeps resolving

qwen36-27b-mtp's KV cache costs 64.00 KiB per token on llama.cpp at f16 (explicit KV, ROCm), measured from a two-point GTT load probe (c=4096 -> 16,821,088,256 bytes gtt_used, c=131,072 -> 25,142,587,392 bytes gtt_used; delta 8,321,499,136 bytes over 126,976 tokens = exactly 64.00 KiB/token). This is LARGER than every other hybrid-architecture full-bench candidate on this board (ornith-35b's 20.00 KiB/token, nemotron3-super's 8.00, deepseek-v4-flash's 7.13), consistent with this candidate carrying MORE full-attention layers than any of them: 16 of 64 blocks (25%, full_attention_interval=4), each with head_count_kv=4 (double ornith-35b's 2), key/value length 256. At the matrix's deepest tested depth (c=131,072), projected GTT use is ~23.4 GiB against the 120 GiB (122,880 MiB) boot window — comfortably inside, with ~97 GiB headroom to spare; fit was never this candidate's constraint.

verified 2026-08-16 · volatility low

Note — the record's own working

METHOD: two-point GTT probe on stock 3653e6d ROCm, qwen36-27b-mtp Q4_K_M, --load-mode none, explicit f16 KV (-ctk f16 -ctv f16), gtt_used sampled from /sys/class/drm/card0/device/mem_info_gtt_used 10s after the server reported healthy (settling convention), server torn down between points. RAW (kvprobe-c4096.log / kvprobe-c131072.log, aihydra ~/bench-results/qwen36-27b-mtp-fullbench/): c=4096 -> gtt_used_after_settle=16,821,088,256 bytes (15.6621 GiB); c=131,072 -> gtt_used_after_settle=25,142,587,392 bytes (23.4159 GiB); delta 8,321,499,136 bytes = 8,126,460.875 KiB / 126,976 tok = 64.0006 KiB/token, i.e. exactly 64 KiB/token to four significant figures. FIRST-PRINCIPLES RECONCILIATION, TWO WAYS. (1) Fixed component: the c=4096 reading (15.6621 GiB) sits within 3.8 MB of the GGUF's own on-disk file size (16,817,244,384 bytes = 15.6584 GiB, llama-bench's own model_size field) — the tightest base-weight agreement recorded in this programme to date (ornith-35b's equivalent check agreed to within 0.040 GiB; this one to within 0.0037 GiB), consistent with this candidate having very little fixed compute-graph or SSM-state overhead relative to raw weight size. (2) Per-token slope from architecture math: KV per token per full-attention layer at f16 = head_count_kv(4) * (key_length(256) + value_length(256)) * 2 bytes = 4,096 bytes; x16 full-attention layers (qwen35.full_attention_interval=4 = 16 of 64 blocks, clm-0073) = 65,536 bytes/token = 64.00 KiB/token — matches the measured GTT delta EXACTLY, not just to an order of magnitude. Not independently re-probed a second time this session (unlike ornith-35b's bit-for-bit reproduction across two probe pairs) — confidence is rated high on the strength of the double first-principles reconciliation (file-size match AND architecture-math match) rather than repro alone; a second probe pair would raise it further but was not run given this candidate's much larger time cost elsewhere in the session (tau2 alone ran 11.3h).

Cited by — computed at build time, never stored

model pages qwen36-27b-mtp