⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.
clm-0080
measured-herehigh ●●●
citable URL: https://halobench.com/records/clm-0080/ — this address never moves; the anchor /records/#clm-0080 keeps resolving
Qwen3-Coder-Next 80B's KV cache costs 24 KiB/token (f16 KV), computed from the GGUF's own architecture metadata rather than a live GTT probe: qwen3next. full_attention_interval=4 over 48 blocks means 12 full-attention layers (GQA, attention.head_count_kv=2, attention.key_length=attention.value_length=256), each costing 2 kv_heads x 256 head_dim x 2(K+V) x 2 bytes(f16) = 2,048 bytes/token, and the other 36 blocks are gated-DeltaNet-class SSM/linear- attention with NO growing KV cache at all. This job's own live FIT probe at c=32768 measured 81,662 MiB GTT (delta from empty ~81,644 MiB against the 84,812,055,968-byte / 78.99 GiB Q8_0 weight footprint), consistent with the metadata-derived KV rate to within measurement noise (weights + ~0.75 GiB KV at c=32768 against the observed delta). Projected GTT use at this candidate's own model-max (c=262,144) is ~85 GiB against the 120 GiB boot window -- comfortably inside, ~35 GiB of headroom to spare, never a fit risk at any depth this bench measured.
Unlike ornith-35b's two-point live-probe KV measurement (kvprobe-ornith-redo.sh, two fresh-process FIT-only loads at different contexts), this candidate's KV cost was derived from the GGUF's own metadata (full_attention_interval, head_count_kv, key/value length) and cross-checked against the single FIT probe this job's screen already captured (c=32768, 24s load, 81,662 MiB GTT) -- a live two-point GTT delta probe (e.g. c=4096 vs c=131072) was not run separately this session; the metadata-derived rate is the primary basis for this claim, with the single live point as a consistency check rather than an independent confirmation. Same architecture family and per-layer KV shape (head_dim 256, 2 kv_heads) as ornith-35b (20 KiB/token, 10 full-attention layers) and the production 122B -- this candidate is 48 blocks / 12 full-attention vs ornith's 40 blocks / 10, so the per-token rate scales exactly with the extra 2 full-attention layers (12 x 2048 = 24,576 bytes vs ornith's 10 x 2048 = 20,480 bytes).