Home › Evidence › Records › clm-0062
⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.

clm-0062

measured-heremed ●●○
citable URL: https://halobench.com/records/clm-0062/ — this address never moves; the anchor /records/#clm-0062 keeps resolving

Nemotron-3-Super-120B-A12B's KV cache costs 8.00 KiB per token at UD-Q4_K_M on llama.cpp — measured byte-exact from a two-point GTT delta, ROCm, f16 KV, --load-mode none, --parallel 1: 224 MiB between a c=4096 and a c=32768 load probe (78,754 vs 78,978 MiB gtt_used), over 28,672 additional tokens (229,376 KiB / 28,672 tok = 8.00 KiB/token exactly). The fixed (weights + compute-graph) component this implies, 76.88 GiB (78,722 MiB), matches the GGUF's own reported model_size (82,533,249,024 bytes = 76.87 GiB) to within 0.01 GiB — a clean cross-check of the method, same discipline as clm-0060's DeepSeek-V4-Flash measurement. THIS DIRECTLY SUPERSEDES THE EARLIER "cannot allocate past d0" READING for this candidate: perf-matrix-gtt120's d32768 cell (2026-08-14) recorded a BadAlloc storm under mmap and inferred d65536/d131072 as OOM without running them — that was clm-0053's mmap double-residency defect (llama-bench's four throughput queue scripts lacked --load-mode none), not a memory-footprint fact about this model. Under the fix, this job's own throughput matrix (run-0308..run-0315) completed d0 and d32768 cleanly on BOTH backends, 3/3 fresh-process reps each, and this KV-cost measurement shows why: at 8.00 KiB/token the model is nowhere near the 122,880 MiB GTT boot budget at any depth this bench tested — headroom past d32768's 78,978 MiB is ~43.9 GiB, enough for several million more tokens by the KV-cost arithmetic alone (before compute-graph/batch-buffer growth at depth is accounted for, which this two-point measurement cannot isolate). The real ceiling for this candidate, like DeepSeek-V4-Flash (clm-0060), is very unlikely to be GTT fit — it is wall-clock: at ~16-18 t/s decode with no speculation path (Mamba rejects draft-mtp, candidate history 2026-06-27), a much deeper matrix cell would cost hours per rep for a number this job's scope did not call for.

verified 2026-08-16 · volatility low
evidence run-0308 run-0309 run-0310 run-0311 run-0312 run-0313 run-0314 run-0315

Note — the record's own working

METHOD — two llama-server load-only probes on aihydra, stock ROCm 3653e6d, UD-Q4_K_M, f16 KV, -ngl 999, -fa on, --load-mode none, --parallel 1, gtt_used read from /sys/class/drm/card0/device/mem_info_gtt_used before/after each load, server torn down between probes. Raw values: c=4096 -> 78,754 MiB (delta from a fresh-boot-equivalent stop 78,737 MiB); c=32,768 -> 78,978 MiB (delta 78,961 MiB). Not captured as formal run/config records — a load probe produces no throughput or capability metric in this schema's sense, the same treatment clm-0056 and clm-0060's phase-A/load-probe measurements received. ARITHMETIC — (78,978 - 78,754) MiB / (32,768 - 4,096) tokens = 224 MiB / 28,672 tokens = 229,376 KiB / 28,672 tokens = 8.00 KiB/token exactly. Fixed component at c=4096: 78,754 MiB - (4,096 * 8.00 KiB / 1024) = 78,754 - 32 MiB = 78,722 MiB = 76.88 GiB. Cross-check against llama-bench's own model_size field (verified on this exact GGUF via this job's matrix JSONs): 82,533,249,024 bytes = 76.87 GiB — agrees to within 0.01 GiB. COMPARISON TO PEERS — 8.00 KiB/token sits close to DeepSeek-V4-Flash's 7.13 KiB/token (clm-0060, an MLA-style compressed cache) and far below Qwen3.8-27B's 64.00 KiB/token (clm-0056, hybrid 3:1 linear-to-full-attention layout). Two architecturally very different models (hybrid Mamba-2 + LatentMoE MoE here, MLA MoE there) land within 12% of each other on KV cost per token, both far cheaper than a conventional hybrid-attention allocation — worth noting as a pattern, not yet a claim about why. SUPERSEDES — perf-matrix-gtt120 (2026-08-14): nemotron3-super-rocm-kvf16-faon- d32768-rep1.stderr shows a sustained "HSA exception: BadAlloc" storm (dozens of lines) starting near process start; the d65536/d131072 cells were never run and instead recorded as `oom-inferred` with basis "measured BadAlloc OOM at d32768... at GTT120". clm-0053 root-caused this defect class (mmap double-residency competing with the KFD system-memory gate for host RAM, measured on this exact candidate) and fixed the four throughput queue scripts that lacked --load-mode none. This job's matrix used the fixed queue script throughout (see cfg-0087/cfg-0088's provenance) and both d0 and d32768 completed clean on both backends — the direct confirmation the fix holds for a fresh nemotron3-super matrix, not just the fix's own original test cell.

Cited by — computed at build time, never stored

model pages nemotron3-super
candidate gate history nemotron3-super