Models › nemotron3-super

Nemotron-3-Super-120B-A12Bbenchedguard 4/4

120B / ~12B activehybrid Mamba-2 MoE — no MTP/speculation path in llama.cppquant held: UD-IQ4_XSfirst measured 2026-06-27latest run 2026-08-11

Verdict

A 120-billion-parameter hybrid Mamba-2 mixture-of-experts activating 12B per token — the largest active footprint in the field, with no speculation path in llama.cpp, so it decodes at its floor clm-0005. On τ² reward it is indistinguishable from the 122B; on the tasks both models get right it needs 37% more turns and 85% more wall-clock to reach the same correct answer clm-0039 — on this lab's headline metric the incumbent leads. The quant confound is resolved: on 12 identical seeded tasks, UD-Q4_K_M and UD-IQ4_XS produce identical reward on every task while IQ4_XS takes +11% more total turns clm-0047 — quantisation cost this model efficiency, never correctness. It is the flattest model with depth measured here, holding its decode almost unchanged to 32k clm-0025, which earns it a second look wherever long-context stability matters more than raw speed. Every cross-model score above remains provisional until re-run under the pinned independent simulator clm-0043 clm-0048.

verdict written 2026-08-13 · every number above stands next to the claim chip that carries it

Best configuration

modelNVIDIA-Nemotron-3-Super-120B-A12B-UD-IQ4_XS-00001-of-00003.gguf · UD-IQ4_XS
engineggml-org/llama.cpp 3653e6d · rocm · host aihydra (igpu)
flags-ngl 999 -fa on -c 32768 --parallel 1 --load-mode none --jinja -rea off -ctk f16 -ctv f16
templatenot recorded at test time
treeupstream — stock

backfilled aged evidence — reconstructed from the archive · config record cfg-0027

decode @ 0
17.51 t/s
run-0039 · CV 0%
decode @ 32k
17.16 t/s
run-0043 · CV 0%
prefill @ 0
225.36 t/s
run-0038 · CV 0.8%
prefill @ 32k
212.18 t/s
run-0042 · CV 0.5%
τ² airline · thinking off
0.625 ±0.237
run-0102 · passed 10/16
turns to done · median (all tasks)
28
run-0102 · successes only: 26 · max 84
wall-clock to done · median (successes only)
3.7 min
run-0102 · failures excluded — they have no done
Wh per correct answer
unmeasured
open question ↓

every cell generated from the record at build time · throughput cells from cfg-0015 (same build 3653e6d, fa on, f16 KV) · provenance: measured-here throughout

③aDecode against context depth

f16 KV · fa onq8_0 KV · stockflash attention off
05101520032k65k131k204.8kdecode t/scontext depth (tokens)17.51 t/s @ depth 0 · run-0069 · CV 0% · N=317.51 t/s @ depth 0 · run-0039 · CV 0% · N=317.48 t/s @ depth 4k · run-0041 · CV 0% · N=317.16 t/s @ depth 32k · run-0043 · CV 0% · N=317.17 t/s @ depth 32k · run-0071 · CV 0% · N=316.21 t/s @ depth 131k · run-0073 · CV 0% · N=317.45 t/s @ depth 0 · run-0075 · CV 0.1% · N=316.25 t/s @ depth 32k · run-0077 · CV 0.1% · N=312.58 t/s @ depth 131k · run-0079 · CV 0.1% · N=317.46 t/s @ depth 0 · run-0063 · CV 0% · N=313.69 t/s @ depth 32k · run-0065 · CV 0% · N=36.41 t/s @ depth 131k · run-0067 · CV 0% · N=3f16 KV · fa on · 16.21q8_0 KV · stock · 12.58flash attention off · 6.41
3 reps per cell · max CV 0.1% · build 3653e6d · records: run-0069 run-0039 run-0041 run-0043 run-0071 run-0073 run-0075 run-0077 run-0079 run-0063 run-0065 run-0067

Other configurations tested — each as a delta against best

variantΔ decodeΔ turns (paired tasks)Δ energynoterecords
flash attention off-60% @131kThe shallow cells hide it entirely — the fa-off arm only collapses at depth. clm-0028
q8_0 KV · stock build-22% @131kNo dequant-patch arm has been run on this model. clm-0022
UD-Q4_K_M · fair-quant arm-10% n=12+83%Reward identical task-for-task on the seeded subset — the finer quant differs only in turn efficiency and window energy, so the IQ4_XS best-config rows understate this model's speed, never its quality. clm-0047

deltas computed at build time from the named runs, paired-task traces and metered windows — a hazard row shows words where a delta would mislead

Open questions

Provenance

bench host aihydra · rocm · ggml-org/llama.cpp 3653e6d
discipline 3 reps per throughput cell · scatter published per cell (max CV 0.8%) · guard chain on every performance series
window 2026-06-27 → 2026-08-11