Models › nemotron3-super
⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.

Nemotron-3-Super-120B-A12Bbenchedguard 4/4

120B / ~12B activehybrid Mamba-2 + LatentMoE MoE — no MTP/speculation path in llama.cppquant held: UD-Q4_K_M (fair-quant, this bench) · UD-IQ4_XS (superseded acquisition quant, kept separate)first measured 2026-06-27latest run 2026-08-19

②Verdict

A 120-billion-parameter hybrid Mamba-2 + LatentMoE mixture-of-experts activating 12B per token, with no speculation path in llama.cpp (Mamba rejects draft-mtp), served at UD-Q4_K_M — the fleet-standard quant this candidate's earlier benching (clm-0035, retracted by clm-0036) was never matched to, forced instead onto IQ4_XS by the old 96 GB box.

That confound is now resolved twice over: the 2026-08-15 screen was the first Nemotron-3-Super capability number at the same quant as its peers, and this full bench adds the first FULL 26-task tau2 result for this candidate, at any quant, run under the protocol's pinned independent simulator clm-0064 — every earlier full run, including the quant-flip pair that first suggested quantisation cost this model only efficiency and never correctness, ran self-play, which this candidate's own history had already flagged as measuring something weaker clm-0047. The pinned run scored 0.7692 (20/26), 263 tool-call messages, 0 empty turns — VALID under the SMOKE gate, and its first five tasks reproduce the screen's own 0.800 exactly on the same seed clm-0064.

Backend is the standout finding of this bench's throughput matrix: ROCm keeps its usual ~27% prefill lead, but Vulkan decodes 8-9% faster at both measured depths, with a later d204800 crossover cell preserving the same phase split (Vulkan decode faster, ROCm prefill faster) under the matching cfg-0108/cfg-0107 deep records clm-0091. Unlike qwen38-27b and deepseek-v4-flash, Vulkan showed ZERO device-loss at this candidate's d32768 fit ceiling in the full bench, and later survived d131072 and d204800 throughput cells as well — an observed case measured here where Vulkan wins decode outright rather than losing the device with depth clm-0063 clm-0091. Because decode dominates a multi-turn agentic session, this job served the tau2 arm on Vulkan, a genuinely new per-model backend call.

The KV cache costs 8.00 KiB/token — close to deepseek-v4-flash's 7.13 and far below qwen38-27b's 64.00 — so this candidate is nowhere near GTT-bound at any depth tested here; the earlier "OOM past d0" reading was a mmap double-residency harness defect on this exact candidate (clm-0053), not a memory fact, and this session's clean matrix on both backends confirms the fix generalises clm-0062.

Energy sits between the cited qwen38-27b and deepseek-v4-flash same-protocol records: 32.43 Wh per correct answer, 0.98 pence at the standing tariff, cheaper than qwen38-27b's 38.29 Wh but more than deepseek-v4-flash's 24.29 Wh — despite carrying a larger active-parameter footprint than those two cited comparators, which is worth noting precisely because it does not fit a simple "more active parameters costs more energy" story clm-0064.

verdict written 2026-08-16 · every number above stands next to the claim chip that carries it

③Best configuration

modelNVIDIA-Nemotron-3-Super-120B-A12B-UD-Q4_K_M-00001-of-00003.gguf · UD-Q4_K_M
engineggml-org/llama.cpp 3653e6d · vulkan · host aihydra (igpu)
flags-ngl 999 -fa on -c 32768 -dev Vulkan0 --load-mode none --parallel 1 --jinja --reasoning-format deepseek -n 4096 --slots
templatenot recorded at test time
treeupstream — stock

config record cfg-0089

decode @ 0
18.20 t/s
run-0313 · CV 0%
decode @ 32k
17.74 t/s
run-0315 · CV 0%
prefill @ 0
213.73 t/s
run-0312 · CV 0.6%
prefill @ 32k
192.83 t/s
run-0314 · CV 0.5%
τ² airline · thinking off
0.769 ±0.162
run-0316 · passed 20/26
Wh per correct answer (τ², thinking off)
32.43 Wh
eng-0116 run-0316 · 0.98 p per answer

every cell generated from the record at build time · throughput cells from cfg-0088 (same build 3653e6d, fa on, f16 KV) · provenance: measured-here throughout

③aDecode against context depth

Vulkan · UD-Q4_K_MROCm · UD-Q4_K_MROCm · UD-IQ4_XS (superseded quant, f16 KV)
05101520032k65k131k204.8kdecode t/scontext depth (tokens)serving context 32k15.88 t/s @ depth 0 · run-049318.20 t/s @ depth 0 · run-0313 · CV 0% · N=317.74 t/s @ depth 32k · run-0315 · CV 0% · N=316.62 t/s @ depth 131k · run-0390 · CV 0% · N=116.71 t/s @ depth 0 · run-0309 · CV 0% · N=316.40 t/s @ depth 32k · run-0311 · CV 0% · N=315.53 t/s @ depth 131k · run-0386 · CV 0% · N=115.01 t/s @ depth 204.8k · run-0388 · CV 0% · N=117.51 t/s @ depth 0 · run-0069 · CV 0% · N=317.51 t/s @ depth 0 · run-0039 · CV 0% · N=317.48 t/s @ depth 4k · run-0041 · CV 0% · N=317.16 t/s @ depth 32k · run-0043 · CV 0% · N=317.17 t/s @ depth 32k · run-0071 · CV 0% · N=316.21 t/s @ depth 131k · run-0073 · CV 0% · N=3Vulkan · UD-Q4_K_M · 16.62ROCm · UD-IQ4_XS (superseded quant, f16 KV) · 16.21ROCm · UD-Q4_K_M · 15.01
1/3 reps per cell · max CV 0.0% · build 3653e6d / 3653e6d · UD-Q4_K_M lines originally covered d0/d32768 only — this job's brief-specified depth set, itself the same depth the earlier IQ4_XS-era queue script recorded as an OOM under mmap (a harness defect, not a memory fact — clm-0062/clm-0053) and which this session's fresh matrix cleared cleanly on both backends. A 2026-08-17 depth-ladder extension (Part 2 deep cells) pushed ROCm to d131072 and d204800, and Vulkan to d131072; HO-006 then filled the matching Vulkan d204800 throughput cell under the same cfg-0108 fingerprint. Both backends stayed clean at every record-backed deep cell (no OOM, no device loss), and Vulkan kept its decode lead over ROCm at both d131072 and d204800 while ROCm kept the prefill lead. The UD-IQ4_XS line predates the mmap fix and covers a different depth range on a different quant; the two series are never merged, only overlaid for visual comparison. · records: run-0493 run-0313 run-0315 run-0390 run-0309 run-0311 run-0386 run-0388 run-0069 run-0039 run-0041 run-0043 run-0071 run-0073

④Other configurations tested — each as a delta against best

variantΔ decodeΔ turns (paired tasks)Δ energynoterecords
Vulkan backend · same quant+9.0% @0—+13%The first model in the cited local measurements where Vulkan decode beats ROCm outright (+8.97% at d0), not a wash and not a ROCm win. Vulkan also showed ZERO device-loss at this candidate's fit ceiling (d32768) — a different outcome from qwen38-27b and deepseek-v4-flash, both of which lost the Vulkan device with depth. ROCm keeps its usual prefill lead (~27%), but decode dominates a multi-turn agentic session, which is why this job served the τ² arm on Vulkan. clm-0063
flash attention off-60% @131k——The shallow cells hide it entirely — the fa-off arm only collapses at depth. clm-0028
q8_0 KV · stock build-22% @131k——No dequant-patch arm has been run on this model. clm-0022
UD-Q4_K_M · fair-quant arm (self-play, pre-dates this bench)—-10% n=12+83%Reward identical task-for-task on the seeded subset — the finer quant differs only in turn efficiency and window energy, so the IQ4_XS best-config rows understate this model's speed, never its quality. Measured SELF-PLAY (agent_llm == user_llm), not the pinned simulator this bench's headline τ² number uses — see the next row. clm-0047
pinned-simulator quant comparability———This bench's 0.7692 headline (run-0316) is the first tau2 result for this candidate, at any quant, run under the protocol's pinned independent simulator. No UD-IQ4_XS run has ever been measured under that simulator, so there is no valid pinned-simulator quant-vs-quant comparison to report here — only the self-play pair above, which the candidate's own history already flags as measuring something weaker than the pinned setup. clm-0064

deltas computed at build time from the named runs, paired-task traces and metered windows — a hazard row shows words where a delta would mislead

⑤Open questions

⑥Provenance

bench host aihydra · vulkan · ggml-org/llama.cpp 3653e6d
discipline 3 reps per throughput cell · scatter published per cell (max CV 0.6%) · guard or written waiver on every performance series
window 2026-06-27 → 2026-08-19