Nemotron-3-Super-120B-A12Bbenchedguard 4/4
②Verdict
A 120-billion-parameter hybrid Mamba-2 + LatentMoE mixture-of-experts activating 12B per token, with no speculation path in llama.cpp (Mamba rejects draft-mtp), served at UD-Q4_K_M — the fleet-standard quant this candidate's earlier benching (clm-0035, retracted by clm-0036) was never matched to, forced instead onto IQ4_XS by the old 96 GB box.
That confound is now resolved twice over: the 2026-08-15 screen was the first Nemotron-3-Super capability number at the same quant as its peers, and this full bench adds the first FULL 26-task tau2 result for this candidate, at any quant, run under the protocol's pinned independent simulator clm-0064 — every earlier full run, including the quant-flip pair that first suggested quantisation cost this model only efficiency and never correctness, ran self-play, which this candidate's own history had already flagged as measuring something weaker clm-0047. The pinned run scored 0.7692 (20/26), 263 tool-call messages, 0 empty turns — VALID under the SMOKE gate, and its first five tasks reproduce the screen's own 0.800 exactly on the same seed clm-0064.
Backend is the standout finding of this bench's throughput matrix: ROCm keeps its usual ~27% prefill lead, but Vulkan decodes 8-9% faster at both measured depths, with a later d204800 crossover cell preserving the same phase split (Vulkan decode faster, ROCm prefill faster) under the matching cfg-0108/cfg-0107 deep records clm-0091. Unlike qwen38-27b and deepseek-v4-flash, Vulkan showed ZERO device-loss at this candidate's d32768 fit ceiling in the full bench, and later survived d131072 and d204800 throughput cells as well — an observed case measured here where Vulkan wins decode outright rather than losing the device with depth clm-0063 clm-0091. Because decode dominates a multi-turn agentic session, this job served the tau2 arm on Vulkan, a genuinely new per-model backend call.
The KV cache costs 8.00 KiB/token — close to deepseek-v4-flash's 7.13 and far below qwen38-27b's 64.00 — so this candidate is nowhere near GTT-bound at any depth tested here; the earlier "OOM past d0" reading was a mmap double-residency harness defect on this exact candidate (clm-0053), not a memory fact, and this session's clean matrix on both backends confirms the fix generalises clm-0062.
Energy sits between the cited qwen38-27b and deepseek-v4-flash same-protocol records: 32.43 Wh per correct answer, 0.98 pence at the standing tariff, cheaper than qwen38-27b's 38.29 Wh but more than deepseek-v4-flash's 24.29 Wh — despite carrying a larger active-parameter footprint than those two cited comparators, which is worth noting precisely because it does not fit a simple "more active parameters costs more energy" story clm-0064.
③Best configuration
| model | NVIDIA-Nemotron-3-Super-120B-A12B-UD-Q4_K_M-00001-of-00003.gguf · UD-Q4_K_M |
| engine | ggml-org/llama.cpp 3653e6d · vulkan · host aihydra (igpu) |
| flags | -ngl 999 -fa on -c 32768 -dev Vulkan0 --load-mode none --parallel 1 --jinja --reasoning-format deepseek -n 4096 --slots |
| template | not recorded at test time |
| tree | upstream — stock |
config record cfg-0089
every cell generated from the record at build time · throughput cells from cfg-0088 (same build 3653e6d, fa on, f16 KV) · provenance: measured-here throughout
③aDecode against context depth
④Other configurations tested — each as a delta against best
| variant | Δ decode | Δ turns (paired tasks) | Δ energy | note | records |
|---|---|---|---|---|---|
| Vulkan backend · same quant | +9.0% @0 | — | +13% | The first model in the cited local measurements where Vulkan decode beats ROCm outright (+8.97% at d0), not a wash and not a ROCm win. Vulkan also showed ZERO device-loss at this candidate's fit ceiling (d32768) — a different outcome from qwen38-27b and deepseek-v4-flash, both of which lost the Vulkan device with depth. ROCm keeps its usual prefill lead (~27%), but decode dominates a multi-turn agentic session, which is why this job served the τ² arm on Vulkan. | clm-0063 |
| flash attention off | -60% @131k | — | — | The shallow cells hide it entirely — the fa-off arm only collapses at depth. | clm-0028 |
| q8_0 KV · stock build | -22% @131k | — | — | No dequant-patch arm has been run on this model. | clm-0022 |
| UD-Q4_K_M · fair-quant arm (self-play, pre-dates this bench) | — | -10% n=12 | +83% | Reward identical task-for-task on the seeded subset — the finer quant differs only in turn efficiency and window energy, so the IQ4_XS best-config rows understate this model's speed, never its quality. Measured SELF-PLAY (agent_llm == user_llm), not the pinned simulator this bench's headline τ² number uses — see the next row. | clm-0047 |
| pinned-simulator quant comparability | — | — | — | This bench's 0.7692 headline (run-0316) is the first tau2 result for this candidate, at any quant, run under the protocol's pinned independent simulator. No UD-IQ4_XS run has ever been measured under that simulator, so there is no valid pinned-simulator quant-vs-quant comparison to report here — only the self-play pair above, which the candidate's own history already flags as measuring something weaker than the pinned setup. | clm-0064 |
deltas computed at build time from the named runs, paired-task traces and metered windows — a hazard row shows words where a delta would mislead
⑤Open questions
- A pinned-simulator UD-IQ4_XS run — without it, the quant-vs-quant comparison stays confined to the self-play pair (identical reward, +11% turns at IQ4_XS), which this candidate's own history already treats as a weaker measurement than the pinned setup this bench's headline number uses clm-0047 clm-0064.
- Whether Vulkan's clean d204800 throughput result holds at the model's advertised max context — qwen38-27b and deepseek-v4-flash both lost the Vulkan device at deeper cells than where they first looked healthy, so "no device-loss through 204800" is still not "immune at any depth" for this candidate either clm-0063 clm-0091.
- Cross-model comparison against the 122B incumbent — the prior verdict's turn-count and wall-clock comparisons (clm-0039) were built on pre-pinned-simulator data and are now provisional pending a matched re-run of the incumbent under the same pinned Haiku simulator this bench used.
- Long-context agentic behaviour above the 32,768-token serving context — unmeasured; this bench's capability run and throughput matrix both stayed at the depth this candidate has been screened at since 2026-08-15.
- Root cause of the task-7/10/14/20/23/24 failures — no shared pattern in tool-call count or duration was visible in this bench's own summary; a transcript read would be needed to say more clm-0064.