Home › Evidence › Records › nemotron3-super

nemotron3-super

rejected
citable URL: https://halobench.com/records/nemotron3-super/ — this address never moves; the anchor /records/#nemotron3-super keeps resolving
Nemotron 3 Super @ IQ4_XS (not quant-equivalent to peers) · kind model · engine igpu · tier candidate
runs on aibeast · config cfg-0004
⌁ current state rejected · ⌁ days in production 0

Lifecycle — append-only

2026-06-27
candidate · run-0003
2026-06-28
rejected — dec-0001 · run-0003
Purpose-built for agentic reasoning and tool use, and the most compromised entry in the field. Its Q4_K quant does not fit: 77 GiB at 4K context, OOM by 16K on an ~82 GiB budget, so it was tested at IQ4_XS — a LOWER quant than every peer. It was also the only model with no speculative decoding available, because llama.cpp does not support the Mamba-hybrid MTP path. Slowest in field at ~28 s. THIS RESULT IS NOT QUANT-EQUIVALENT and should not be read as a like-for-like verdict on the model — only on what this box could actually run. See clm-0005.