clm-0005
measured-herehigh ●●●
citable URL: https://halobench.com/records/clm-0005/ — this address never moves; the anchor /records/#clm-0005 keeps resolving
Nemotron 3 Super's benchmark result is NOT quant-equivalent to its peers and must not be read as a like-for-like verdict on the model.
verified 2026-06-28 · volatility low
evidence run-0003
Note — the record's own working
Every model was tested at the best quant that fits 128K in ~82 GiB. For the field that was a Q4_K-class quant; for Nemotron it was IQ4_XS, because its Q4_K needs 77 GiB at 4K context and OOMs by 16K. It was additionally the only entry with no speculative decoding available — llama.cpp does not support the Mamba-hybrid MTP path — while peers could use it. So it was handicapped twice, and came last on speed (~28 s). The comparison is honest about what THIS BOX can run; it is not honest about the model. This caveat travels with the result rather than sitting in a footnote, which is the whole reason it is a record and not a sentence.
Cited by — computed at build time, never stored
model pages nemotron3-super
docs benchmark-protocol