⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.
clm-0063
measured-heremed ●●○
citable URL: https://halobench.com/records/clm-0063/ — this address never moves; the anchor /records/#clm-0063 keeps resolving
On gfx1151 at UD-Q4_K_M / f16 KV, Nemotron-3-Super-120B-A12B is the FIRST model in this programme where Vulkan decode measurably beats ROCm: +8.97% at d0 (18.2043 vs 16.7065 t/s) and +8.19% at d32768 (17.7386 vs 16.3962 t/s), a stable signature by depth. ROCm keeps its usual prefill lead (+27.1% at d0, +26.9% at d32768) — so the per-model, per-phase backend rule (clm-0050) holds again, but this is the first candidate where the two phases point to DIFFERENT backends being "the better one" rather than the same backend winning both or Vulkan losing outright to a device-loss crash. Unlike qwen38-27b (clm-0054) and deepseek-v4-flash (clm-0059), Vulkan showed ZERO device-loss at d32768 here — 3/3 fresh-process reps clean on both depths tested, no vk::DeviceLostError, no amdgpu ring reset. Because decode dominates wall-clock in a multi-turn agentic conversation (prefill is amortised across turns via prefix-cache reuse; decode is not), this job picked Vulkan as the serving backend for the 26-task tau2 arm — a genuinely new per-model call, not an inherited default.
METHOD — nemotron3-super-fullbench-matrix: stock llama.cpp 3653e6d (ROCm) / 3653e6d6d (Vulkan, prefix-matches per house convention), UD-Q4_K_M, f16 KV (candidate's serving config, cfg-0072/cfg-0089 — not swept here), fa on, -ngl 999, --load-mode none, pp1024/tg256, llama-bench defaults otherwise (-b 2048/-ub 512, -t 16). N=3 fresh-process reps at d0/d32768 (this job's brief-specified depth set — "up to the fit ceiling"; see clm-0062 for why d32768 is nowhere near this candidate's actual memory ceiling post-fix). Every cell's coefficient of variation <=0.92% — clean per protocol §0's ~3% scatter-flag threshold.
pp1024 / tg256 by depth (median of 3):
| depth | ROCm | Vulkan | vk/rocm pp | vk/rocm tg | |---|---|---|---|---| | 0 | 271.5460 / 16.7065 | 213.7281 / 18.2043 | 0.787 | 1.090 | | 32,768 | 244.7670 / 16.3962 | 192.8306 / 17.7386 | 0.788 | 1.082 |
THE DECISION — cfg-0089 (the tau2 serving config) was set to Vulkan on this measurement, not the 2026-08-15 screen's ROCm pick (cfg-0072, made when no throughput data existed and both backends looked equally healthy at the FIT stage). This job's house 4-item guard (run-0307) then ran fresh against the Vulkan config and passed 4/4 clean before the tau2 arm launched, per protocol's guard-substitutions rule — the backend switch got its own capability check rather than inheriting the screen's ROCm-side guard evidence.
WHY THIS DIFFERS FROM THE PRIOR PATTERN — qwen38-27b and deepseek-v4-flash both showed Vulkan collapsing or crashing with depth, so ROCm was the only real option past d0. Nemotron-3-Super's fit ceiling in this matrix (d32768) is shallower than where those models' Vulkan arms broke (>=32,768/>=131,072 respectively) only in the sense that this job did not test deeper cells (clm-0062) — whether Vulkan would eventually lose the device on this candidate at a much greater depth is UNTESTED, and the tau2 arm's serving context (32768) sits exactly at the deepest cell this matrix cleared, not comfortably past it. This is noted as an open question on the model page, not glossed over.