Home › Evidence › Records › cfg-0108
⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.

cfg-0108

backfilled
citable URL: https://halobench.com/records/cfg-0108/ — this address never moves; the anchor /records/#cfg-0108 keeps resolving
modelNVIDIA-Nemotron-3-Super-120B-A12B-UD-Q4_K_M-00001-of-00003.gguf · UD-Q4_K_M
engineggml-org/llama.cpp 3653e6d6d · vulkan · host aihydra (igpu)
flags-ngl 999 -fa 1 -b 2048 -ub 512 -ctk f16 -ctv f16 -t 16 --load-mode none
templatenot recorded at test time
treeupstream — stock
aged evidence — reconstructed from the archive. Part 2 deep-cell session (2026-08-17), same fingerprint as cfg-0088 (identical llama-bench invocation: build 3653e6d6d, Vulkan, UD-Q4_K_M, f16 KV, -fa 1, --load-mode none -- only depth differs). d131072 attempt: SURVIVED cleanly on the first of a 2-attempt cap, no vk::DeviceLostError, no GPU wedge -- notable given this box's #25664 device-loss dataset at comparable depth on other candidates (laguna-s-21, deepseek-v4-flash). A fresh config id is minted per house convention. Ingested from llama-bench, which measures throughput and not footprint.

Capability basis

unmeasured — performance numbers on this config stand on a guard alone, not an established capability

Runs on this config (4)

rundatekindsuitemetrics
run-03892026-08-17performancellama-bench@1prefill_tps=148.617391
run-03902026-08-17performancellama-bench@1decode_tps=16.622963
run-04922026-08-19performancellama-bench@1prefill_tps=130.812678
run-04932026-08-19performancellama-bench@1decode_tps=15.883158

Cited by — computed at build time, never stored

model pages nemotron3-super