Home › Evidence › Records › clm-0064
⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.

clm-0064

measured-heremed ●●○
citable URL: https://halobench.com/records/clm-0064/ — this address never moves; the anchor /records/#clm-0064 keeps resolving

Nemotron-3-Super-120B-A12B UD-Q4_K_M completed the FULL standard 26-task tau2 airline set cleanly (rc=0, not a wall-bound cut) at 0.7692 mean reward (20/26), 263 tool-call messages, 429 assistant messages, 0 empty assistant turns, 0 infrastructure errors -- VALID under the SMOKE gate. This is the FIRST tau2 result for this candidate, at ANY quant, run under the protocol's PINNED independent simulator (openrouter/anthropic/claude-haiku-4.5) -- every prior full tau2 number on this candidate (the IQ4_XS acquisition-era run-0102, and the quant-flip pair run-0107/run-0108) ran SELF-PLAY, agent_llm == user_llm, which the candidate's own open-questions flag as measuring something weaker ("self-play flattens what the benchmark measures"). The 2026-08-15 screen's 5-task SMOKE (4/5 = 0.800) was the only other pinned-simulator number this candidate had; this run's first five tasks (task-ids 0-4) score 4/5 = 0.800 again on the identical seed and task set, cross-confirming the screen was not a fluke -- SAME reward, though not necessarily the same per-task pattern, since the screen's own per-task breakdown is not separately recorded. VULKAN, not the screen's ROCm serving pick -- this job's own throughput matrix measured Vulkan +8-9% faster decode with zero device-loss at this candidate's fit ceiling (clm-0063), so the capability-establishing run itself ran on the newly-measured better backend, cfg-0089. Energy: 648.60 Wh across the whole 15,968 s (4.44 h) window, mean 146.23 W (delta 136.13 W over the 10.1 W idle floor) -- 32.43 Wh per correct answer, 0.98 pence at 30.3 p/kWh, whole-session (includes simulator wait time end to end). Against this programme's other two full-bench candidates on the SAME denominator: 1.34x qwen38-27b's 38.29 Wh/correct answer (draft-mtp dense, cheaper) but 1.15x more than deepseek-v4-flash's 24.29 Wh/correct answer (a small-active-param MoE) -- despite this candidate having the LARGEST active-parameter footprint measured on this box (12B active vs deepseek's ~6B-equivalent MoE activation and qwen38's dense 27B all-active), it lands in the middle of the three on Wh/correct, not at either extreme, which is a genuinely interesting finding worth noting rather than assuming active-parameter count predicts energy ranking cleanly.

verified 2026-08-16 · volatility medium
evidence run-0316

Note — the record's own working

METHOD -- tau2-bench airline, harness 668d3bc (same version as qwen38-27b's run-0271 and deepseek-v4-flash's run-0287), --num-trials 1 --seed 42 --max-steps 200 --max-concurrency 1, user simulator pinned per protocol. Agent: local llama-server endpoint on cfg-0089 (Vulkan, f16 KV mandatory per the screen's serving config, -c 32768, --parallel 1, --load-mode none, no speculation -- Mamba rejects draft-mtp), agent sampling temperature 0.0 / max_tokens 4096 -- identical to the 2026-08-15 screen's SMOKE arm. Task-ids 0-25 explicit (the STANDARD 26-task set this programme adopted after qwen38-27b's wall-bound experience). THE RUN COMPLETED IN FULL: 02:32:56Z to 06:59:04Z, 15,968 s, rc=0, well inside the 28,800 s (8 h) safety ceiling. Slowest task (10) took 1,629.6 s (27.2 min); several tasks ran past 700-800 s. No pattern in the six failures (tasks 7, 10, 14, 20, 23, 24) suggesting a systematic cliff -- all resolved via user_stop, spread across both short (tc=6) and long (tc=19) tool-call counts, and across the full duration range (404.8 s to 1,629.6 s). No cut_at_max_steps, no infrastructure error, in any of the 26 tasks. SIMULATOR PROVENANCE MATTERS HERE MORE THAN USUAL: this is a genuinely new measurement, not a re-run of an existing number under a different simulator -- there is no earlier pinned-simulator FULL run to compare 0.7692 against for this candidate. The only same-simulator anchor is the 2026-08-15 SMOKE's 5-task subset (0.800), which this run reproduces exactly on tasks 0-4. ENERGY (same-session join, eng-0116): counter interpolated from raw ~10s HA history samples at each run boundary (not the coarser 5-minute statistics bucket, given the multi-hour window's need for precision at the exact run-meta timestamps): start 20.99128201 kWh, end 21.63988203 kWh, delta 0.64860002 kWh = 648.60 Wh. Mean power 146.23 W sits between this job's own throughput-matrix cells at the same depth (rocm-d32768 147.74 W / eng-0113, vulkan-d32768 149.47 W / eng-0115) -- a sanity-check the tau2 arm's power draw is not anomalous relative to raw decode-bound throughput at the same serving context. CROSS-MODEL ENERGY CAVEAT -- not a controlled A/B: three different quants (Q4_K_M here, Q8_0 for qwen38-27b, IQ3_XXS for deepseek-v4-flash), three different architectures (hybrid Mamba-2+LatentMoE MoE / dense hybrid-attention / MLA MoE), different backends (Vulkan here, ROCm for both peers), and qwen38-27b ran with draft-mtp speculation while this candidate and deepseek-v4-flash both ran plain decode. Reported because it is the only same-shape (26-task, same simulator, same protocol, same wh_per_task denominator) energy comparison available in this programme, not because the confound list has been controlled for.

Cited by — computed at build time, never stored

model pages nemotron3-super
candidate gate history nemotron3-super