HALOBENCH_
The reference lab for local AI on Strix Halo-class hardware: what to run, how to configure it, and what it truly costs — every number wall-measured, provenance-graded, and corrected in public.
last metered draw 137.6 W (2026-08-11, eng-0064)standing cost £2.24/moduty cycle 79%models live 1build-time snapshot, 2026-08-13 19:19Z
The topology — what this lab actually runs
topology as of 2026-08-13 19:19Z — build-time snapshot of the record, not a live read
[aibeast]dark since 2026-07-23
[aihydra]active
retired: 1 · rejected: 2 · live: 1 — every exit on a recorded measurement
states: live · candidate ·retired · rejected · dashed = planned / dark · node → its model or record page
Four measurements
9.8x
prefix-cache reuse on this box — an 8,000-token prefix costs 25.99 s cold, 2.65 s warm
clm-0027 · verified 2026-08-08
+70.3%
decode recovered at 204,800 tokens — production's own context — by one cherry-picked KV-dequant commit
clm-0022 · verified 2026-08-08
6.81 Wh
per correct τ² answer on the 122B at f16 KV — 0.21 p at 30.3 p/kWh
clm-0042 · verified 2026-08-10
42%
energy cut at 200k context by the KV dequant patch — 146.1 Wh stock against 85.3 Wh patched
clm-0041 · verified 2026-08-10
Field snapshot → full comparison
| model | params | decode @32k | τ² airline | guard |
|---|---|---|---|---|
| Qwen3.6-35B-A3B | 35B / 3B active | 42.45 t/s | not run | guard 4/4 |
| gpt-oss-120b | 117B / ~5B active | 41.70 t/s | not run | guard failed |
| Qwen3.5-122B-A10B (MTP) | 122B / 10B active | 18.17 t/s | 0.545 ±0.208 n=22 | guard 4/4 |
| Nemotron-3-Super-120B-A12B | 120B / ~12B active | 17.16 t/s | 0.625 ±0.237 n=16 | guard 4/4 |
best measured configuration per model · every row's full fingerprint on /compare · exclusions stay visible: an excluded row is data, not an omission
Start with a question → all questions
Is thinking mode worth it?
conclusion
Per-model, not global. A net negative on the 122B — it deadlocks on 2 of 5 hard tasks rather than degrading clm-0033 — so any global default is wrong for part of the field.
From the log → the full notebook
- 2026-08-13claimOn gfx1151 at f16 KV, stock Vulkan beats stock ROCm in EVERY cell of a matched matrix (one binary commit 3653e6d, one model, depths 0 to 131,072): decode +19-21% at every depth, prefill +4% to +20% growing with depth to 65k. clm-0050
- 2026-08-13claimQuantised KV on gfx1151 splits three ways by build. clm-0051
- 2026-08-13gate7 candidates → screened: gemma4-12b, glm-45-air, ling-30-flash, llama4-scout, maple-preview, nemotron35-lightning-30b, qwen36-27b-mtp
- 2026-08-13gatenemotron35-lightning-30b: listed → acquired — Official ggml-org Q4_K_M downloaded to aihydra (25,430,738,944 bytes, exactly the listed size) and sha256-verified against the HF LFS oid (6110e2e2e6cd324e6ee69ddced5a6b34fad6c94ca9827222a1e420fb92e3c90b).
- 2026-08-12gatedeepseek-v4-flash: listed → listed — Identity now concrete via the r/LocalLLaMA Strix Halo guide thread: DeepSeek V4 Flash 0731, deepseek4 arch, 256 experts/6 active + 1 shared, MIT licence, 1M native context. clm-0031
- 2026-08-12gatenemotron35-lightning-30b: listed — Surfaced via an operator-shared r/AIDeveloperNews link ("NVIDIA has launched Nemotron 3.5 Lightning") plus a follow-up r/StrixHalo post on a community ROCmFP4 requant with hardware-matched Strix Halo numbers.
- 2026-08-11claimA second independent Strix Halo source (llama.cpp PR #26856 + its Reddit write-up) reports Vulkan ahead of ROCm on decode at depth by ~10.5% on a clean same-binary comparison — same direction as clm-0031's +55% but a fifth the magnitude, confirming that figure was mostly build-gap and private patches. clm-0044
- 2026-08-11claimThe KV dequant patch removes most of quantised KV's agentic cost, not just its speed cost: on identical seeded tasks, patched q8_0 takes +9.3% more turns than patched f16 (234 vs 214 over 11 paired tasks) where the stock build cost +39% (clm-0038). clm-0045
- 2026-08-11claimThe community BF16 flash-attention predictions (clm-0044) reproduce on this hardware under single-binary methodology: at 32k depth, stock Vulkan leads ROCm +17.4% prefill / +17.1% decode at f16 KV, and the PR-26856 patch inverts prefill to ROCm +46.9% with bf16 KV. clm-0046
- 2026-08-11claimNemotron's quant confound resolves cleanly: on 12 identical seeded tasks, UD-Q4_K_M and UD-IQ4_XS produce IDENTICAL reward on every task (0.583 both), but IQ4_XS takes +11% more total turns (630 vs 567). clm-0047
- 2026-08-11claimSimulator identity changes what tau2 measures. clm-0048
- 2026-08-11claimThe 122B has a reproducible rule-precedence bug in policy application: on tau2 airline task 9 it cancels a partially-flown reservation in 7 of 8 trials, every failure with the identical signature — cancel_reservation called without ever checking flight status. clm-0049
- 2026-08-11gatemuse-glimmer-30b: blocked → screened — Screened on min-62bf73d (its minimum build, anchor-calibrated: +2.8%/+0.9% vs fleet baseline, identical tau2 capability).
- 2026-08-11runs9 runs landed on cfg-0029, cfg-0030, cfg-0025, cfg-0031, cfg-0028, cfg-0027 (tau2-bench-airline) cfg-0029 cfg-0030 cfg-0025 cfg-0031 cfg-0028 cfg-0027
How to read our numbers → method
measured-here wall-measured on this lab's hardware
community reported elsewhere, not yet reproduced here
vendor a vendor's own claim — never grounds for a verdict
unmeasured an honest gap, shown rather than hidden
retracted withdrawn in public, kept visible with its correction