HALOBENCH_

The reference lab for local AI on Strix Halo-class hardware: what to run, how to configure it, and what it truly costs — every number wall-measured, provenance-graded, and corrected in public.

last metered draw 137.6 W (2026-08-11, eng-0064)standing cost £2.24/moduty cycle 79%models live 1build-time snapshot, 2026-08-13 19:19Z

The topology — what this lab actually runs

topology as of 2026-08-13 19:19Z — build-time snapshot of the record, not a live read

[aibeast]dark since 2026-07-23

hardware failure — RMA in progress · inc-0005
5 configs · 4 runs benched here · latest 2026-07-20

[aihydra]active

idle floor 10.1 W · resident-quiet 13.6 W · duty cycle 79% · idle-aihydra-2026-08
26 configs · 107 runs benched here · latest 2026-08-11
  • no production nodes recorded — bench duty
retired: 1 · rejected: 2 · live: 1 — every exit on a recorded measurement
states: live · candidate ·retired · rejected · dashed = planned / dark · node → its model or record page

Four measurements

9.8x
prefix-cache reuse on this box — an 8,000-token prefix costs 25.99 s cold, 2.65 s warm
clm-0027 · verified 2026-08-08
+70.3%
decode recovered at 204,800 tokens — production's own context — by one cherry-picked KV-dequant commit
clm-0022 · verified 2026-08-08
6.81 Wh
per correct τ² answer on the 122B at f16 KV — 0.21 p at 30.3 p/kWh
clm-0042 · verified 2026-08-10
42%
energy cut at 200k context by the KV dequant patch — 146.1 Wh stock against 85.3 Wh patched
clm-0041 · verified 2026-08-10

Field snapshot → full comparison

modelparamsdecode @32kτ² airlineguard
Qwen3.6-35B-A3B35B / 3B active42.45 t/snot runguard 4/4
gpt-oss-120b117B / ~5B active41.70 t/snot runguard failed
Qwen3.5-122B-A10B (MTP)122B / 10B active18.17 t/s0.545 ±0.208 n=22guard 4/4
Nemotron-3-Super-120B-A12B120B / ~12B active17.16 t/s0.625 ±0.237 n=16guard 4/4

best measured configuration per model · every row's full fingerprint on /compare · exclusions stay visible: an excluded row is data, not an omission

Start with a question → all questions

Does KV quantisation cost quality?

conclusion

With a one-commit engine patch, mostly no: +70.3% generation speed at production depth clm-0022, 42% less energy clm-0041, and the quality cost shrinks from +39% extra conversation turns clm-0038 to +9.3% clm-0045.

Vulkan or ROCm on Strix Halo?

conclusion

Depth decides. At 32k, stock Vulkan leads ROCm +17.4% prefill and +17.1% decode at f16 KV, and a pending BF16 flash-attention patch inverts prefillclm-0046. Community reports agree on direction and disagree on magnitude clm-0044 clm-0031.

Is thinking mode worth it?

conclusion

Per-model, not global. A net negative on the 122B — it deadlocks on 2 of 5 hard tasks rather than degrading clm-0033 — so any global default is wrong for part of the field.

From the log → the full notebook

  • 2026-08-13claimOn gfx1151 at f16 KV, stock Vulkan beats stock ROCm in EVERY cell of a matched matrix (one binary commit 3653e6d, one model, depths 0 to 131,072): decode +19-21% at every depth, prefill +4% to +20% growing with depth to 65k. clm-0050
  • 2026-08-13claimQuantised KV on gfx1151 splits three ways by build. clm-0051
  • 2026-08-13gate7 candidates → screened: gemma4-12b, glm-45-air, ling-30-flash, llama4-scout, maple-preview, nemotron35-lightning-30b, qwen36-27b-mtp
  • 2026-08-13gatenemotron35-lightning-30b: listed → acquired — Official ggml-org Q4_K_M downloaded to aihydra (25,430,738,944 bytes, exactly the listed size) and sha256-verified against the HF LFS oid (6110e2e2e6cd324e6ee69ddced5a6b34fad6c94ca9827222a1e420fb92e3c90b).
  • 2026-08-12gatedeepseek-v4-flash: listed → listed — Identity now concrete via the r/LocalLLaMA Strix Halo guide thread: DeepSeek V4 Flash 0731, deepseek4 arch, 256 experts/6 active + 1 shared, MIT licence, 1M native context. clm-0031
  • 2026-08-12gatenemotron35-lightning-30b: listed — Surfaced via an operator-shared r/AIDeveloperNews link ("NVIDIA has launched Nemotron 3.5 Lightning") plus a follow-up r/StrixHalo post on a community ROCmFP4 requant with hardware-matched Strix Halo numbers.
  • 2026-08-11claimA second independent Strix Halo source (llama.cpp PR #26856 + its Reddit write-up) reports Vulkan ahead of ROCm on decode at depth by ~10.5% on a clean same-binary comparison — same direction as clm-0031's +55% but a fifth the magnitude, confirming that figure was mostly build-gap and private patches. clm-0044
  • 2026-08-11claimThe KV dequant patch removes most of quantised KV's agentic cost, not just its speed cost: on identical seeded tasks, patched q8_0 takes +9.3% more turns than patched f16 (234 vs 214 over 11 paired tasks) where the stock build cost +39% (clm-0038). clm-0045
  • 2026-08-11claimThe community BF16 flash-attention predictions (clm-0044) reproduce on this hardware under single-binary methodology: at 32k depth, stock Vulkan leads ROCm +17.4% prefill / +17.1% decode at f16 KV, and the PR-26856 patch inverts prefill to ROCm +46.9% with bf16 KV. clm-0046
  • 2026-08-11claimNemotron's quant confound resolves cleanly: on 12 identical seeded tasks, UD-Q4_K_M and UD-IQ4_XS produce IDENTICAL reward on every task (0.583 both), but IQ4_XS takes +11% more total turns (630 vs 567). clm-0047
  • 2026-08-11claimSimulator identity changes what tau2 measures. clm-0048
  • 2026-08-11claimThe 122B has a reproducible rule-precedence bug in policy application: on tau2 airline task 9 it cancels a partially-flown reservation in 7 of 8 trials, every failure with the identical signature — cancel_reservation called without ever checking flight status. clm-0049
  • 2026-08-11gatemuse-glimmer-30b: blocked → screened — Screened on min-62bf73d (its minimum build, anchor-calibrated: +2.8%/+0.9% vs fleet baseline, identical tau2 capability).
  • 2026-08-11runs9 runs landed on cfg-0029, cfg-0030, cfg-0025, cfg-0031, cfg-0028, cfg-0027 (tau2-bench-airline) cfg-0029 cfg-0030 cfg-0025 cfg-0031 cfg-0028 cfg-0027

How to read our numbers → method

measured-here wall-measured on this lab's hardware
community reported elsewhere, not yet reproduced here
vendor a vendor's own claim — never grounds for a verdict
unmeasured an honest gap, shown rather than hidden
retracted withdrawn in public, kept visible with its correction