⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.

HALOBENCH_

The reference lab for local AI on Strix Halo-class hardware: what to run, how to configure it, and what it truly costs — every number wall-measured, provenance-graded, and corrected in public.

last metered draw 186.56 W (2026-09-06, eng-0280)standing cost £2.24/mo (Aug 7-10 baseline, idle-aihydra-2026-08)duty cycle 79% (Aug 7-10, unrepresentative, idle-aihydra-2026-08)models live 1 (derived from node lifecycle records)build-time snapshot, 2026-09-30 09:19Z

The topology — what this lab actually runs

topology as of 2026-09-30 09:19Z — build-time snapshot of the record, not a live read

[aibeast]active

5 configs · 4 runs benched here · latest 2026-07-20

[aihydra]active

idle floor 10.1 W · resident-quiet 13.6 W · duty cycle 79% · idle-aihydra-2026-08
182 configs · 668 runs benched here · latest 2026-09-06
retired: 2 · rejected: 3 · live: 1 — every exit on a recorded measurement
states: live · candidate ·retired · rejected · dashed = planned / dark · node → its model or record page

The fleet, liveas of 2026-09-30 09:17Z

aibeastproductionserving
model
fast-rocmfp4 · Q4_0_ROCMFP4_FAST
power
15 W
uptime
3h 1m
decode
21.4 t/s
slots
0/1
aihydralab-benchserving
lane
lab / bench
power
28 W
model
/home/aihydra/models/qwen38-27b/Qwen3.8-27B-UD-Q4_K_XL.gguf

live snapshot, refreshed hourly by an out-of-band collector on the tailnet · an offline box degrades to its last-known values, never a blank · this strip says how old it is

Four measurements

9.8x
claim-recorded prefix-cache reuse on this box — an 8,000-token prefix costs 25.99 s cold, 2.65 s warm
clm-0027 · verified 2026-08-08
+70.3%
decode recovered at 204,800 tokens — production's own context — by one cherry-picked KV-dequant commit
clm-0022 · verified 2026-08-08
6.81 Wh
matched-task derived per correct τ² answer on the 122B at f16 KV — 0.21 p at 30.3 p/kWh
clm-0042 · verified 2026-08-10
42%
energy cut at 200k context by the KV dequant patch — 146.1 Wh stock against 85.3 Wh patched
clm-0041 · verified 2026-08-10

Field snapshot → full comparison

modelparamsdecode @32kτ² airlineguard
Ornith-1.0-35B (ornith-ai / deepreinforce-ai)35B / A3B-class MoE46.22 t/s0.885 ±0.123 n=26guard 4/4
Qwen3.6-35B-A3B35B / 3B active42.45 t/snot runguard 4/4
gpt-oss-120b117B / ~5B active41.70 t/snot runguard failed
Qwen3-Coder-Next (Qwen)80B total / ~3B active (A3B-class MoE, 512 experts / 10 active + 1 shared)36.97 t/s0.538 ±0.192 n=26guard 4/4
Ling-3.0-flash124B / 5.1B active32.89 t/s0.500 ±0.192 n=26guard 4/4
Qwen3.8-27B27B dense19.23 t/s0.917 ±0.111 n=24no guard
Qwen3.5-122B-A10B (MTP)122B / 10B active18.17 t/s0.545 ±0.208 n=22guard 4/4
Nemotron-3-Super-120B-A12B120B / ~12B active17.74 t/s0.769 ±0.162 n=26guard 4/4
Laguna S 2.1 (poolside)118B / ~8B active15.53 t/s0.692 ±0.177 n=26guard 4/4
DeepSeek-V4-Flash-0731256 experts / 6 active + 1 shared (deepseek4 arch)12.35 t/s0.846 ±0.139 n=26guard 4/4
Qwen3.6-27B (staged as "qwen36-27b-mtp" — the artifact has no MTP path)27B dense10.86 t/s0.700 ±0.201 n=20guard 4/4
Deep-Thought-Posttrain (tsfrm)361.8M dense—0.357 ±0.251 n=14guard 2/4
Qwen3.8-Flash-Next (Unsloth first look + KingJones full-STRIX)~180B stored / 125B MoE (512 experts, ~6B active) + 51B N-gram/PLE tables + 4B MTP head + vision—0.792 ±0.162 n=24no guard

best measured configuration per model · every row's full fingerprint on /compare · exclusions stay visible: an excluded row is data, not an omission

Start with a question → all questions

Does KV quantisation cost quality?

conclusion

With a one-commit engine patch, mostly no: +70.3% generation speed at production depth clm-0022, 42% less energy clm-0041, and the quality cost shrinks from +39% extra conversation turns clm-0038 to +9.3% clm-0045.

Vulkan or ROCm on Strix Halo?

conclusion

Depth decides. At 32k, stock Vulkan leads ROCm +17.4% prefill and +17.1% decode at f16 KV, and a pending BF16 flash-attention patch inverts prefillclm-0046. Community reports agree on direction and disagree on magnitude clm-0044 clm-0031.

Is thinking mode worth it?

conclusion

Per-model, not global. A net negative on the 122B — it deadlocks on 2 of 5 hard tasks rather than degrading clm-0033 — so any global default is wrong for part of the field.

From the log → the full notebook

  • 2026-09-06claimOn its proposed dense-WORKER serving config (cfg-0183: UD-Q4_K_XL, the pr27311 leak-fix build, self-speculative draft-mtp n_max 2, --parallel 4, reasoning_effort low, greedy, fronted by the slotpin proxy), Qwen3.8-27B scored 91.7% on tau2-bench airline tasks 0-25 — 22 of 24 scored passed, run end-to-end through the proxy under a real 4-slot concurrent workload, with 0 empty assistant turns over 106 minutes. clm-0129
  • 2026-09-06claimThe Qwen3.8-27B dense worker earns its slot on THROUGHPUT UNDER CONCURRENCY, not single-stream speed, and only on the leak-fixed build. clm-0130
  • 2026-09-06claimThe Qwen3.8-27B worker config cost 15.18 Wh per correct answer on the tau2 0-25 run (333.87 Wh whole-session across 22 correct, mean 189.6 W, 0.46 pence at 30.3 p/kWh; eng-0278) — well under the 38.29 Wh per correct of the earlier single-slot Q8_0 run (eng-0081). clm-0131
  • 2026-09-06claimThe reproduced pwilkin/ilintar Strix Halo stack (IQ4_XS-imatrix + DFlash2 draft on the retained-PM4 runtime, cfg-0185) scored 83.3% on tau2 airline 0-25 — 20 of 24 scored passed, 4-slot, reasoning low, 2 cloud-user-sim infra errors excluded — versus the worker's 91.7% (run-0658) on the identical task set, harness (668d3bc), seed, reasoning and concurrency. clm-0132
  • 2026-09-06claimThe pwilkin candidate wins single-stream throughput decisively but only ties the worker under concurrency. clm-0133
  • 2026-09-06claimThe pwilkin candidate cost 10.26 Wh per correct answer on its tau2 0-25 run (205.2 Wh whole-session across 20 correct, mean 187 W, 0.31 pence at 30.3 p/kWh; eng-0280) — below the worker's 15.18 (eng-0278). clm-0134
  • 2026-09-06claimThe retained-PM4 runtime is a real but small, lossless dispatch optimisation — NOT the source of the pwilkin speed advantage. clm-0135
  • 2026-09-06gateqwen38-27b: benched → benched — WORKER-CONFIG CAMPAIGN (2026-09, aihydra) — a full HaloBench on a DIFFERENT serving profile from the 2026-08 screen's cfg-0066, aimed at the dense-worker role (capability + throughput beside a faster interactive model), not a re-measurement of the Q8_0 pick. clm-0129 clm-0130 clm-0131 run-0658 run-0660 cfg-0183 con-0018
  • 2026-09-06gateqwen38-27b: benched → benched — THIRD-PARTY STACK SCREEN (pwilkin/ilintar + halo-box), gate unchanged. clm-0132 clm-0133 clm-0134 clm-0135 run-0665 run-0668 con-0019
  • 2026-09-06runs1 run landed on cfg-0185 (tau2-bench-airline) run-0665
  • 2026-09-06runs1 run landed on cfg-0185 (empty-output-monitor) run-0666
  • 2026-09-06runs1 run landed on cfg-0185 (slotpin-mslot) run-0667
  • 2026-09-06runs4 runs landed on cfg-0186 (decode-depth-served) cfg-0186
  • 2026-09-06runs1 run landed on cfg-0187 (pm4-isolation) run-0672

How to read our numbers → method

measured-here wall-measured on this lab's hardware
community reported elsewhere, not yet reproduced here
vendor a vendor's own claim — never grounds for a verdict
unmeasured an honest gap, shown rather than hidden
retracted withdrawn in public, kept visible with its correction