Home › Evidence › Records › clm-0124
⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.

clm-0124

measured-herehigh ●●●superseded
citable URL: https://halobench.com/records/clm-0124/ — this address never moves; the anchor /records/#clm-0124 keeps resolving

Qwen3.8-Flash-Next UD-Q4_K_XL (Qwen4-arch preview MoE: 125B MoE / ~6B active + 51B N-gram/PLE tables + vision) on the Unsloth qwen4exp Vulkan build (cfg-0177, commit 250b6144, gfx1151/RADV) — first full benchmark. ADMITTED CLAIM 1 (capability leader): on the standard tau2 airline full-26 suite (seed 42, claude-haiku-4.5 simulator, deterministic reward) it scored mean_reward 0.9231, 24/26 tasks at reward 1.0 (run-0632). This LEADS the tau2 airline board — the prior best was Ornith-1.0-35B UD-Q4_K_XL at 22/26 (0.846). Tool-calling was clean: 184 tool calls, 0 empty-argument calls, 1 tool-error message, 24.7 messages/task. Failures = tasks 7 and 20. ADMITTED CLAIM 2 (energy leader per correct answer): wall-metered join (eng-0275, HA counter-difference) = 256.08 Wh total over 6605 s, 237.55 Wh active above the 10.1 W idle floor, = 9.90 Wh per correct answer (0.300 p @ 30.3 p/kWh). This is ~4x cheaper per correct answer than the ho003 stock-f16 tau2 control (42.06 Wh/correct), because it solves more than twice as many tasks in less wall time. ADMITTED CLAIM 3 (footprint): GTT-resident 76.7 GiB with ~43 GiB free for KV. The 26.8 GiB n-gram/PLE table (per_layer_token_embd, iq4_nl) is placed on CPU automatically by the Vulkan backend regardless of -ot, so it never occupies GTT; KV scales ~48 MiB per 1000 tokens (hybrid QSA attention). Performance: prefill 277-313 tok/s at 5-8K, decode ~22 tok/s short falling to 8.5 at 131K; TTFT ~2.9 s. BLOCKED CLAIMS (explicitly NOT made): single trial (N=1), no variance bound. reasoning_effort=low, not the model's default xhigh — a higher-reasoning result is untested and may differ (both tasks 7 and 20 may be reasoning-recoverable). The runtime is a draft/WIP PR (#27742), not a released/immutable build; the GGUF's advertised LICENSE artifact 404s and conversion-base provenance is unpinned, so this is a LAB result with NO weight redistribution or public/production recommendation. MTP speculative decoding is unavailable (the GGUF lacks MTP head layers) and n-gram speculation did not help this quant. 262K context hits a Vulkan workgroup-count assertion (cap -c <= 262140). No UD-Q2 control or cross-host pair run yet.

superseded by clm-0125 — the corrected statement lives there; this record keeps its URL and full text
verified 2026-08-27 · volatility medium
evidence cfg-0177 run-0632 eng-0275
corrections 2026-08-28: The first-look observations remain published under their original record ids, but this claim no longer carries current capability-leader, energy-per-correct, production-throughput, or role-fit authority. The run lacked the protocol screen, formal guard, repetitions, retained exact full-run serving transport, and an unambiguous admitted energy denominator. clm-0125 records the interim evidence boundary while a protocol-valid intended-config supersession is prepared.

Note — the record's own working

First-look benchmark of Qwen3.8-Flash-Next, run under lab custody of AI Hydra. Capability + energy admitted; the WIP- runtime / license / N=1 / reasoning-low caveats are load-bearing and are carried into the candidate page verdict.

Cited by — computed at build time, never stored

claims clm-0125
candidate gate history qwen38-flash-next