⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.
clm-0082
measured-heremed ●●○
citable URL: https://halobench.com/records/clm-0082/ — this address never moves; the anchor /records/#clm-0082 keeps resolving
Qwen3-Coder-Next 80B Q8_0 completed the FULL standard 26-task tau2 AIRLINE set cleanly (rc=0, not a wall-bound cut) at 0.5385 mean reward (14/26), 192 tool-call messages, 390 assistant messages, 0 empty assistant turns, 0 infrastructure errors -- VALID under the SMOKE gate, served on Vulkan (this bench's own throughput-matrix pick, clm-0079) at plain decode (no MTP/ speculation exists in this GGUF, clm-0081). This is the LOWEST capability score of any full standard-26-task tau2 comparison run on this box, below qwen38-27b's 0.577 -- expected and unsurprising: this is a CODER model, purpose-built and marketed for agentic coding workflows (Terminal-Bench, SWE- bench-style tasks), being scored on tau2's AIRLINE customer-service domain, not its home domain. The airline benchmark is this programme's cross-model comparability instrument -- the same 26 tasks, same pinned simulator, same protocol every other full-bench candidate on this board was scored against -- and that comparability is exactly why the number is worth publishing, but it should be read as "how a coding specialist performs on an out-of-domain agentic task," not as this model's ceiling on the work it is actually built for. Energy tells the sharper story: 67.054 Wh across the run window (mean 103.69 W, delta 93.59 W over the 10.1 W idle floor), whole-session -- 4.7896 Wh per correct answer (0.1451 pence at 30.3 p/kWh), the NEW LOWEST Wh-per-correct- answer of any full standard-26-task tau2 comparison on this board, undercutting ornith-35b's previous-leading 6.4834 Wh by 26%. This candidate therefore holds the LOWEST capability score AND the LOWEST energy-per-correct-answer of this programme's full-bench series SIMULTANEOUSLY -- the inverse of ornith-35b's joint-highest result, and a genuinely new combination this board has not shown before. The efficiency figure is real and measured on the identical protocol as every other comparator, but it is not a capability win, and the verdict does not present it as one.
METHOD -- tau2-bench airline, harness 668d3bcd135c02aa3438f987ef45735b7c163ee3 (same version as every prior full-bench arm on this board), --num-trials 1 --seed 42 --max-steps 200 --max-concurrency 1, user simulator pinned per protocol (openrouter/anthropic/claude-haiku-4.5, temperature 0). Agent: local llama-server endpoint on cfg-0111 (Vulkan, default f16 KV, -c 32768, --parallel 1, --load-mode none, no speculation), agent sampling temperature 0.0 / max_tokens 4096. Task-ids 0-25 explicit (the standard 26-task set). Guarded by run-0395 (house 4-item guard, 4/4, against this exact serving config, run fresh immediately before this arm since backend is the `binary` lever the guard exists to catch on a config swap).
A 5-task SMOKE screen (run-0432, ROCm, cfg-0110) preceded this arm and scored 0.800 (4/5) -- directionally consistent (below the board's higher scorers, not a screen/full-bench reversal), 2.4205 Wh/correct on its own 5-task window (eng-0167, NOT a like-for-like comparator against the standard-26-task figure above, same scoping every prior candidate's smoke window carries).
FAILED TASKS (12 of 26, reward 0.0): tasks 0, 7, 8, 10, 11, 14, 15, 18, 20, 21, 23, 24. No single shared pattern in tool-call count or duration is visible from this bench's own arm summary alone (durations range 44s-152s across both passed and failed tasks); a transcript read would be needed to characterise the failure mode further, and this claim does not speculate beyond the aggregate. Every task resolved via user_stop (the simulator ending the conversation, judging it unresolved) -- 0 tasks cut at max_steps across all 26.
EFFICIENCY-LEADER CLAIM SCREENED AGAINST EVERY OTHER MODEL PAGE'S wh_correct_energy BEFORE PUBLISHING, per house policy: ornith-35b eng-0134 = 6.4834 Wh/correct, laguna-s-21 eng-0124 = 7.2588 Wh/correct, deepseek-v4-flash eng-0107 = 24.29 Wh/correct, nemotron3-super eng-0116 = 32.43 Wh/correct, qwen38-27b eng-0081 = 38.29 Wh/correct -- this candidate leads all five on the SAME protocol (full standard 26-task set, pinned haiku-4.5 simulator, same energy-join method, same idle baseline). ONE LOWER NUMBER EXISTS ON THE SITE: qwen35-122b's eng-0034 records 5.713 Wh/correct, but that run is a 5-task SMOKE window predating the standard-26-task convention -- a different scale on a different protocol, not a like-for-like comparison, and this claim does not assert an unqualified site-wide superlative on the strength of it (same scoping every prior leader's claim already established).
CAVEAT ON EVERY CROSS-MODEL COMPARISON ABOVE -- none of these are controlled A/B: different models, different architectures, quant tiers (Q8_0 here vs Q4_K_M/IQ3_XXS/IQ4_XS/UD-Q4_K_XL for the others), and different backends (Vulkan here and for nemotron3-super/ornith-35b; ROCm for deepseek-v4-flash and laguna-s-21). The lower Wh/correct here is plausibly explained in large part by the much lower absolute token cost of a lower-reward run (fewer turns needed to reach a user_stop on a task the model does not solve is not the same as efficient problem-solving) -- this claim does NOT assert this candidate is intrinsically more energy-efficient PER UNIT OF CAPABILITY than ornith-35b or any other comparator; it reports the measured Wh-per-correct-answer figure on the standard protocol, with the capability-score caveat stated in the same breath, exactly as required.