⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.
clm-0067
measured-heremed ●●○
citable URL: https://halobench.com/records/clm-0067/ — this address never moves; the anchor /records/#clm-0067 keeps resolving
Laguna-S-2.1 Q4_K_M completed the FULL standard 26-task tau2 airline set cleanly (rc=0, not a wall-bound cut) at 0.6923 mean reward (18/26), 187 tool-call messages, 0 empty assistant turns, 0 infrastructure errors -- VALID under the SMOKE gate. Wall time 3,444s (57.4 min) against a 28,800s (8h) safety ceiling never approached. Energy: 130.66 Wh across the run window (mean 136.6 W, delta 126.5 W over the 10.1 W idle floor), whole-session, joined same-session from HA history -- 7.2588 Wh per correct answer (0.2199 pence at 30.3 p/kWh), the LOWEST Wh-per-correct-answer of any full standard-26-task tau2 comparison run on this box to date: cheaper than deepseek-v4-flash's 24.29 Wh (clm-0061), nemotron3-super's 32.43 Wh, and qwen38-27b's 38.29 Wh (clm-0057) -- roughly a third of deepseek-v4-flash's figure and under a fifth of qwen38-27b's, despite plain decode throughout (no speculation/MTP exists for this architecture) at a raw floor of 15.53 t/s (d32768, the serving depth).
METHOD -- tau2-bench airline, harness 668d3bcd135c02aa3438f987ef45735b7c163ee3 (same version as qwen38-27b/deepseek-v4-flash/nemotron3-super's full-bench arms), --num-trials 1 --seed 42 --max-steps 200 --max-concurrency 1, user simulator pinned per protocol (openrouter/anthropic/claude-haiku-4.5, temperature 0, Anthropic-only routing). Agent: local llama-server endpoint on cfg-0092 (ROCm, f16 KV, -c 32768, --parallel 1, --load-mode none, no speculation), agent sampling temperature 0.0 / max_tokens 4096. Task-ids 0-25 explicit (the standard 26-task set). Guarded by run-0328 (house 4-item guard, 4/4, against this exact serving config, run fresh immediately before this arm).
ALL NUMBERS IN THIS CLAIM VERIFIED DIRECTLY FROM RAW ARTIFACTS, not copied from the delegate's MODEL-PAGE.md draft: tool_calls/assistant_msgs/empty/mean_reward summed and averaged from tau2-full-20260816T091034Z/armcompare.json's per_task records (187/388/0/0.692307692...); duration range (46.9s task 19 to 370.6s task 23) read from the same file; termination_reason == "user_stop" on all 26 confirmed. Energy independently recomputed from aihydra's sensor.hardware_ai_hydra_energy raw ~10s HA history for the exact run-meta.jsonl window (09:11:47Z-10:09:11Z), NOT copied from the draft's "resume delegate" join -- see eng-0124. This run's own numbers reproduced the draft's headline arithmetic (7.26 Wh/correct, 5.03 Wh/task, 136.6 W mean) to within rounding; the one discrepancy found anywhere in this job's energy join was the much smaller guard window (eng-0123), not this headline figure.
FAILED TASKS (8 of 26, reward 0.0): task 1 (7 tc, 71.2s), task 5 (6 tc, 105.2s), task 7 (12 tc, 148.1s), task 10 (12 tc, 184.3s), task 14 (8 tc, 242.1s), task 16 (5 tc, 151.5s), task 21 (10 tc, 166.6s), task 23 (8 tc, 370.6s -- the longest task in the run). Every failure resolved via user_stop -- the simulator ending the conversation judging the task unresolved, not a harness timeout or agent error.
EFFICIENCY-LEADER CLAIM SCREENED AGAINST EVERY OTHER MODEL PAGE'S wh_correct_energy BEFORE PUBLISHING (per the job brief's explicit instruction): deepseek-v4-flash eng-0107 = 24.29 Wh/correct, nemotron3-super eng-0116 = 32.43 Wh/correct, qwen38-27b eng-0081 = 38.29 Wh/correct -- laguna-s-21 leads all three on the SAME protocol (full standard 26-task set, pinned haiku-4.5 simulator, same energy-join method). ONE LOWER NUMBER EXISTS ON THE SITE: qwen35-122b's eng-0034 records 5.713 Wh/correct, but that run is a 5-task SMOKE window from 2026-08-09 -- a different scale (5 vs 26 tasks) on a different, pre-"standard-26-task-convention" protocol (the standard set was adopted post-qwen38, per run-0287/run-0329's own notes) -- not a like-for-like comparison, and this claim deliberately does NOT assert an unqualified site-wide superlative on the strength of it. The comparison above is scoped to full standard-26-task tau2 runs only.
Speculation/MTP: not applicable -- llama.cpp source has no nextn/draft-tensor wiring for LLM_ARCH_LAGUNA (checked src/llama-model.cpp, src/models/laguna.cpp directly). Laguna has no MTP architecture support at all in this tree, unlike qwen38-27b's draft-mtp arm -- this candidate's efficiency lead is achieved at a raw plain-decode floor (15.53 t/s at the serving depth), the same class of result deepseek-v4-flash's clm-0061 noted for its own MoE-vs-dense comparison: a small active-parameter footprint (~8B active of 117.6B total) apparently costs less energy per correct answer than faster dense/speculated decode, even before accounting for absolute speed.
CAVEAT ON EVERY CROSS-MODEL COMPARISON ABOVE -- none of these are controlled A/B: different models, different architectures (hybrid SWA/global MoE here vs deepseek's MLA-compressed MoE, nemotron's Mamba-hybrid, qwen38's dense+MTP), different quant tiers, different backends (ROCm here and for deepseek-v4-flash; Vulkan for nemotron3-super). Reported because it is the only same-shape (26-task, same simulator, same protocol, same energy-join method) set of data points available in this programme, not because the confound list has been controlled for.