Home › Evidence › Records › clm-0076
⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.

clm-0076

measured-herehigh ●●●
citable URL: https://halobench.com/records/clm-0076/ — this address never moves; the anchor /records/#clm-0076 keeps resolving

qwen36-27b-mtp is the WORST full-bench candidate on this board by energy efficiency, by a wide margin, despite passing the cheap house guard twice (4/4, no cliff) and a directionally-consistent SMOKE screen (0.80 mean, 2026-08-13). The standard 26-task tau2 airline set could NOT be completed in one 8h session — rc=124 at 21/26 task attempts (20 scored, one an unrecoverable infrastructure_error with zero messages) — and needed a same-build/same-config CONTINUATION session to finish the remaining 5 tasks. COMBINED across both sessions: 25 of 26 tasks scored (task 15 excluded as an infrastructure_error, not a 0), 16 passed, mean 0.6400 — above qwen38-27b's 0.577 but well below deepseek-v4-flash (0.846), nemotron3-super (0.769), laguna-s-21 (0.6923) and ornith-35b (0.8846). This is NOT a clean full-26-task completion like every other benched candidate on this board and that caveat travels with every comparison above. The energy figure is the decisive finding. Combined tau2 wall time across both sessions was 40,836s (11.34h) for the SAME 26-task set ornith-35b completed in 4,193s (70 min) — roughly 10x longer — driven by this candidate's raw ~11-12 t/s decode floor (a third of ornith-35b's ~46 t/s, clm-0074) COMPOUNDED by repeated harness-level retries from a malformed-output failure class this session observed directly ("AssistantMessage must have either content or tool_calls", zero content and zero tool_calls in one turn). Combined energy: 2055.19 Wh across both windows (whole-session, same wall-meter method and idle baseline as every other board entry) = 128.45 Wh per correct answer. That is more than 3x qwen38-27b's previous-worst 38.29 Wh/correct, and roughly 20x ornith-35b's leading 6.48 Wh/correct — the widest efficiency gap this programme has measured between two full-bench candidates on the same protocol.

verified 2026-08-17 · volatility low
evidence run-0357 run-0358 run-0359 run-0360 eng-0141 eng-0143 eng-0144

Note — the record's own working

METHOD — tau2-bench airline, harness 668d3bcd135c02aa3438f987ef45735b7c163ee3 (same version as every other full-bench arm on this board), --num-trials 1 --seed 42 --max-steps 200 --max-concurrency 1, user simulator pinned per protocol (openrouter/anthropic/claude-haiku-4.5, temperature 0). Agent: local llama-server endpoint on cfg-0099 (ROCm, explicit f16 KV, -c 32768, --parallel 1, --load-mode none, no speculation — none exists in this GGUF, clm-0073), agent sampling temperature 0.0 / max_tokens 4096. Guarded twice, fresh each session (run-0357, run-0359), both 4/4 — no cliff detected on either check. ARM 1 (run-0358): task-ids 0-25 requested, 21 attempted before the 8h TAU2_TIMEOUT safety ceiling (rc=124). Task 15's final attempt returned reward=null, termination_reason=infrastructure_error, 0 messages — excluded from tasks_total per the SMOKE validity gate's own logic (a zero-message result cannot be scored either way), not counted as a pass or a fail. 20 scored, 14 passed, mean 0.7000. ARM 2 (run-0360): task-ids 21-25, same build/config/seed as arm 1, run in a fresh session after arm 1's timeout, house guard re-run fresh first (protocol §9's "split across sessions" safe lever — same config, bookkeeping only). rc=0, all 5 completed. 2 passed, 3 failed, mean 0.4000. COMBINED: (14+2)/(20+5) = 16/25 = 0.6400 mean reward across the 25 scored tasks. RETRY PATTERN: at least 3 distinct tasks (6, 7, 8, 10, 14, 21, 23, 25 across both arms — 8 of the 25 scored/attempted tasks, 32%) needed one or more internal harness retries before landing on a scored attempt, each retry re-running a multi-turn conversation from the start and multiplying wall time. Task 14 needed 3 retries before scoring 0.0; task 15's final (unretryable) attempt failed outright; task 23 (arm 2) accumulated roughly 90 minutes of cumulative wall time across attempts before scoring 0.0 — the single most expensive task this bench measured across every candidate to date. The malformed-output signature ("AssistantMessage must have either content or tool_calls. Got AssistantMessage") indicates the model intermittently emits a fully empty turn (no text, no tool call) at this depth/config — a genuine failure mode distinct from, and additional to, the architecture/MTP corrections in clm-0073. ENERGY, computed the same way as every other board entry: wall-meter counter-difference on sensor.hardware_ai_hydra_energy (aihydra, Home Assistant), raw ~10s-resolution history pulled fresh this session and linearly interpolated to each run-meta window, against the 10.1 W box-idle-all-empty baseline. Arm 1 (eng-0141): 1446.5888 Wh over 28,800s (mean 180.82 W, delta 170.72 W), 103.3278 Wh/correct on its own 14-passed denominator. Arm 2 (eng-0143): 608.6042 Wh over 12,036s (mean 182.04 W, delta 171.94 W), 304.3021 Wh/correct on its own thin 2-passed denominator (not representative alone — the COMBINED figure is the one that belongs on the model page). Combined: 1446.5888 + 608.6042 = 2055.1930 Wh / 16 correct = 128.4496 Wh/correct, 3.8920 pence at the standing 30.3 p/kWh tariff. EFFICIENCY-COMPARISON SCREENED AGAINST EVERY OTHER BENCHED MODEL PAGE'S wh_correct_energy BEFORE PUBLISHING, per house policy: ornith-35b eng-0134 = 6.4834 Wh/correct, laguna-s-21 eng-0124 = 7.2588, deepseek-v4-flash eng-0107 = 24.29, nemotron3-super eng-0116 = 32.43, qwen38-27b eng-0081 = 38.29 — qwen36-27b-mtp's 128.45 Wh/correct trails all five, and by a materially wider margin than separates any adjacent pair among them (the next-widest gap, qwen38-27b to nemotron3-super, is ~1.2x; this candidate to qwen38-27b is >3.3x). CAVEAT ON THE CROSS-MODEL COMPARISON: not a controlled A/B — different models, different quant tiers, different backends (ROCm here; Vulkan for ornith-35b and nemotron3-super, ROCm for deepseek-v4-flash and laguna-s-21). Reported because it is the only same-shape (26-task set attempted, same simulator, same protocol, same energy-join method) set of data points available in this programme. The BOUND-CUT status of this run (25/26 scored rather than a clean 26/26) is the one confound specific to this candidate among the six, and is why every number above states its own denominator explicitly rather than assuming "26" the way the other five pages can.

Cited by — computed at build time, never stored

model pages qwen36-27b-mtp