Home › Evidence › Records › clm-0072

clm-0072

measured-heremed ●●○
citable URL: https://halobench.com/records/clm-0072/ — this address never moves; the anchor /records/#clm-0072 keeps resolving

Ornith-1.0-35B Q8_0 completed the FULL standard 26-task tau2 airline set cleanly (rc=0, not a wall-bound cut) at 0.8846 mean reward (23/26), 202 tool-call messages, 0 empty assistant turns, 0 infrastructure errors -- VALID under the SMOKE gate. This is the NEW LEADER among full standard-26-task tau2 comparisons run on this box, ahead of deepseek-v4-flash's 0.846, nemotron3-super's 0.769, laguna-s-21's 0.6923 and qwen38-27b's 0.577. Wall time 4,193s (69.9 min) against a 28,800s (8h) safety ceiling never approached. Energy: 149.12 Wh across the run window (mean 128.0 W, delta 117.9 W over the 10.1 W idle floor), whole-session, joined same-session from HA history -- 6.4834 Wh per correct answer (0.1964 pence at 30.3 p/kWh), ALSO the lowest Wh-per-correct-answer of any full standard-26-task tau2 comparison run on this box: cheaper than laguna-s-21's 7.2588 Wh (clm-0067), deepseek-v4-flash's 24.29 Wh (clm-0061), nemotron3-super's 32.43 Wh, and qwen38-27b's 38.29 Wh (clm-0057) -- achieved on Vulkan (this bench's own throughput-matrix pick, clm-0070) at plain decode (no MTP/speculation exists in this GGUF, clm-0070's own tensor scan) against a raw ~46.2 t/s decode floor at the serving depth (d32768) -- both the highest capability score AND the lowest energy-per-correct-answer of any full-26-task comparison on this box, a joint win this programme has not seen before (every prior leader on one metric has trailed on the other).

verified 2026-08-16 · volatility medium
evidence run-0345

Note — the record's own working

METHOD -- tau2-bench airline, harness 668d3bcd135c02aa3438f987ef45735b7c163ee3 (same version as qwen38-27b/deepseek-v4-flash/nemotron3-super/laguna-s-21's full-bench arms), --num-trials 1 --seed 42 --max-steps 200 --max-concurrency 1, user simulator pinned per protocol (openrouter/anthropic/claude-haiku-4.5, temperature 0). Agent: local llama-server endpoint on cfg-0095 (Vulkan, default f16 KV, -c 32768, --parallel 1, --load-mode none, no speculation), agent sampling temperature 0.0 / max_tokens 4096. Task-ids 0-25 explicit (the standard 26-task set). Guarded by run-0338 (house 4-item guard, 4/4, against this exact serving config -- Vulkan, not the ROCm config the screen's run-0330 guard covered -- run fresh immediately before this arm, since backend is itself the `binary` lever this guard exists to catch on a config swap). All numbers computed directly from raw artifacts (results.json, run-meta.jsonl, tier3-armcompare.py's own arm summary), no delegate/draft to reconcile against for this job. FAILED TASKS (3 of 26, reward 0.0): task 7 (11 tc, 156.2s), task 14 (8 tc, 294.7s), task 23 (11 tc, 332.4s -- the longest task in the run, and it succeeded on retry 1 per the harness's own retry log before landing at reward 0.0 on the scored attempt). Every failure resolved via user_stop -- the simulator ending the conversation judging the task unresolved, not a harness timeout or agent error. cut_at_max_steps: 0 across all 26 tasks. EFFICIENCY-LEADER CLAIM SCREENED AGAINST EVERY OTHER MODEL PAGE'S wh_correct_energy BEFORE PUBLISHING, per house policy: laguna-s-21 eng-0124 = 7.2588 Wh/correct, deepseek-v4-flash eng-0107 = 24.29 Wh/correct, nemotron3-super eng-0116 = 32.43 Wh/correct, qwen38-27b eng-0081 = 38.29 Wh/correct -- ornith-35b leads all four on the SAME protocol (full standard 26-task set, pinned haiku-4.5 simulator, same energy-join method, same idle baseline). ONE LOWER NUMBER EXISTS ON THE SITE: qwen35-122b's eng-0034 records 5.713 Wh/correct, but that run is a 5-task SMOKE window predating the standard-26-task convention -- a different scale on a different protocol, not a like-for-like comparison, and this claim does not assert an unqualified site-wide superlative on the strength of it (same scoping laguna-s-21's clm-0067 already established). Speculation/MTP: not applicable -- this job's own from-scratch GGUF header parser found zero nextn/mtp/eagle/medusa/draft tensors among all 733 tensors in the Q8_0 GGUF (clm-0070). Ornith-1.0-35B has no speculative-decode path exercised in this bench -- this candidate's dual lead (capability AND energy) is achieved at a raw plain-decode floor (46.22 t/s at the serving depth, Vulkan), the highest of any full-bench candidate's serving-depth decode figure in this programme to date. CAVEAT ON EVERY CROSS-MODEL COMPARISON ABOVE -- none of these are controlled A/B: different models, different architectures, quant tiers (Q8_0 here vs Q4_K_M/IQ3_XXS/ IQ4_XS/UD-Q4_K_XL for the others -- a HIGHER-precision quant than every other full-bench candidate in this programme, per protocol §9's "higher quant does not fit, so test lower" grey-zone logic run in reverse: this candidate specifically fits Q8_0 comfortably), and different backends (Vulkan here and for nemotron3-super; ROCm for deepseek-v4-flash and laguna-s-21). Reported because it is the only same-shape (26-task, same simulator, same protocol, same energy-join method) set of data points available in this programme, not because the confound list has been controlled for -- and the quant-tier gap in particular means this is NOT presented as evidence that Ornith-1.0-35B is intrinsically more capable or efficient than the other candidates at matched precision.

Cited by — computed at build time, never stored

model pages ornith-35b
candidate gate history ornith-35b