Ornith-1.0-35B (ornith-ai / deepreinforce-ai)benchedguard 4/4
②Verdict
A hybrid mixture-of-experts (qwen35moe architecture, 10 of 40 blocks carrying full-attention layers via full_attention_interval=4, the other 30 SSM/linear- attention with no growing KV cache — same architecture family as this lab's production 122B) served at Q8_0 — the candidate's own designated primary screening quant, a PROTOCOL_OVERRIDE against protocol.json's quants.expected list, matching the same override rationale used for qwen38-27b — with no speculation path: this job's own from-scratch GGUF header parser found zero nextn/mtp/eagle/medusa/draft tensors among all 733 tensors.
Backend is a genuine finding, not a formality: this job's own throughput matrix has Vulkan leading decode at every depth (55.70/46.22/32.37 t/s at d0/d32768/d131072 vs ROCm's 47.66/39.75/27.48) and prefill at d0/d32768 (1063.7/679.2 vs 865.3/525.2), with ROCm only marginally ahead on prefill at the deepest cell (+2.6%) — and, notably, Vulkan did NOT lose the GPU device at d131072, breaking a run of device-loss occurrences this silicon has shown on every other hybrid/MoE full-bench candidate tested at comparable depth in the cited local records (laguna-s-21, deepseek-v4-flash) clm-0070. The KV cache costs a measured 20.00 KiB/token — smaller than laguna-s-21's 48.0, with nemotron3-super's 8.00 and deepseek-v4-flash's 7.13 smaller still, so ornith's SSM-heavy layer mix (only 10 of 40 blocks full-attention) keeps its absolute figure competitive — and this figure carries HIGH confidence: reproduced bit-for-bit across two independent fresh-process probe pairs and reconciled to within 0.04 GiB against the GGUF's own model_size from first principles clm-0071.
Capability clears the standard bar decisively: the FULL 26-task tau2 airline run completed end to end (rc=0, not wall-bound) at 0.8846 mean reward (23/26), 202 tool-call messages, 0 empty turns — VALID under the SMOKE gate, and stronger than the cited same-protocol full-26-task comparisons available to this claim (deepseek-v4-flash's 0.846, nemotron3-super's 0.769, laguna-s-21's 0.6923 and qwen38-27b's 0.577) clm-0072. Energy is the standout finding a second time over: 6.4834 Wh per correct answer, 0.1964 pence at the standing tariff — lower than the cited same-protocol records for laguna-s-21, deepseek-v4-flash, nemotron3-super and qwen38-27b — paired with a raw ~46.2 t/s decode floor (Vulkan, serving depth) and no speculation at all clm-0072.
This same-protocol energy claim was screened against the then-published benched models' Wh-per-correct-answer records before being published; the one lower figure on the site at publication time (qwen35-122b, a 5-task smoke window that pre-dates this programme's standard-26-task convention) is a different scale and protocol, not a like-for-like comparator, and this verdict does not claim to beat it. The one confound worth naming plainly: Q8_0 is a HIGHER-precision quant than every cited full-bench candidate quant records (Q4_K_M/IQ3_XXS/IQ4_XS/UD-Q4_K_XL for the others) — this candidate happens to fit Q8_0 comfortably at its size, and this verdict does not claim Ornith-1.0-35B would still lead at matched quant precision. A later same-model UD-Q4_K_XL full-26 arm partly answers that caveat with measured Q4 capability (22/26, mean 0.8461538461538461, run-0499), while preserving the boundary: that Q4 result is not inherited from Q8_0 and is not a Q4/Q8 equivalence or energy-ranking claim clm-0093.
③Best configuration
| model | Ornith-1.0-35B-Q8_0.gguf · Q8_0 |
| engine | ggml-org/llama.cpp 3653e6d · vulkan · host aihydra (igpu) |
| flags | -ngl 999 -fa on -c 32768 -dev Vulkan0 --load-mode none --parallel 1 --jinja --reasoning-format deepseek -n 4096 --slots |
| template | not recorded at test time |
| tree | upstream — stock |
config record cfg-0095
every cell generated from the record at build time · throughput cells from cfg-0096 (same build 3653e6d, fa on, f16 KV) · provenance: measured-here throughout
③aDecode against context depth
④Other configurations tested — each as a delta against best
| variant | Δ decode | Δ turns (paired tasks) | Δ energy | note | records |
|---|---|---|---|---|---|
| Vulkan backend · same quant | +18% @131k | — | -15% | Unlike laguna-s-21 (ROCm-favoured outright) but matching nemotron3-super and deepseek-v4-flash's pattern, Vulkan wins decode at every depth (+17% at d0, +16% at d32768, +18% at d131072) and prefill at d0/d32768 (+23%, +29%). ROCm is only marginally ahead on prefill at the deepest cell (+2.6%). Vulkan is the backend this bench serves the capability arm on. | clm-0070 |
| Vulkan device-loss precedent — broken this time | — | — | — | Vulkan completed d131072 cleanly on its first attempt (of a 2-attempt cap) — no vk::DeviceLostError, no GPU device wedge. This breaks a run of device-loss occurrences on this silicon at comparable depth (laguna-s-21 clm-0065, deepseek-v4-flash clm-0059, and earlier clm-0054) — the first hybrid/MoE full-bench candidate in the cited local measurements where Vulkan survived its own deepest matrix cell. | clm-0070 |
| KV cache cost (fit) | — | — | — | 20.00 KiB/token (f16 KV, ROCm) — smaller than laguna-s-21's 48.0, nemotron3- super's 8.00 is still smaller, and deepseek-v4-flash's 7.13 smaller still. Reproduced bit-for-bit across two independent fresh-process probe pairs (the delta, not just the absolute readings, matched exactly) and reconciled from first principles against the GGUF's own architecture metadata (10 of 40 blocks carry full attention; the other 30 are SSM/linear-attention with no growing KV cache) — a high-confidence KV-cost measurement for this local record. Projected GTT use at c=131,072 is ~36.6 GiB against the 120 GiB boot window, ~83 GiB of headroom to spare. | clm-0071 |
| Wh-per-correct-answer with high tau2 reward | — | — | — | 6.4834 Wh/correct is lower than the cited same-protocol full standard-26-task tau2 comparator records on this page (laguna-s-21 7.2588, deepseek-v4-flash 24.29, nemotron3-super 32.43, qwen38-27b 38.29), and it is paired with a high capability score in the same cited set (0.8846 mean reward). One lower Wh figure exists on the site (qwen35-122b, 5.71 Wh/correct) but from a 5-task smoke window that pre-dates the standard-26-task convention — not a like-for-like comparison, and not what this row claims to beat. | clm-0072 |
| UD-Q4_K_XL measured full-26 capability arm | — | — | — | The later HO-005 comparability arm used the acquired UD-Q4_K_XL artifact on the same stock 3653e6d6d Vulkan/f16-KV serving boundary and ran the full task-id 0..25 tau2 airline set after a fresh 4/4 guard. It measured 22/26 and mean_reward 0.8461538461538461 with 185 tool-call messages, 0 empty- argument tool calls, 0 empty assistant turns, 0 infrastructure errors, and 0 max-step cuts (run-0499). hborchestrator then joined wall-meter energy as eng-0203: 146.5378 Wh total, 6.6608 Wh/correct over the 22 correct tasks, 2.0182 pence/correct at the standing tariff. This is measured Q4_K_XL evidence only; it does not inherit the Q8_0 run-0345 result and does not establish Q4/Q8 equivalence or an energy-ranking claim. | clm-0093 |
deltas computed at build time from the named runs, paired-task traces and metered windows — a hazard row shows words where a delta would mislead
⑤Open questions
- ANSWERED 2026-08-17: Vulkan's clean run was re-tested at d204800 (depth-ladder backfill) and survived again, then extended the same day to d262144 (this candidate's real model-max) — also clean, no device-lost, no wedge. Vulkan has now completed its entire five-depth ladder (0/32768/131072/204800/262144) on this candidate without a single device-loss incident, and keeps its decode lead over ROCm at every depth including the new ceiling clm-0070.
- Root cause of the 3 failed tasks (7, 14, 23) — no shared pattern in tool-call count or duration is visible from this bench's own summary; a transcript read would be needed to say more, and task 23 is notable for succeeding on an internal harness retry before landing at reward 0.0 on the scored attempt clm-0072.
- Whether the capability-and-energy finding survives at matched quant precision — Q8_0 is a higher-precision quant than every other full-bench candidate compared here. PARTIALLY ANSWERED 2026-08-20: the UD-Q4_K_XL comparability arm acquired alongside this quant first passed a bounded preflight/smoke and one model-max d262144 throughput cell (cfg-0129, run-0494..run-0497, clm-0092), then a reviewer-admitted full-26 tau2 successor measured 22/26 and mean_reward 0.8461538461538461 at Q4_K_XL (run-0499, clm-0093), with joined wall-meter energy of 146.5378 Wh total / 6.6608 Wh per correct task for that Q4 arm (eng-0203). That answers measured Q4 capability at the serving-depth protocol boundary, but the result still does not establish Q4/Q8 equivalence or a Q4-vs-Q8 energy ranking.
- The vision claim remains unconfirmed — the official repo ships no mmproj file and the README makes no mention of vision/multimodal capability; unsloth's mmproj-F16 was acquired but never screened (same open gap the candidate record's own note names, and out of scope for this throughput/capability bench).
- Whether this candidate's throughput/capability lead holds at a deeper serving context — this bench's capability run stayed at the 32,768-token band the 2026-08-16 screen already characterised; the matrix's own d131072 cell only measured throughput, not a capability or energy result at that depth.