Models › ornith-35b

Ornith-1.0-35B (ornith-ai / deepreinforce-ai)benchedguard 4/4

35B / A3B-class MoEhybrid MoE (qwen35moe) · 10/40 full-attention + 30/40 SSM/linear-attention layers · ~35B total / A3B-class active — no MTP/speculation path in this GGUFquant held: Q8_0 · designated primary/full-bench quant (PROTOCOL_OVERRIDE, not in protocol.json quants.expected); UD-Q4_K_XL now has its own measured full-26 capability record at 22/26, mean 0.8461538461538461 (run-0499), not Q8 inheritance or Q4/Q8 equivalencefirst measured 2026-08-16latest run 2026-08-17

Verdict

A hybrid mixture-of-experts (qwen35moe architecture, 10 of 40 blocks carrying full-attention layers via full_attention_interval=4, the other 30 SSM/linear- attention with no growing KV cache — same architecture family as this lab's production 122B) served at Q8_0 — the candidate's own designated primary screening quant, a PROTOCOL_OVERRIDE against protocol.json's quants.expected list, matching the same override rationale used for qwen38-27b — with no speculation path: this job's own from-scratch GGUF header parser found zero nextn/mtp/eagle/medusa/draft tensors among all 733 tensors.

Backend is a genuine finding, not a formality: this job's own throughput matrix has Vulkan leading decode at every depth (55.70/46.22/32.37 t/s at d0/d32768/d131072 vs ROCm's 47.66/39.75/27.48) and prefill at d0/d32768 (1063.7/679.2 vs 865.3/525.2), with ROCm only marginally ahead on prefill at the deepest cell (+2.6%) — and, notably, Vulkan did NOT lose the GPU device at d131072, breaking a run of device-loss occurrences this silicon has shown on every other hybrid/MoE full-bench candidate tested at comparable depth in the cited local records (laguna-s-21, deepseek-v4-flash) clm-0070. The KV cache costs a measured 20.00 KiB/token — smaller than laguna-s-21's 48.0, with nemotron3-super's 8.00 and deepseek-v4-flash's 7.13 smaller still, so ornith's SSM-heavy layer mix (only 10 of 40 blocks full-attention) keeps its absolute figure competitive — and this figure carries HIGH confidence: reproduced bit-for-bit across two independent fresh-process probe pairs and reconciled to within 0.04 GiB against the GGUF's own model_size from first principles clm-0071.

Capability clears the standard bar decisively: the FULL 26-task tau2 airline run completed end to end (rc=0, not wall-bound) at 0.8846 mean reward (23/26), 202 tool-call messages, 0 empty turns — VALID under the SMOKE gate, and stronger than the cited same-protocol full-26-task comparisons available to this claim (deepseek-v4-flash's 0.846, nemotron3-super's 0.769, laguna-s-21's 0.6923 and qwen38-27b's 0.577) clm-0072. Energy is the standout finding a second time over: 6.4834 Wh per correct answer, 0.1964 pence at the standing tariff — lower than the cited same-protocol records for laguna-s-21, deepseek-v4-flash, nemotron3-super and qwen38-27b — paired with a raw ~46.2 t/s decode floor (Vulkan, serving depth) and no speculation at all clm-0072.

This same-protocol energy claim was screened against the then-published benched models' Wh-per-correct-answer records before being published; the one lower figure on the site at publication time (qwen35-122b, a 5-task smoke window that pre-dates this programme's standard-26-task convention) is a different scale and protocol, not a like-for-like comparator, and this verdict does not claim to beat it. The one confound worth naming plainly: Q8_0 is a HIGHER-precision quant than every cited full-bench candidate quant records (Q4_K_M/IQ3_XXS/IQ4_XS/UD-Q4_K_XL for the others) — this candidate happens to fit Q8_0 comfortably at its size, and this verdict does not claim Ornith-1.0-35B would still lead at matched quant precision. A later same-model UD-Q4_K_XL full-26 arm partly answers that caveat with measured Q4 capability (22/26, mean 0.8461538461538461, run-0499), while preserving the boundary: that Q4 result is not inherited from Q8_0 and is not a Q4/Q8 equivalence or energy-ranking claim clm-0093.

verdict written 2026-08-16 · every number above stands next to the claim chip that carries it

Best configuration

modelOrnith-1.0-35B-Q8_0.gguf · Q8_0
engineggml-org/llama.cpp 3653e6d · vulkan · host aihydra (igpu)
flags-ngl 999 -fa on -c 32768 -dev Vulkan0 --load-mode none --parallel 1 --jinja --reasoning-format deepseek -n 4096 --slots
templatenot recorded at test time
treeupstream — stock

config record cfg-0095

decode @ 0
55.70 t/s
run-0340 · CV 0%
decode @ 32k
46.22 t/s
run-0342 · CV 0%
prefill @ 0
1063.66 t/s
run-0339 · CV 0.1%
prefill @ 32k
679.22 t/s
run-0341 · CV 0.2%
τ² airline · thinking off
0.885 ±0.123
run-0345 · passed 23/26
turns to done · median (all tasks)
26
run-0345 · successes only: 26 · max 56
wall-clock to done · median (successes only)
1.7 min
run-0345 · failures excluded — they have no done
Wh per correct answer (τ², thinking off)
6.48 Wh
eng-0134 run-0345 · 0.20 p per answer

every cell generated from the record at build time · throughput cells from cfg-0096 (same build 3653e6d, fa on, f16 KV) · provenance: measured-here throughout

③aDecode against context depth

VulkanROCm
0102030405060065k131k204.8k262.1kdecode t/scontext depth (tokens)serving context 32k55.70 t/s @ depth 0 · run-0340 · CV 0% · N=346.22 t/s @ depth 32k · run-0342 · CV 0% · N=332.37 t/s @ depth 131k · run-0344 · CV 0% · N=126.53 t/s @ depth 204.8k · run-0382 · CV 0% · N=123.31 t/s @ depth 262.1k · run-0384 · CV 0% · N=147.66 t/s @ depth 0 · run-0333 · CV 1.1% · N=339.75 t/s @ depth 32k · run-0335 · CV 0.1% · N=327.48 t/s @ depth 131k · run-0337 · CV 0% · N=122.23 t/s @ depth 204.8k · run-0378 · CV 0% · N=119.37 t/s @ depth 262.1k · run-0380 · CV 0% · N=1Vulkan · 23.31ROCm · 19.37
1/3 reps per cell · max CV 1.1% · build 3653e6d / 3653e6d · Both backends completed all three original tested depths cleanly (0, 32768, 131072) — unlike laguna-s-21, no Vulkan device loss occurred at this candidate's deepest matrix cell. Vulkan leads decode at every depth and prefill at d0/d32768; ROCm is only marginally ahead on prefill at d131072 (+2.6%). Depth-ladder backfill (2026-08-17) extended ROCm to d204800 and d262144 (model-max) and Vulkan to d204800, then a same-day follow-up took Vulkan the rest of the way to d262144 too: Vulkan SURVIVED every cell cleanly through model-max (no device-lost, no wedge), with no device-loss incident across the cited whole-depth ladder, and it keeps its decode lead at every depth including the new ceiling (23.31 vs ROCm's 19.37 t/s at d262144, +20.3%). · records: run-0340 run-0342 run-0344 run-0382 run-0384 run-0333 run-0335 run-0337 run-0378 run-0380

Other configurations tested — each as a delta against best

variantΔ decodeΔ turns (paired tasks)Δ energynoterecords
Vulkan backend · same quant+18% @131k-15%Unlike laguna-s-21 (ROCm-favoured outright) but matching nemotron3-super and deepseek-v4-flash's pattern, Vulkan wins decode at every depth (+17% at d0, +16% at d32768, +18% at d131072) and prefill at d0/d32768 (+23%, +29%). ROCm is only marginally ahead on prefill at the deepest cell (+2.6%). Vulkan is the backend this bench serves the capability arm on. clm-0070
Vulkan device-loss precedent — broken this timeVulkan completed d131072 cleanly on its first attempt (of a 2-attempt cap) — no vk::DeviceLostError, no GPU device wedge. This breaks a run of device-loss occurrences on this silicon at comparable depth (laguna-s-21 clm-0065, deepseek-v4-flash clm-0059, and earlier clm-0054) — the first hybrid/MoE full-bench candidate in the cited local measurements where Vulkan survived its own deepest matrix cell. clm-0070
KV cache cost (fit)20.00 KiB/token (f16 KV, ROCm) — smaller than laguna-s-21's 48.0, nemotron3- super's 8.00 is still smaller, and deepseek-v4-flash's 7.13 smaller still. Reproduced bit-for-bit across two independent fresh-process probe pairs (the delta, not just the absolute readings, matched exactly) and reconciled from first principles against the GGUF's own architecture metadata (10 of 40 blocks carry full attention; the other 30 are SSM/linear-attention with no growing KV cache) — a high-confidence KV-cost measurement for this local record. Projected GTT use at c=131,072 is ~36.6 GiB against the 120 GiB boot window, ~83 GiB of headroom to spare. clm-0071
Wh-per-correct-answer with high tau2 reward6.4834 Wh/correct is lower than the cited same-protocol full standard-26-task tau2 comparator records on this page (laguna-s-21 7.2588, deepseek-v4-flash 24.29, nemotron3-super 32.43, qwen38-27b 38.29), and it is paired with a high capability score in the same cited set (0.8846 mean reward). One lower Wh figure exists on the site (qwen35-122b, 5.71 Wh/correct) but from a 5-task smoke window that pre-dates the standard-26-task convention — not a like-for-like comparison, and not what this row claims to beat. clm-0072
UD-Q4_K_XL measured full-26 capability armThe later HO-005 comparability arm used the acquired UD-Q4_K_XL artifact on the same stock 3653e6d6d Vulkan/f16-KV serving boundary and ran the full task-id 0..25 tau2 airline set after a fresh 4/4 guard. It measured 22/26 and mean_reward 0.8461538461538461 with 185 tool-call messages, 0 empty- argument tool calls, 0 empty assistant turns, 0 infrastructure errors, and 0 max-step cuts (run-0499). hborchestrator then joined wall-meter energy as eng-0203: 146.5378 Wh total, 6.6608 Wh/correct over the 22 correct tasks, 2.0182 pence/correct at the standing tariff. This is measured Q4_K_XL evidence only; it does not inherit the Q8_0 run-0345 result and does not establish Q4/Q8 equivalence or an energy-ranking claim. clm-0093

deltas computed at build time from the named runs, paired-task traces and metered windows — a hazard row shows words where a delta would mislead

Open questions

Provenance

bench host aihydra · vulkan · ggml-org/llama.cpp 3653e6d
discipline 3/1 reps per throughput cell · scatter published per cell (max CV 0.2%) · guard chain on every performance series
window 2026-08-16 → 2026-08-17