Home › Compare
⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.

The field

Every benched model's best measured configuration, side by side, on the same hardware. A row states its own fingerprint; a number you can't trace to a run is a bug. Click a column header to re-sort — the table is complete and sorted without it.

held constant across every row
aihydra · Strix Halo 128 GB · vulkan/rocm/HIP/gfx1151 · llama.cpp 3653e6d/7077abb/c530ea7/36e9acd · fa on · f16 KV · --parallel 1 · reps per cell published · page cache dropped between models · guard before every series
provisional — read before citing
One arm still ran self-play (agent = user simulator) and is not comparable cross-model — Qwen3.5-122B-A10B clm-0043; the pinned-simulator re-run is outstanding clm-0048. The scatter marks it as a caution diamond, off the frontier; the rest ran the pinned Haiku simulator. Within-model deltas are unaffected.
capability vs cost — the efficiency frontier
on the efficiency frontierdominatedself-play — not comparable
1002005001000200050000.00.20.40.60.81.0time to correct (s, log) — whole-session ÷ correct, lower is better ↑τ² airline score (pass rate) — higher is better →Ornith-1.0-35B · score 0.885 · 182 · run-0345Qwen3-Coder-Next · score 0.538 · 166 · run-0433Ling-3.0-flash · score 0.500 · 139 · run-0439Qwen3.8-27B · score 0.917 · 288 · run-0658Qwen3.5-122B-A10B · score 0.545 · 1371 · SELF-PLAY (confounded) · run-0100Nemotron-3-Super-120B-A12B · score 0.769 · 798 · run-0316Laguna S 2.1 · score 0.692 · 191 · run-0329DeepSeek-V4-Flash-0731 · score 0.846 · 509 · run-0287Qwen3.6-27B · score 0.700 · 2057 · run-0358Deep-Thought-Posttrain · score 0.357 · 299 · run-0372Qwen3.8-Flash-Next · score 0.792 · 542 · run-0639Qwen3.6-27BQwen3.6-27BQwen3.5-122B-A10B ·self-playQwen3.5-122B-A10BNemotron-3-Super-120B-A12BNemotron-3-Super-120B-A12BQwen3.8-Flash-NextQwen3.8-Flash-NextDeepSeek-V4-Flash-0731DeepSeek-V4-Flash-0731Deep-Thought-PosttrainDeep-Thought-PosttrainQwen3.8-27BQwen3.8-27BLaguna S 2.1Laguna S 2.1Ornith-1.0-35BOrnith-1.0-35BQwen3-Coder-NextQwen3-Coder-NextLing-3.0-flashLing-3.0-flash
the dashed line is the Pareto frontier — no model beats a frontier model on both capability and time · records: run-0345 run-0433 run-0439 run-0658 run-0100 run-0316 run-0329 run-0287 run-0358 run-0372 run-0639
raw decode speed — the field, ranked
decode t/s @ 32k · best config per model01020304050ornith-35bornith-35bQ8_0Q8_046.22 · ornith-35b · Q8_0 · CV 0% · N=3 · run-034246.22qwen36-35bqwen36-35bUD-Q4_K_XLUD-Q4_K_XL42.45 · qwen36-35b · UD-Q4_K_XL · CV 0.1% · N=3 · run-003542.45gpt-oss-120bgpt-oss-120bguard failed — stale a…guard failed — stale answers41.70 · gpt-oss-120b · guard failed — stale answers · CV 0% · N=3 · run-004941.70qwen3-coder-next-80bqwen3-coder-next-80bQ8_0Q8_036.97 · qwen3-coder-next-80b · Q8_0 · CV 0% · N=1 · run-042736.97ling-30-flashling-30-flashQ4_K_MQ4_K_M32.89 · ling-30-flash · Q4_K_M · CV 0.2% · N=3 · run-043832.89qwen38-27bqwen38-27bguard failed — see runguard failed — see run19.23 · qwen38-27b · guard failed — see run · run-066119.23qwen35-122bqwen35-122bUD-Q4_K_MUD-Q4_K_M18.17 · qwen35-122b · UD-Q4_K_M · CV 0.1% · N=3 · run-001918.17nemotron3-supernemotron3-superUD-Q4_K_MUD-Q4_K_M17.74 · nemotron3-super · UD-Q4_K_M · CV 0% · N=3 · run-031517.74laguna-s-21laguna-s-21Q4_K_MQ4_K_M15.53 · laguna-s-21 · Q4_K_M · run-032015.53deepseek-v4-flashdeepseek-v4-flashUD-IQ3_XXSUD-IQ3_XXS12.35 · deepseek-v4-flash · UD-IQ3_XXS · CV 0% · N=3 · run-029112.35qwen36-27b-mtpqwen36-27b-mtpQ4_K_MQ4_K_M10.86 · qwen36-27b-mtp · Q4_K_M · CV 0% · N=3 · run-034910.86
1/3 reps per cell · max CV 0.2% · build 3653e6d / 3653e6d / 7077abb / c530ea7 · the guard-failed bar renders dashed, not deleted — exclusions are data · records: run-0342 run-0035 run-0049 run-0427 run-0438 run-0661 run-0019 run-0315 run-0320 run-0291 run-0349
modelactive / totalguarddec @0dec @32kpp @0τ² airlineturns med (all tasks)Wh / correctverified
Ornith-1.0-35B (ornith-ai / deepreinforce-ai)
Q8_0 · vulkan · 3653e6d · benched 2026-08-16
35B / A3B-class MoE4/455.7046.221063.660.885 ±0.123 · n=26266.482026-08-17
Qwen3.6-35B-A3B
UD-Q4_K_XL · rocm · 3653e6d · benched 2026-08-09
35B / 3B active4/451.0142.451079.38not run—unmeasured2026-08-19
gpt-oss-120b
UD-Q4_K_XL · rocm · 3653e6d · benched 2026-08-08
excluded — returns stale answers from previous requests clm-0025
117B / ~5B activestale answers55.4541.70450.83not run — guard failed—unmeasured2026-08-13
Qwen3-Coder-Next (Qwen)
Q8_0 · vulkan · 3653e6d · benched 2026-08-17
Scored on tau2's airline domain, not this model's home domain (agentic coding) — 0.5385 mean reward (14/26) is weak beside the cited full-bench airline comparisons, expected for a coding specialist on an out-of-domain customer-service task, not a verdict on its coding ability. clm-0082
80B total / ~3B active (A3B-class MoE, 512 experts / 10 active + 1 shared)4/444.1336.97614.000.538 ±0.192 · n=26304.792026-08-17
Ling-3.0-flash
Q4_K_M · rocm · 7077abb · benched 2026-08-18
corrected post-PR-26608 GGUF; ROCm plain decode recommended; bounded MTP activates but does not improve speed clm-0015 clm-0083 clm-0086
124B / 5.1B active4/436.0832.89409.550.500 ±0.192 · n=26—5.322026-08-19
Qwen3.8-27B
UD-Q4_K_XL · rocm · c530ea7 · benched 2026-08-15
27B denseno guard—19.23—0.917 ±0.111 · n=24—15.182026-09-06
Qwen3.5-122B-A10B (MTP)
UD-Q4_K_M · rocm · 3653e6d · benched 2026-08-08 · draft-mtp available (NextN heads in this GGUF)
122B / 10B active4/421.9118.17321.720.545 ±0.208 · n=22225.71 n=5 smoke2026-08-23
Nemotron-3-Super-120B-A12B
UD-Q4_K_M · vulkan · 3653e6d · benched 2026-08-09
fair-quant, pinned-simulator bench (UD-Q4_K_M) — UD-IQ4_XS never ran under the pinned simulator, so no matched quant-vs-quant number exists at this bench's headline metric clm-0047 clm-0064
120B / ~12B active4/418.2017.74213.730.769 ±0.162 · n=26—32.432026-08-19
Laguna S 2.1 (poolside)
Q4_K_M · rocm · 3653e6d · benched 2026-08-16
118B / ~8B active4/423.6815.53300.670.692 ±0.177 · n=26—7.262026-08-17
DeepSeek-V4-Flash-0731
UD-IQ3_XXS · rocm · 3653e6d · benched 2026-08-16
256 experts / 6 active + 1 shared (deepseek4 arch)4/415.3012.35140.800.846 ±0.139 · n=262424.292026-08-18
Qwen3.6-27B (staged as "qwen36-27b-mtp" — the artifact has no MTP path)
Q4_K_M · rocm · 3653e6d · benched 2026-08-16
27B dense4/412.0810.86357.840.700 ±0.201 · n=20—128.452026-08-17
Deep-Thought-Posttrain (tsfrm)
F16 · rocm · 3653e6d · benched 2026-08-17
SMOKE-INVALID (zero tool calls across all 26 simulations). Do not compare this candidate's tau2 mean or Wh-per-correct-answer figure to any other model on this site — both are computed but flagged invalid by the same gate that disqualified npu-lfm2's vacuous 1.00. clm-0078 run-0372 eng-0146
361.8M dense2/4———0.357 ±0.251 · n=1484.552026-08-17
Qwen3.8-Flash-Next (Unsloth first look + KingJones full-STRIX)
Q4_0_ROCmFP4_STRIX · HIP/gfx1151 · 36e9acd · benched 2026-08-27
The KingJones full-STRIX ROCmFP4/HIP result is bounded, not merged with the historical Unsloth series, and not publisher-equivalent. Its incomplete PLE adapter, engine scatter, zero-byte outputs and full-window refusal remain visible exclusions. clm-0127 con-0017
~180B stored / 125B MoE (512 experts, ~6B active) + 51B N-gram/PLE tables + 4B MTP head + visionno guard———0.792 ±0.162 · n=24—20.542026-08-30

decode/prefill: t/s, llama-bench, scatter published per cell on each model page · τ²-bench airline; thinking mode is per model, stated on each model page's metric card; ± is a 95% binomial interval; n below 10 carries the smoke badge — clm-0036's noise finding made structural · turns = median conversation turns to a completed task across ALL tasks, failures included — this lab's headline efficiency metric · Wh / correct joins energy records to capability runs; the column ships even while sparse, because an honest gap in the differentiating unit beats hiding the unit · exclusions render muted with a red edge, but present: exclusions are data

Not yet in the field:1 listed0 acquired11 screened0 blocked4 NPU-lane (separate artifacts — numbers never transfer)→ the candidates board, with every gate reason