The field
Every benched model's best measured configuration, side by side, on the same hardware. A row states its own fingerprint; a number you can't trace to a run is a bug. Click a column header to re-sort — the table is complete and sorted without it.
| model | active / total | guard | dec @0 | dec @32k | pp @0 | τ² airline | turns med (all tasks) | Wh / correct | verified |
|---|---|---|---|---|---|---|---|---|---|
| Ornith-1.0-35B (ornith-ai / deepreinforce-ai) Q8_0 · vulkan · 3653e6d · benched 2026-08-16 | 35B / A3B-class MoE | 4/4 | 55.70 | 46.22 | 1063.66 | 0.885 ±0.123 · n=26 | 26 | 6.48 | 2026-08-17 |
| Qwen3.6-35B-A3B UD-Q4_K_XL · rocm · 3653e6d · benched 2026-08-09 | 35B / 3B active | 4/4 | 51.01 | 42.45 | 1079.38 | not run | — | unmeasured | 2026-08-19 |
| gpt-oss-120b UD-Q4_K_XL · rocm · 3653e6d · benched 2026-08-08 excluded — returns stale answers from previous requests clm-0025 | 117B / ~5B active | stale answers | 55.45 | 41.70 | 450.83 | not run — guard failed | — | unmeasured | 2026-08-13 |
| Qwen3-Coder-Next (Qwen) Q8_0 · vulkan · 3653e6d · benched 2026-08-17 Scored on tau2's airline domain, not this model's home domain (agentic coding) — 0.5385 mean reward (14/26) is weak beside the cited full-bench airline comparisons, expected for a coding specialist on an out-of-domain customer-service task, not a verdict on its coding ability.
clm-0082 | 80B total / ~3B active (A3B-class MoE, 512 experts / 10 active + 1 shared) | 4/4 | 44.13 | 36.97 | 614.00 | 0.538 ±0.192 · n=26 | 30 | 4.79 | 2026-08-17 |
| Ling-3.0-flash Q4_K_M · rocm · 7077abb · benched 2026-08-18 | 124B / 5.1B active | 4/4 | 36.08 | 32.89 | 409.55 | 0.500 ±0.192 · n=26 | — | 5.32 | 2026-08-19 |
| Qwen3.8-27B UD-Q4_K_XL · rocm · c530ea7 · benched 2026-08-15 | 27B dense | no guard | — | 19.23 | — | 0.917 ±0.111 · n=24 | — | 15.18 | 2026-09-06 |
| Qwen3.5-122B-A10B (MTP) UD-Q4_K_M · rocm · 3653e6d · benched 2026-08-08 · draft-mtp available (NextN heads in this GGUF) | 122B / 10B active | 4/4 | 21.91 | 18.17 | 321.72 | 0.545 ±0.208 · n=22 | 22 | 5.71 n=5 smoke | 2026-08-23 |
| Nemotron-3-Super-120B-A12B UD-Q4_K_M · vulkan · 3653e6d · benched 2026-08-09 | 120B / ~12B active | 4/4 | 18.20 | 17.74 | 213.73 | 0.769 ±0.162 · n=26 | — | 32.43 | 2026-08-19 |
| Laguna S 2.1 (poolside) Q4_K_M · rocm · 3653e6d · benched 2026-08-16 | 118B / ~8B active | 4/4 | 23.68 | 15.53 | 300.67 | 0.692 ±0.177 · n=26 | — | 7.26 | 2026-08-17 |
| DeepSeek-V4-Flash-0731 UD-IQ3_XXS · rocm · 3653e6d · benched 2026-08-16 | 256 experts / 6 active + 1 shared (deepseek4 arch) | 4/4 | 15.30 | 12.35 | 140.80 | 0.846 ±0.139 · n=26 | 24 | 24.29 | 2026-08-18 |
| Qwen3.6-27B (staged as "qwen36-27b-mtp" — the artifact has no MTP path) Q4_K_M · rocm · 3653e6d · benched 2026-08-16 | 27B dense | 4/4 | 12.08 | 10.86 | 357.84 | 0.700 ±0.201 · n=20 | — | 128.45 | 2026-08-17 |
| Deep-Thought-Posttrain (tsfrm) F16 · rocm · 3653e6d · benched 2026-08-17 | 361.8M dense | 2/4 | — | — | — | 0.357 ±0.251 · n=14 | 8 | 4.55 | 2026-08-17 |
| Qwen3.8-Flash-Next (Unsloth first look + KingJones full-STRIX) Q4_0_ROCmFP4_STRIX · HIP/gfx1151 · 36e9acd · benched 2026-08-27 | ~180B stored / 125B MoE (512 experts, ~6B active) + 51B N-gram/PLE tables + 4B MTP head + vision | no guard | — | — | — | 0.792 ±0.162 · n=24 | — | 20.54 | 2026-08-30 |
decode/prefill: t/s, llama-bench, scatter published per cell on each model page · τ²-bench airline; thinking mode is per model, stated on each model page's metric card; ± is a 95% binomial interval; n below 10 carries the smoke badge — clm-0036's noise finding made structural · turns = median conversation turns to a completed task across ALL tasks, failures included — this lab's headline efficiency metric · Wh / correct joins energy records to capability runs; the column ships even while sparse, because an honest gap in the differentiating unit beats hiding the unit · exclusions render muted with a red edge, but present: exclusions are data