The field
Every benched model's best measured configuration, side by side, on the same hardware. A row states its own fingerprint; a number you can't trace to a run is a bug. Click a column header to re-sort — the table is complete and sorted without it.
| model | active / total | guard | dec @0 | dec @32k | pp @0 | τ² airline | turns med (all tasks) | Wh / correct | verified |
|---|---|---|---|---|---|---|---|---|---|
| Qwen3.6-35B-A3B UD-Q4_K_XL · rocm · 3653e6d · benched 2026-08-09 | 35B / 3B active | 4/4 | 51.01 | 42.45 | 1079.38 | not run | — | unmeasured | 2026-08-08 |
| gpt-oss-120b UD-Q4_K_XL · rocm · 3653e6d · benched 2026-08-08 excluded — returns stale answers from previous requests clm-0025 | 117B / ~5B active | stale answers | 55.45 | 41.70 | 450.83 | not run — guard failed | — | unmeasured | 2026-08-08 |
| Qwen3.5-122B-A10B (MTP) UD-Q4_K_M · rocm · 3653e6d · benched 2026-08-08 · draft-mtp available (NextN heads in this GGUF) | 122B / 10B active | 4/4 | 21.91 | 18.17 | 321.72 | 0.545 ±0.208 · n=22 | 22 | 5.71 n=5 smoke | 2026-08-11 |
| Nemotron-3-Super-120B-A12B UD-IQ4_XS · rocm · 3653e6d · benched 2026-08-09 quant confound resolved — reward identical at UD-Q4_K_M, +11% total turns at IQ4_XS clm-0047 | 120B / ~12B active | 4/4 | 17.51 | 17.16 | 225.36 | 0.625 ±0.237 · n=16 | 28 | unmeasured | 2026-08-11 |
decode/prefill: t/s, llama-bench, scatter published per cell on each model page · τ²-bench airline, thinking off; ± is a 95% binomial interval; n below 10 carries the smoke badge — clm-0036's noise finding made structural · turns = median conversation turns to a completed task across ALL tasks, failures included — this lab's headline efficiency metric · Wh / correct joins energy records to capability runs; the column ships even while sparse, because an honest gap in the differentiating unit beats hiding the unit · exclusions render muted with a red edge, but present: exclusions are data