Home › Models
⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.

Models

One page per benched model: best configuration fully specified, every other configuration a delta against it, and every number joined from the record at build time. Side-by-side numbers live in the field table.

Ornith-1.0-35B (ornith-ai / deepreinforce-ai)

35B / A3B-class MoE · Q8_0 · designated primary/full-bench quant (PROTOCOL_OVERRIDE, not in protocol.json quants.expected); UD-Q4_K_XL now has its own measured full-26 capability record at 22/26, mean 0.8461538461538461 (run-0499), not Q8 inheritance or Q4/Q8 equivalence

benched guard 4/4

decode @32k 46.22 t/s · latest run 2026-08-17

Qwen3.6-35B-A3B

35B / 3B active · UD-Q4_K_XL

benched guard 4/4

decode @32k 42.45 t/s · latest run 2026-08-19

gpt-oss-120b

117B / ~5B active · UD-Q4_K_XL

benched guard failed

decode @32k 41.70 t/s · latest run 2026-08-13

excluded — returns stale answers from previous requests

Qwen3-Coder-Next (Qwen)

80B total / ~3B active (A3B-class MoE, 512 experts / 10 active + 1 shared) · Q8_0 · designated primary quant, fit-projected and confirmed comfortable against the 120 GiB GTT window

benched guard 4/4

decode @32k 36.97 t/s · latest run 2026-08-17

Ling-3.0-flash

124B / 5.1B active · Q4_K_M (corrected post-PR-26608 reference conversion)

benched guard 4/4

decode @32k 32.89 t/s · latest run 2026-08-19

Qwen3.8-27B

27B dense · Q8_0 (unsloth) · UD-Q4_K_XL comparability arm

benched no guard

decode @32k 19.23 t/s · latest run 2026-09-06

Qwen3.5-122B-A10B (MTP)

122B / 10B active · UD-Q4_K_M (unsloth, rev a91c2f7)

benched guard 4/4

decode @32k 18.17 t/s · latest run 2026-08-23

Nemotron-3-Super-120B-A12B

120B / ~12B active · UD-Q4_K_M (fair-quant, this bench) · UD-IQ4_XS (superseded acquisition quant, kept separate)

benched guard 4/4

decode @32k 17.74 t/s · latest run 2026-08-19

Laguna S 2.1 (poolside)

118B / ~8B active · Q4_K_M · sole staged quant

benched guard 4/4

decode @32k 15.53 t/s · latest run 2026-08-17

DeepSeek-V4-Flash-0731

256 experts / 6 active + 1 shared (deepseek4 arch) · UD-IQ3_XXS (unsloth) · sole staged quant, forced by the ~104 GB weight footprint

benched guard 4/4

decode @32k 12.35 t/s · latest run 2026-08-18

Qwen3.6-27B (staged as "qwen36-27b-mtp" — the artifact has no MTP path)

27B dense · Q4_K_M · house-standard quant (protocol.json quants.expected)

benched guard 4/4

decode @32k 10.86 t/s · latest run 2026-08-17

Deep-Thought-Posttrain (tsfrm)

361.8M dense · F16 · the only GGUF variant this candidate ships — no quantized version exists upstream (PROTOCOL_OVERRIDE against protocol.json quants.expected)

benched guard 2/4

latest run 2026-08-17

LLaDA2.2-flash

100B / 13B active · Q4_K_S (Akicou diffuse.* GGUF; loads only under diffuse-cpp)

screened no guard

latest run 2026-08-22

Qwen3.8-Flash-Next (Unsloth first look + KingJones full-STRIX)

~180B stored / 125B MoE (512 experts, ~6B active) + 51B N-gram/PLE tables + 4B MTP head + vision · Two non-transferable series · Unsloth UD-Q4_K_XL/Vulkan first look · KingJones Q4_0_ROCmFP4_STRIX/HIP bounded campaign

benched no guard

latest run 2026-08-30