Model candidate field — the whole list, and nothing pruned by assumption
Reconstructed 2026-08-10 after I answered “what other models should be added?” from memory, produced five, and pruned two of those on assumptions. A candidate field already existed in
docs/14-model-backend-benchmark.md(June 2026) and I had lost sight of it. This page exists so the answer stops depending on recall.
The standing rule
Nothing leaves this list without a measurement. Not on a model card’s own admission,
not on a vendor benchmark, and not on my framing of what role a model is “for”. Those are
claims, and the entire point of the harness is that claims get tested. Ruling a model out
before benchmarking reproduces the exact failure the retraction in clm-0036 was about:
a conclusion resting on something other than evidence.
Two specific exclusions I made and am withdrawing:
- Maple-Preview — excluded because its own model card concedes weak agentic performance. That is the vendor’s claim about their own model, which is the kind of thing we measure rather than accept.
- Gemma-4-12B — excluded because I had labelled it “an orchestrator” and orchestrators should not be scored on τ². The label was mine. The correct response to “τ² won’t characterise routing” is to build a routing benchmark, not to skip the model.
Status
In the aihydra capability matrix (4)
| model | note |
|---|---|
| Qwen3.5-122B-A10B (MTP) | production incumbent; deadlocks with thinking on (clm-0033) |
| gpt-oss-120b | stale-answer bug reproduced 6/6 (clm-0025); scored 0.00 thinking-off |
| Qwen3.6-35B-A3B | reflex-tier candidate |
| Nemotron-3-Super-120B-A12B | ⚠ ran at IQ4_XS while peers ran Q4_K_M — never quant-equivalent |
⚑ The Nemotron quant confound is live. It was forced to IQ4_XS by a 77 GiB-at-4K OOM on
the 96 GB box. aihydra’s 104 GiB GTT removes that constraint, so Q4_K_M is now testable —
and until it is, any Nemotron-vs-peer comparison is confounded by quantisation. This sits
directly underneath the thinking result clm-0035 called the standout.
On disk, benchmarked once, then dropped out of view (1)
Qwen3.6-27B-MTP — 16 GB, on aihydra now. 27B dense, 64 blocks, 262 144 context, GQA 24/4. Vendor claim 77.2 SWE-bench (≈Sonnet 4.5 coding) at 27B; MTP accelerates dense models 1.4–2× against 1.15–1.25× for MoE; Gated DeltaNet linear attention for cheap long context. Passed the Phase A tool-call gate 100% at 16–22 s time-to-correct.
Two reasons it matters more than its size suggests: it is a dense architecture, and dense is where MTP’s larger speed-up (1.4–2× against 1.15–1.25× for MoE) applies. It has zero rows in the aihydra matrix.
Never acquired (11)
| candidate | shape | the claim to test |
|---|---|---|
| Laguna S 2.1 (poolside) | 118B / ~8B hybrid global+SWA MoE | Terminal-Bench 70.2% vs Qwen’s 41.6%. Q4_K_M 75.2 GB, 36/48 SWA layers = tiny KV. Blocked on llama.cpp #25165 (fork-only until merged); SWA shares our #25913 checkpoint problem, so the sidecar experience applies. Available on OpenRouter at $0.10/$0.20 per M for feel-testing now |
| Ling-3.0-flash (Ant) | 124B / 5.1B | Near-ideal controlled peer to the 122B — same footprint, half the active params. Needs non-stock llama.cpp; MTP head ships INACTIVE, which would rig any decode comparison toward Qwen if not handled |
| GLM-4.7 Flash/Air | fits-tier | Strong agentic/coding lineage; ~40 t/s ROCm reported. Pick the variant that fits |
| GLM-4.5-Air | — | Appeared in an earlier chat-tier group config |
| Gemma 4 26B-A4B | 26B / 3.8B | ~85 t/s at near-31B quality — the “efficient fast daily”. Check tool-call format |
| Gemma-4-12B | 12B | Runs as a working orchestrator in domdoss/Warden (clm-0034). Needs the routing benchmark that does not exist yet — and should still be screened like everything else |
| Nemotron-Cascade 2 | 30B / 3B | Gold-medal competitive-programming reasoning at tiny active count |
| LFM2-24B (Liquid) | 24B | Speed outlier with huge prefill — the “is fast enough also good enough?” probe |
| Llama 4 Scout | 109B / 17B | The in-range Llama 4, ~61 GB Q4, large context |
| Maple-Preview (DeepGrove) | 20.2B-A1.49B ternary, 5.31 GB | Mac-mini second lane. Its “on-device weight adaptation (dreaming)” is undocumented at source |
| DeepSeek-V4-Flash | — | Community report claims Vulkan beating ROCm, 55% faster decode at depth. Fit against 104 GiB GTT unverified |
⚑ DeepSeek-V4-Flash is partly mis-framed as a model. Its most valuable claim is about the backend, and that is testable on models we already hold — cheaper and more decisive than acquiring it. Acquire it as a model too; just do not let that block the Vulkan test.
How 16 candidates get measured without a month of nights
The problem with the field is cost, not merit: a powered τ² arm is ~2–3 h, so 16 models × 2 thinking arms is not a schedule. The answer is a cheap screen everything passes through, with promotion on evidence — not exclusion on assumption. Doc 14 already had the instinct, gating on tool-call success.
- Fit probe — does it load at target context, and what is the real GTT high-water mark. Minutes. Catches the Nemotron-class OOM before it forces an unequal quant.
- Guard, 4 checks — coherence, tool-call, needle, isolation. Catches gpt-oss-class defects. Minutes.
- Short agentic screen — τ² small-n as a smoke test, which is what it is actually good for, plus tool-call gate. ~30 min.
- Powered τ² at n=50 — only for survivors, and only where a decision depends on it.
Steps 1–3 are ~30 min/model, so the entire field screens in roughly one night. Screening is not judging: a model that fails step 2 has measured evidence against it, which is the whole difference from what I did above.
Quant equivalence is a precondition, not a detail
Comparisons are only meaningful at matched quantisation. Nemotron at IQ4_XS against peers at
Q4_K_M was never a fair test, and the field below carries the same hazard wherever a model
only fits at a lower quant. Record the quant in every row, and where a model is forced down,
say so in the result rather than in a footnote. Open quant questions already noted in
runbooks/second-halo-integration-plan.md: Nemotron at Q4_K, the 122B at Q5_K_M/Q6, gpt-oss
at higher quant, and 35B Q8 vs Q4.