Docs › candidates

Model candidate field — the whole list, and nothing pruned by assumption

Reconstructed 2026-08-10 after I answered “what other models should be added?” from memory, produced five, and pruned two of those on assumptions. A candidate field already existed in docs/14-model-backend-benchmark.md (June 2026) and I had lost sight of it. This page exists so the answer stops depending on recall.

The standing rule

Nothing leaves this list without a measurement. Not on a model card’s own admission, not on a vendor benchmark, and not on my framing of what role a model is “for”. Those are claims, and the entire point of the harness is that claims get tested. Ruling a model out before benchmarking reproduces the exact failure the retraction in clm-0036 was about: a conclusion resting on something other than evidence.

Two specific exclusions I made and am withdrawing:

  • Maple-Preview — excluded because its own model card concedes weak agentic performance. That is the vendor’s claim about their own model, which is the kind of thing we measure rather than accept.
  • Gemma-4-12B — excluded because I had labelled it “an orchestrator” and orchestrators should not be scored on τ². The label was mine. The correct response to “τ² won’t characterise routing” is to build a routing benchmark, not to skip the model.

Status

In the aihydra capability matrix (4)

model note
Qwen3.5-122B-A10B (MTP) production incumbent; deadlocks with thinking on (clm-0033)
gpt-oss-120b stale-answer bug reproduced 6/6 (clm-0025); scored 0.00 thinking-off
Qwen3.6-35B-A3B reflex-tier candidate
Nemotron-3-Super-120B-A12B ran at IQ4_XS while peers ran Q4_K_M — never quant-equivalent

The Nemotron quant confound is live. It was forced to IQ4_XS by a 77 GiB-at-4K OOM on the 96 GB box. aihydra’s 104 GiB GTT removes that constraint, so Q4_K_M is now testable — and until it is, any Nemotron-vs-peer comparison is confounded by quantisation. This sits directly underneath the thinking result clm-0035 called the standout.

On disk, benchmarked once, then dropped out of view (1)

Qwen3.6-27B-MTP — 16 GB, on aihydra now. 27B dense, 64 blocks, 262 144 context, GQA 24/4. Vendor claim 77.2 SWE-bench (≈Sonnet 4.5 coding) at 27B; MTP accelerates dense models 1.4–2× against 1.15–1.25× for MoE; Gated DeltaNet linear attention for cheap long context. Passed the Phase A tool-call gate 100% at 16–22 s time-to-correct.

Two reasons it matters more than its size suggests: it is a dense architecture, and dense is where MTP’s larger speed-up (1.4–2× against 1.15–1.25× for MoE) applies. It has zero rows in the aihydra matrix.

Never acquired (11)

candidate shape the claim to test
Laguna S 2.1 (poolside) 118B / ~8B hybrid global+SWA MoE Terminal-Bench 70.2% vs Qwen’s 41.6%. Q4_K_M 75.2 GB, 36/48 SWA layers = tiny KV. Blocked on llama.cpp #25165 (fork-only until merged); SWA shares our #25913 checkpoint problem, so the sidecar experience applies. Available on OpenRouter at $0.10/$0.20 per M for feel-testing now
Ling-3.0-flash (Ant) 124B / 5.1B Near-ideal controlled peer to the 122B — same footprint, half the active params. Needs non-stock llama.cpp; MTP head ships INACTIVE, which would rig any decode comparison toward Qwen if not handled
GLM-4.7 Flash/Air fits-tier Strong agentic/coding lineage; ~40 t/s ROCm reported. Pick the variant that fits
GLM-4.5-Air Appeared in an earlier chat-tier group config
Gemma 4 26B-A4B 26B / 3.8B ~85 t/s at near-31B quality — the “efficient fast daily”. Check tool-call format
Gemma-4-12B 12B Runs as a working orchestrator in domdoss/Warden (clm-0034). Needs the routing benchmark that does not exist yet — and should still be screened like everything else
Nemotron-Cascade 2 30B / 3B Gold-medal competitive-programming reasoning at tiny active count
LFM2-24B (Liquid) 24B Speed outlier with huge prefill — the “is fast enough also good enough?” probe
Llama 4 Scout 109B / 17B The in-range Llama 4, ~61 GB Q4, large context
Maple-Preview (DeepGrove) 20.2B-A1.49B ternary, 5.31 GB Mac-mini second lane. Its “on-device weight adaptation (dreaming)” is undocumented at source
DeepSeek-V4-Flash Community report claims Vulkan beating ROCm, 55% faster decode at depth. Fit against 104 GiB GTT unverified

DeepSeek-V4-Flash is partly mis-framed as a model. Its most valuable claim is about the backend, and that is testable on models we already hold — cheaper and more decisive than acquiring it. Acquire it as a model too; just do not let that block the Vulkan test.

How 16 candidates get measured without a month of nights

The problem with the field is cost, not merit: a powered τ² arm is ~2–3 h, so 16 models × 2 thinking arms is not a schedule. The answer is a cheap screen everything passes through, with promotion on evidence — not exclusion on assumption. Doc 14 already had the instinct, gating on tool-call success.

  1. Fit probe — does it load at target context, and what is the real GTT high-water mark. Minutes. Catches the Nemotron-class OOM before it forces an unequal quant.
  2. Guard, 4 checks — coherence, tool-call, needle, isolation. Catches gpt-oss-class defects. Minutes.
  3. Short agentic screen — τ² small-n as a smoke test, which is what it is actually good for, plus tool-call gate. ~30 min.
  4. Powered τ² at n=50 — only for survivors, and only where a decision depends on it.

Steps 1–3 are ~30 min/model, so the entire field screens in roughly one night. Screening is not judging: a model that fails step 2 has measured evidence against it, which is the whole difference from what I did above.

Quant equivalence is a precondition, not a detail

Comparisons are only meaningful at matched quantisation. Nemotron at IQ4_XS against peers at Q4_K_M was never a fair test, and the field below carries the same hazard wherever a model only fits at a lower quant. Record the quant in every row, and where a model is forced down, say so in the result rather than in a footnote. Open quant questions already noted in runbooks/second-halo-integration-plan.md: Nemotron at Q4_K, the 122B at Q5_K_M/Q6, gpt-oss at higher quant, and 35B Q8 vs Q4.