Home › Evidence › Candidates

Candidates

Every model under consideration, and the recorded reason for every gate transition. The standing rule this board enforces: nothing leaves the pipeline without a measurement — vendor claims and role assumptions are admissible reasons to list a candidate and never grounds to drop one, so rejected cards show their evidence. Expand a card for its full gate history. NPU-lane cards run a different artifact at a different quantisation on a different runtime — their numbers never transfer to or from the GPU lane (4 on the board).

listed (6)

DeepSeek-V4-Flash-0731

deepseek-v4-flash · 256 experts / 6 active + 1 shared (deepseek4 arch) · moe · UD-IQ3_XXS

A community report on this model independently confirmed our mmap/GTT double-residency finding and demonstrated a 120 GiB GTT ceiling in production - so the report is credible on details we can check. A second, richer community measurement (2026-08-12) puts it at 23-24 t/s steady-state decode on our exact silicon with a speculative drafter, and the community is openly asking a quality question our rig can answer with data.

in gate since 2026-08-12 (1d)
gate history (3)
  • 2026-08-09 · listed

    Community report worth testing directly.

    evidence: clm-0031

  • 2026-08-10 · listed → listed

    Partly MIS-FRAMED as a model addition. Its most consequential claim - Vulkan beating ROCm on 3 of 4 cells, including 55% faster decode at depth - is about the BACKEND, and is testable on models we already hold, more cheaply and more decisively. Acquire it as a model too, but do not let acquisition block the Vulkan test. Fit against 104 GiB GTT is unverified.

    evidence: clm-0031

  • 2026-08-12 · listed → listed

    Identity now concrete via the r/LocalLLaMA Strix Halo guide thread: DeepSeek V4 Flash 0731, deepseek4 arch, 256 experts/6 active + 1 shared, MIT licence, 1M native context. Unsloth UD-IQ3_XXS is ~98 GB (a commenter says 104 GB, unanswered) - tight but plausible against the 104 GiB GTT budget; the thread's own numbers were produced on 126976 MiB GTT with a 4 GB BIOS carve-out, so fit at OUR boot params still needs a probe before any download decision. Author's steady-state on this exact silicon: 23-24 t/s decode with an 11 GB bf16 DSpark drafter (1.46x over plain decode), ~236 t/s prefill, on Nathan's Vulkan fork - headline 27 t/s is a peak-window cherry by his own caveat. KV note for our protocol: his KL-divergence data on q8_0 KV shows a fat tail (99.9%-ile KLD 64x the mean) and he recommends f16 KV for agentic work - consistent with clm-0035's direction. Open community question our tau2 rig could settle: antirez Q2-Q4 quants vs Unsloth UD-IQ3_XXS quality (currently argued on vibes in both directions). Backend decider spun out as coverage.md item 0; this candidate stays acquisition-gated on the fit probe.

    evidence: clm-0031

  • note

    The thread's quantised-drafter crash (invalid token = -1 on a Q2K draft) is a llama.cpp bug with a claimed GGUF-header fix ("Confirmed works", no patch link) - if real it cuts the draft footprint 11 GB -> ~6.5 GB, which matters at this model's size. Drafter must otherwise be bf16. COMMUNITY HAZARD (thefrontierlab.ai field report, 2026-08-13, r/StrixHalo 1vn9kn4): run f16 KV ALWAYS on this model. Quantising the K cache trips an open upstream llama.cpp issue - an incoherence rotation that disables the sparse-attention dispatch (dense fallback) and corrupts output on CPU/CUDA. The reporter could not reproduce the corruption on Vulkan (sparse buffers still allocate) but still measured ~-13 to -14% on BOTH prefill and decode from KV quant - opposite sign to every GQA model in his nine-model sweep. His build refused mixed cache types, so K-only isolation is untested. Unverified here; screen with f16 KV only.

EmbeddingGemma (NPU)

npu-embeddinggemma · ~300M · dense · FLM manifest (not GGUF) · FastFlowLM
npu lane

Semantic triage for the continuous escalation loop. Embeddings let the loop compare an incoming item against memory before deciding whether to escalate, rather than judging on surface features alone - the difference between "this mentions a deadline" and "this contradicts something we already recorded".

in gate since 2026-08-10 (3d)
gate history (2)
  • 2026-08-10 · listed

    Listed as the highest-value NPU workload we have: a real recurring production job that does not need the big model and currently steals slot time from one that does.

  • 2026-08-10 · listed → listed

    RATIONALE CORRECTED - the reason above is false. qmd's re-embedding runs on the MAC MINI and never competed for the inference box's GPU slot at all; the operator caught it. I asserted a contention I had not checked, which is the same failure as dropping a candidate on an unverified claim, just aimed at keeping one instead of cutting one. The candidate stays listed on the corrected rationale in why_listed above.

  • note

    NPU runs are QUEUED BEHIND GPU work, never concurrent. The NPU shares the same LPDDR5X pool and memory bandwidth as the GPU on Strix Halo, so a concurrent NPU run would both contend for bandwidth and perturb any GPU number measured alongside it - which would quietly corrupt the GPU series rather than merely slowing it. ⚠ amd_iommu=off is worth +5-12% on the GPU and DISABLES the NPU entirely. Those are mutually exclusive strategies, so the second lane has a standing cost that has to be counted against whatever it buys.

LFM2 (Liquid, NPU)

npu-lfm2 · small variant, TBD · moe · FLM manifest (not GGUF) · FastFlowLM
npu lane

The only family in FLM's catalog that also sits in our GPU field, which makes it the one candidate that could answer whether NPU-versus-GPU is a fair trade at all - same family, both lanes, measured separately.

in gate since 2026-08-10 (3d)
gate history (1)
  • 2026-08-10 · listed

    Listed for cross-lane comparability of the FAMILY, explicitly not for metric transfer. The GPU entry is LFM2-24B-A2B at Q4_K_M; the NPU artifact will be a different size at FLM's own quantisation, so the two produce independent numbers that may be compared but never merged.

  • note

    NPU runs are QUEUED BEHIND GPU work, never concurrent. The NPU shares the same LPDDR5X pool and memory bandwidth as the GPU on Strix Halo, so a concurrent NPU run would both contend for bandwidth and perturb any GPU number measured alongside it - which would quietly corrupt the GPU series rather than merely slowing it. ⚠ amd_iommu=off is worth +5-12% on the GPU and DISABLES the NPU entirely. Those are mutually exclusive strategies, so the second lane has a standing cost that has to be counted against whatever it buys.

Qwen3-4B-Thinking-2507 (NPU)

npu-qwen3-4b-thinking · 4B · dense · FLM manifest (not GGUF) · FastFlowLM
npu lane

The triage brain for the continuous escalation loop. FLM's validated headline model, and thinking-capable at 4B, which makes it the actual escalate-or-hold decision maker rather than merely a proof the lane runs. Bring-up and first real candidate are the same model here, which is convenient but incidental.

in gate since 2026-08-10 (3d)
gate history (2)
  • 2026-08-10 · listed

    Listed as the NPU lane's bring-up model: if this does not run, nothing else on the lane matters. Deliberately chosen for low risk rather than interest.

  • 2026-08-10 · listed → listed

    REFRAMED against the actual design intent. The original list was bring-up shaped - does the lane work - with use-case reasons retrofitted. The standing intent is a CONTINUOUS TRIAGE LOOP watching inputs and escalating to the big model, and this is the one candidate that is genuinely the loop's decision maker rather than a component near it. Its evaluation is therefore escalation precision/recall over sustained operation, NOT tau2 task success.

    evidence: clm-0034

  • note

    NPU runs are QUEUED BEHIND GPU work, never concurrent. The NPU shares the same LPDDR5X pool and memory bandwidth as the GPU on Strix Halo, so a concurrent NPU run would both contend for bandwidth and perturb any GPU number measured alongside it - which would quietly corrupt the GPU series rather than merely slowing it. ⚠ amd_iommu=off is worth +5-12% on the GPU and DISABLES the NPU entirely. Those are mutually exclusive strategies, so the second lane has a standing cost that has to be counted against whatever it buys.

Whisper (NPU)

npu-whisper · TBD · dense · FLM manifest (not GGUF) · FastFlowLM
npu lane

STT already runs on wardenmac via whisper.cpp and is confirmed working end to end. Moving it to the NPU is the cheapest possible test of whether the lane can take a real production workload off an already-loaded machine.

in gate since 2026-08-10 (3d)
gate history (2)
  • 2026-08-10 · listed

    Listed because there is an existing working baseline to compare against, which is rare - most NPU numbers would have nothing to be judged relative to.

  • 2026-08-10 · listed → listed

    Role clarified for the escalation loop: Whisper is a genuine CONTINUOUS INPUT SOURCE, which is what the loop consumes, not a model the loop chooses between. It tests sustained residency and non-interference - can something stay resident on the NPU for hours without perturbing concurrent GPU work - which is a loop requirement no one-shot benchmark covers.

  • note

    NPU runs are QUEUED BEHIND GPU work, never concurrent. The NPU shares the same LPDDR5X pool and memory bandwidth as the GPU on Strix Halo, so a concurrent NPU run would both contend for bandwidth and perturb any GPU number measured alongside it - which would quietly corrupt the GPU series rather than merely slowing it. ⚠ amd_iommu=off is worth +5-12% on the GPU and DISABLES the NPU entirely. Those are mutually exclusive strategies, so the second lane has a standing cost that has to be counted against whatever it buys.

Qwen3.8-27B

qwen38-27b · 27B (dense expected) · unknown · Q4_K_M

Successor to the Qwen3.6-27B-MTP line, a dense architecture — the case where MTP's larger dense speed-up (1.4-2x, against 1.15-1.25x for MoE) applies. If the 27B line holds its claimed coding strength - the 3.6 generation claimed 77.2 SWE-bench at 27B - then a same-size successor is the highest-leverage candidate we have: reflex-tier footprint with executor-tier capability.

in gate since 2026-08-10 (3d)
gate history (2)
  • 2026-08-10 · listed

    Listed on the operator's expectation of an imminent release. Provenance is honest about itself: this is an expectation, not a vendor benchmark and not a measurement, and it is a fine reason to LIST - listing costs nothing and commits us to nothing.

  • 2026-08-10 · listed → listed

    Confirmed the line EXISTS but is not yet runnable for us. HF carries huginnfork/Qwen3.8-27B-FP8, huginnfork/Qwen3.8-27B-NVFP4A16 and neroued/Qwen3.8-27B-nvfp4-NInfer, so the 27B weights are out - but the only Qwen3.8 GGUFs are 4B and 1.2B distills. No 27B GGUF means nothing we can run on llama.cpp yet. A daily two-tier watch is now live on aihydra (ANY = line exists, early warning; GGUF = actionable) writing to ~/.model-watch.NEW, so the first 27B GGUF surfaces without anyone having to remember to check.

  • note

    Watch it against the incumbent it would replace rather than in isolation: Qwen3.6-27B-MTP is already on disk with zero rows, so screening the two together costs barely more than screening either alone and answers the generation question directly. ⚠ Do not inherit the 3.6 generation's MTP assumption. MTP support is per-model and has to be confirmed in the GGUF rather than assumed from the line - Ling-3.0-flash ships its MTP head INACTIVE (clm-0015), which would silently distort any decode comparison, and Nemotron's MTP turned out incompatible with draft-mtp entirely because of Mamba. Two precedents is enough to make this a check, not a footnote. Vision audit (2026-08-10): Qwen positions the whole 3.8 generation, including this 27B, as multimodal-capable per its own announcement materials - consistent with the 3.5/3.6 pattern already confirmed vision-capable on this box (see the qwen35-122b, qwen36-27b-mtp and qwen36-35b notes). But that is a parent-line claim only: as recorded above, no 27B GGUF conversion exists anywhere yet, vision or otherwise, so there is no artifact to check an mmproj against and nothing for llama.cpp to load. When a 27B conversion surfaces on the existing watch, check for a companion mmproj in the same repo before assuming one exists.

acquired (0)

none at this gate

screened (13)

Gemma-4-12B

gemma4-12b · 12B · dense · Q4_K_M

Runs as a working ORCHESTRATOR in domdoss/Warden, which demonstrates that routing and brief-writing have a far lower capability bar than task execution. If a 12B can route, the reflex tier is much cheaper than we assumed.

in gate since 2026-08-13 (0d)
gate history (3)
  • 2026-08-09 · listed

    Evidence from a comparable project that a 12B suffices for the orchestrator role.

    evidence: clm-0034

  • 2026-08-10 · listed → listed

    I had excluded this from the field on the grounds that orchestrators should not be scored on tau2. That was my own role label, not a measurement, and the operator correctly rejected it. Re-listed. The right response to 'tau2 will not characterise routing' is to BUILD a routing benchmark, not to skip the model.

    evidence: clm-0034

  • 2026-08-13 · listed → screened

    Downloaded overnight and sha256-verified (companion mmproj-F16.gguf fetched alongside; this screen is TEXT-ONLY per the house tier, projector deliberately not loaded). Screened on aihydra, build min-62bf73d (62bf73d25) — the newest protocol-registered build — under the Aug-12 protocol (pinned haiku-4.5 simulator). FIT ok, loaded in 8s at 32K. GUARD 3/4, NOT clean: coherence PASS, toolcall PASS, isolation PASS (skipped under --parallel 1), needle FAIL — NEEDLE LOST at depth 8000 (tok=5364): in place of the needle the model returned a run of '<unused49>' control-token spam. SMOKE 5-task tau2 airline: rc=124 — the 3600s wall cap expired with 0/5 tasks completed, task 1 still running at 3,600s; no results file, so mean=n/a, cut count n/a. The server log shows a single runaway generation at n_decoded=13,525 and climbing when killed — a degenerate-output profile consistent with the needle spam, exactly the gpt-oss-class defect GUARD exists to catch. Caveat a screen cannot resolve: whether this is the checkpoint, the Q4_K_M quant, or a chat-template/tokenizer gap for the 12B's "unified encoder-free" variant in current llama.cpp. The orchestrator-role question from listing is untouched either way — a model that spams control tokens at depth cannot route. Smoke is a cliff-detector, not a ranking (noise ~0.40, clm-0036). Run-meta: screening/gemma4-12b 2026-08-13T06:52:41Z->07:53:06Z (3625s, rc=124).

  • note

    Vision audit (2026-08-10): parent IS vision-capable - Gemma 4 ships multimodal across all five sizes, and the 12B specifically uses a distinct "unified encoder-free" architecture (raw image/audio patches projected directly, no separate CLIP tower) rather than the standard vision path used by the 26B-A4B/31B sizes. Our llama.cpp build (commit 3653e6d) already registers both paths - clip.cpp defines PROJECTOR_TYPE_GEMMA4UV/GEMMA4UA for this unified 12B variant alongside GEMMA4V/GEMMA4A for the rest of the line. Not staged, so no mmproj question applies yet; if acquired, confirm the specific HF repo publishes a companion mmproj (the 26B-A4B unsloth repo does - see gemma4-26b) before assuming vision comes along for free.

GLM-4.5-Air

glm-45-air · fits-tier · moe · Q4_K_M

Appeared in an earlier chat-tier group config alongside gpt-oss-120b and qwen-fast, so it was once considered deployable - but it was never benchmarked and never explained.

in gate since 2026-08-13 (0d)
gate history (2)
  • 2026-08-10 · listed

    Surfaced by search rather than recall. Listed to be resolved: either it supersedes GLM-4.7-Flash for our purposes or it does not, and right now nothing on record says which.

  • 2026-08-13 · listed → screened

    Downloaded overnight and sha256-verified (2-part Q4_K_M GGUF, 68 GiB; server pointed at part 00001). Screened on aihydra, build min-62bf73d (62bf73d25), Aug-12 protocol (pinned haiku-4.5 simulator). FIT ok, loaded in 45s at 32K — the fit question is settled, it runs with headroom. GUARD 3/4, NOT clean: coherence PASS, toolcall PASS, isolation PASS (skipped under --parallel 1), needle FAIL — NEEDLE LOST at depth 8000 (tok=5363): the model returned a run of '?' characters in place of the needle. SMOKE 5-task tau2 airline: rc=124 — the 3600s wall cap expired with 0/5 completed, task 1 still running; no results file, mean=n/a, cut count n/a. Server log shows a runaway generation at n_decoded=9,500 and climbing when killed. IMPORTANT shared-profile caveat: this is the SAME failure signature gemma4-12b produced an hour earlier on the same build (needle lost to junk-token spam + smoke runaway) while qwen36-27b-mtp and nemotron35-lightning-30b screened CLEAN on this identical build/harness last night — so the harness demonstrably can pass, but two unrelated archs failing identically in one session leaves model-defect-vs-build-interaction (FA/rope/template on gemma4+glm4moe paths) genuinely undecided. Do not treat this screen as a capability verdict on the checkpoint without a cross-build repro; do not treat it as harness noise either — the runaway is real and measured. The supersedes-GLM-4.7-Flash question from listing stays open. Smoke is a cliff-detector, not a ranking (noise ~0.40, clm-0036). Run-meta: screening/glm-45-air 2026-08-13T07:53:25Z->08:54:32Z (3667s, rc=124).

  • note

    Vision audit (2026-08-10): parent is text-only. Zhipu ships vision separately as GLM-4.5V; GLM-4.5-Air itself is the agentic/tool-use text line. Not staged, so no artifact to check, but the parent-level answer alone settles it: text-only, not a vision-lane candidate.

Ling-3.0-flash (Ant Group)

ling-30-flash · 124B / 5.1B active · moe · Q4_K_M

A near-ideal CONTROLLED comparison for the 122B: same footprint, half the active parameters. That isolates active-parameter count as a variable, which nothing else in the field does.

in gate since 2026-08-13 (0d)
gate history (3)
  • 2026-07-26 · listed

    The only candidate that functions as a controlled experiment rather than another data point.

  • 2026-08-10 · listed → listed

    Still listed, not acquired. Two hazards to handle first: it needs a non-stock llama.cpp, and its MTP head ships INACTIVE - benchmarking it as shipped would silently rig any decode comparison in Qwen's favour.

    evidence: clm-0015

  • 2026-08-13 · listed → screened

    Downloaded overnight and sha256-verified (2-part Q4_K_M_STOCK GGUF, 70 GiB; server pointed at part 00001). Screening attempted on aihydra, build min-62bf73d (62bf73d25), Aug-12 protocol. FIT FAILED after 4s and the verbatim error IS the screening result (laguna precedent): "llama_model_load: error loading model: unknown model architecture: 'bailingmoe3'". Secondary probe on 77a9a66eb, the only by-date-newer registered build, fails with the identical line — unsupported on every registered general-purpose binary. No GUARD, no SMOKE. This measures what the 2026-08-10 entry predicted ("needs a non-stock llama.cpp"): the controlled active-parameter comparison against the 122B stays the most scientifically interesting question in the field, and it is now concretely blocked on upstream bailingmoe3 support landing in a protocol-registered build. The shipped-INACTIVE MTP head hazard from that entry still stands unmeasured and must be handled before any future decode comparison. Run-meta: screening/ling-30-flash 2026-08-13T07:53:17Z->07:53:25Z (8s, rc=1).

  • note

    Vision audit (2026-08-10): parent is text-only. Ant Group splits its model families by modality - Ling (text MoE, this candidate), Ring (reasoning), and Ming (multimodal: vision/speech/audio) are three separate lines, not variants of one checkpoint. Ling-3.0-flash carries no vision capability to inherit. Verdict: text-only, not a vision-lane candidate.

Llama 4 Scout

llama4-scout · 109B / 17B active · moe · Q4_K_M

The in-range Llama 4 at ~61 GB Q4 with a large context window - a fits-tier representative of the Llama lineage.

in gate since 2026-08-13 (0d)
gate history (3)
  • 2026-06-20 · listed

    Marked optional at listing: ~half gpt-oss speed at 17B active, and not clearly better than the gpt-oss control in its size class.

  • 2026-08-10 · listed → listed

    Still listed. 'Optional' was a priority call, not a rejection, and it stays in the field until something measured removes it.

  • 2026-08-13 · listed → screened

    Downloaded overnight and sha256-verified (2-part Q4_K_M GGUF, ~61 GiB; server pointed at part 00001; the vision-note mmproj question remains untouched — text-only screen). Screened on aihydra, build min-62bf73d (62bf73d25), Aug-12 protocol (pinned haiku-4.5 simulator). FIT ok, loaded in 40s at 32K. GUARD 3/4: coherence PASS, needle PASS at depth 8000 (tok=5361 'chartreuse-viper-88'), isolation PASS (skipped under --parallel 1), toolcall FAIL — probe raised HTTPError 500. SMOKE 5-task tau2 airline: rc=0 in 22 min but completed=0/5, mean=n/a, cut_at_max_steps=0 — every task died on the same server-side rejection. Verbatim error, repeated throughout the server log: {"error":{"code":500,"message":"The model produced output that does not match the expected peg-native format","type":"server_error"}} — and the common_chat_peg_parse warning shows WHAT was rejected: 'unparsed peg-native output: [sophia_silva_7557, get_user_details(user_id="sophia_silva_7557")]', i.e. well-formed Llama-4 pythonic tool-call syntax the build's peg-native grammar will not accept. Distinct profile from the gemma4-12b/glm-45-air runaway-generation failures the same night: here decode is sane and needle retrieval works — the failure is confined to tool-call format handling, which reads as a chat-template/parser dialect mismatch in the build (fixable harness-side) rather than model incapability. Do not score the model on this smoke; re-screen once a registered build parses Llama-4 pythonic calls. Smoke is a cliff-detector, not a ranking (noise ~0.40, clm-0036). Run-meta: screening/llama4-scout 2026-08-13T08:54:34Z->09:17:41Z (1387s, rc=0).

  • note

    Vision audit (2026-08-10): parent IS vision-capable, and natively so - Llama 4 uses early fusion (image patches tokenized alongside text from pretraining, no bolt-on encoder), and Scout specifically is documented for visual recognition, image reasoning, and captioning. Our llama.cpp build (commit 3653e6d) already carries real runtime support: LLM_ARCH_LLAMA4 is registered and clip.cpp implements PROJECTOR_TYPE_LLAMA4 end-to-end. Not staged yet (still the ~61GB Q4 download from listing), so the mmproj question doesn't apply until then - but unlike the Qwen/Gemma/Muse-Glimmer cases, this one is runtime-verified-ready pending acquisition rather than a hazard to route around. When it comes off the list: confirm whichever HF repo we pull from actually publishes a companion mmproj before assuming it ships bundled with the quant.

DeepGrove Maple-Preview

maple-preview · 20.2B / 1.49B active, ternary · moe · TQ2_0-head-Q4_K

A ternary 20.2B at just 5.31 GB, MIT-licensed. That footprint would make it viable as a Mac-mini second lane where nothing else in the field fits at all.

in gate since 2026-08-13 (0d)
gate history (3)
  • 2026-07-30 · listed

    Ternary quantisation at a footprint that opens a second lane on hardware we already own.

    evidence: clm-0018

  • 2026-08-10 · listed → listed

    I had excluded this because its own model card concedes underperformance on agentic benchmarks. That is the vendor's claim about their own model, which is exactly the kind of thing this harness exists to test rather than accept. Exclusion withdrawn. Its undocumented 'on-device weight adaptation (dreaming)' remains unverified at source and is a separate question from whether the model is any good.

    evidence: clm-0018

  • 2026-08-13 · listed → screened

    Downloaded overnight and sha256-verified (5.9 GB single artifact; quant field corrected to the shipped TQ2_0-head-Q4_K — DeepGrove publishes no Q4_K_M, the ternary singleton is the only artifact). Screening attempted on aihydra, build min-62bf73d (62bf73d25), Aug-12 protocol; the off-list quant required a recorded PROTOCOL_OVERRIDE of the quant gate (screening-only, entering no same-quant comparison — override text lands on the run-meta line). FIT FAILED after 4s and the verbatim error IS the screening result (laguna precedent): "llama_model_load: error loading model: unknown model architecture: 'maple'". Secondary probe on 77a9a66eb, the only by-date-newer registered build, fails with the identical line — the arch is unsupported on every registered general-purpose binary, not just the screening default. No GUARD, no SMOKE. The Mac-mini second-lane question from listing stays open, but it is now measurably blocked on upstream llama.cpp 'maple' arch support rather than on anything we control. Run-meta: screening/maple-preview 2026-08-13T07:53:08Z->07:53:16Z (8s, rc=1).

  • note

    Vision audit (2026-08-10): parent is text-only. DeepGrove's own materials describe Maple-Preview as a natively-trained ternary reasoning model with no vision component; the model card's acknowledged agentic weaknesses (already noted above) are about tool use, not modality. Verdict: text-only, not a vision-lane candidate.

NVIDIA Nemotron 3.5 Lightning 30B-A3B

nemotron35-lightning-30b · 30B / ~3B active · hybrid · Q4_K_M

NVIDIA's own model card confirms a hybrid Mamba-2 + Attention + MoE "execution layer" design (128 routed experts, 6 active + 1 shared, native MTP baked into the checkpoint, 1M token context) pitched at exactly Warden's workload: long-running agent loops doing tool calling, state management and result validation. It is the direct small-footprint sibling of the already-benched nemotron3-super, and llama.cpp gained Nemotron MTP support (#26725, merged 2026-08-10) after nemotron3-super's "MTP incompatible with Mamba" finding was recorded - that finding may now be stale for this whole family. An official Q4_K_M GGUF exists at 23.69 GiB, no fork required, and NVIDIA's own benchmark table gives a real (if middling) agentic-axis result: SWE-bench Verified 51.56, Terminal-Bench 2.1 24.58, tau3-bench-Banking 9.28 - all below our incumbent Qwen3.6-35B-A3B on NVIDIA's own comparison, but a genuine measurement all the same.

in gate since 2026-08-13 (0d)
gate history (3)
  • 2026-08-12 · listed

    Surfaced via an operator-shared r/AIDeveloperNews link ("NVIDIA has launched Nemotron 3.5 Lightning") plus a follow-up r/StrixHalo post on a community ROCmFP4 requant with hardware-matched Strix Halo numbers. Independently verified: a real NVIDIA release (nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-{BF16,NVFP4}, OpenMDW-1.1 licence, released 2026-08-11, commercial use permitted), with an official ggml-org GGUF repo publishing a stock Q4_K_M (25,430,738,944 bytes / 23.69 GiB) alongside BF16/NVFP4/Q8_0 - no fork needed for the standard quant. Architecture (nemotron_h_moe: hybrid Mamba-2 + Attention + MoE) has been in mainline llama.cpp since PR #20411 (2026-03-11), long before any of our deployed builds, and native MTP support for the Nemotron family landed via #26725 merged 2026-08-10T08:25Z - BEFORE our 62bf73d build (2026-08-10T11:07Z), so that build should load this model with MTP working. Our older 3653e6d build (2026-08-07) predates the MTP commit and would load the model (base arch is old enough) but without the MTP speedup. DFlash-specific draft-model support (#26905) merged 2026-08-11T13:16Z, after 62bf73d, so DFlash itself is NOT confirmed available on either resolvable deployed build - 77a9a66eb did not resolve against ggml-org/llama.cpp mainline history via the GitHub API and its date is unconfirmed; treat its DFlash/MTP support as unverified until the operator confirms its lineage. NVIDIA's own published benchmark table places this model BELOW our incumbent Qwen3.6-35B-A3B on nearly every agentic metric (SWE-bench Verified 51.56 vs 70.12, SWE-bench Multilingual 39.33 vs 63.40, Terminal-Bench 2.1 24.58 vs 44.38, tau3-bench-Banking 9.28 vs 10.52, BrowseComp 36.97 vs 48.74, PinchBench 85.37 vs 88.07) - listing on the strength of a real, vendor-published agentic measurement and clean upstream architecture support, not a capability-leadership claim.

  • 2026-08-13 · listed → acquired

    Official ggml-org Q4_K_M downloaded to aihydra (25,430,738,944 bytes, exactly the listed size) and sha256-verified against the HF LFS oid (6110e2e2e6cd324e6ee69ddced5a6b34fad6c94ca9827222a1e420fb92e3c90b). Mirrored to the NAS gguf-library.

  • 2026-08-13 · acquired → screened

    Screened same night on build min-62bf73d (62bf73d25) — the newest protocol-registered build that post-dates Nemotron MTP #26725 — under the house FIT/GUARD/SMOKE tier (Aug-12 protocol, pinned haiku-4.5 simulator). FIT ok, loaded in 8s at 32K. GUARD 4/4 PASS including needle at depth 8000 (tok=5370 'chartreuse-viper-88'). SMOKE 5-task tau2 airline: completed=5/5, mean=0.00, cut_at_max_steps=3, 7 min wall — the cascade2-30b profile (needle pass + agentic flail with non-termination), and directionally consistent with NVIDIA's own below-incumbent agentic table. MTP caveat: 62bf73d loads the model but the server logged the layer-52 MTP head as unused ("model has unused tensor blk.52.nextn.enorm.weight ... -- ignoring", likewise hnorm/eh_proj/ shared_head_norm), so this screen ran the base decode path with NO MTP engagement; the "does the small sibling's MTP work now" question from listing is still open and needs whatever draft/spec flag or newer build actually wires the nextn head. Smoke is a cliff-detector, not a ranking (noise ~0.40, clm-0036). Run-meta: screening/nemotron35-lightning-30b 2026-08-13T03:06:34Z->03:14:38Z (484s, rc=0).

  • note

    Marketing-claim caveat: the r/AIDeveloperNews post (linking a non-canonical promo site, aideveloper44.com, not nvidia.com) claims the model was "trained specifically for popular agent harnesses (like OpenClaw and Hermes Agent)". This is NOT corroborated anywhere in NVIDIA's own model card, which describes only generic "agentic, tool-use" training data (synthetic CLI/web/coding agent corpora from gpt-oss-120b, Qwen3-Coder-480B, GLM-4.7-Flash) with no named-harness claim. Treat the OpenClaw mention as unverified vendor-adjacent marketing copy, not a primary-source fact - flagging because the coincidence with our own deployed harness is exactly the kind of detail that should be checked rather than taken at face value. Community Strix Halo numbers (r/StrixHalo, MrWidmore888, hardware: Framework Ryzen AI Max 395+, gfx1151, 128 GB unified memory - same hardware class as aihydra): these are measured on julianmb's ROCmFP4 community requant, NOT the official Q4_K_M, and are NOT directly inheritable as a screening result for that reason (same pattern as muse-glimmer-30b). Vulkan0 backend, FlashAttention, q8_0 KV cache, full unified-memory offload: STRIX_LEAN preset (Q4_0_ROCMFP4_STRIX_LEAN, ~4.38 bpw, 15.73 GiB) - pp512 1,299.7 t/s, tg128 85.6 t/s, wikitext-2 perplexity 5.9936±0.0358; FAST preset (~4.25 bpw, 15.66 GiB) - pp512 1,310.5 t/s, tg128 86.0 t/s; COHERENT preset (agentic/coding tuned, ~4.70 bpw, 16.74 GiB) - pp512 1,290.4 t/s, tg128 81.6 t/s. Poster notes Vulkan beat ROCm by ~21% on prompt processing on this box. These ROCmFP4 tensor types are NOT stock llama.cpp - loading them needs charlie12345/ROCmFPX (independently confirmed real on GitHub, same fork already carried as a blocker for muse-glimmer-30b) via julianmb's shoutout in the post. The official stock Q4_K_M path does not need this fork at all, so the fork question only applies if chasing this specific community quant rather than NVIDIA's own artifact. Fit: official Q4_K_M is 23.69 GiB - comfortable in either the 84 or 104 GiB GTT budget with large headroom for KV/context even before touching the 1M-token window (compare nemotron3-super, which was forced to IQ4_XS and still OOMed by 16K on the smaller box). Same architecture family as nemotron3-super (nemotron_h_moe) but roughly 4x smaller total / 4x smaller active params - this is a genuine niche question, not a clean duplicate: nemotron3-super occupies the large-agentic-reasoning slot (120B/12B active, already benched), while this candidate would sit in the reflex/executor tier alongside qwen38-27b and qwen36-27b-mtp - except MoE rather than dense, and with MTP genuinely native to the checkpoint rather than bolted on. Worth screening against BOTH neighbours: nemotron3-super (does the small sibling's MTP actually work now that #26725 is upstream, re-testing the "Mamba rejects MTP" finding on a Q4_K_M-clean artifact) and Qwen3.6-35B-A3B (the true incumbent this would need to beat on tool-calling, per NVIDIA's own numbers it currently does not).

Qwen3.6-27B-MTP

qwen36-27b-mtp · 27B dense · dense · Q4_K_M

A dense architecture — the case where MTP's larger dense speed-up (1.4-2x, against 1.15-1.25x for MoE) applies. Vendor claims 77.2 SWE-bench at 27B, near Sonnet 4.5 coding. Gated DeltaNet linear attention for cheap long context.

in gate since 2026-08-13 (0d)
gate history (5)
  • 2026-06-20 · listed

    Dense architecture + dense-MTP multiplier + a strong vendor coding number.

  • 2026-06-27 · listed → acquired

    Q4_K_M, 16 GB, on aihydra now.

  • 2026-06-27 · acquired → screened

    Phase A: 16-22s with MTP active, 100% tool-call gate. MTP works but does not rescue latency - reasoning tokens dominate.

  • 2026-08-10 · screened → acquired

    REGRESSED to acquired, not dropped. It has zero rows in the aihydra matrix and I lost track of it entirely, describing it to the operator as an accidental find. Its Phase A screen was on the OLD box under a different harness, so it needs re-screening here rather than inheriting a stale pass.

  • 2026-08-13 · acquired → screened

    RE-SCREENED on aihydra under the house FIT/GUARD/SMOKE tier (build 3653e6d, Aug-12 protocol: pinned simulator openrouter/anthropic/claude-haiku-4.5). FIT ok, loaded in 16s at 32K. GUARD 4/4 PASS — coherence, toolcall (get_booking {"reference":"ABC123"}), needle at depth 8000 (tok=5368 'chartreuse-viper-88' retrieved), isolation (skipped by design at --parallel 1; #25992 only manifests above 1). SMOKE 5-task tau2 airline: completed=5/5, mean=0.80, cut_at_max_steps=0, 11 min wall. Equal to laguna-s21's 0.80 — the joint-strongest smoke on this box — but a smoke test is a cliff-detector, not a ranking (noise ~0.40, clm-0036). Run-meta: screening/qwen36-27b-mtp 2026-08-13T02:54:10Z->03:06:26Z (736s, rc=0).

  • note

    Vision audit (2026-08-10): parent IS vision-capable - Qwen's own materials describe Qwen3.6-27B as natively multimodal in a single unified checkpoint (same as Qwen3.6-35B-A3B), handling images and video alongside text. Confirmed on our staged artifact: Qwen3.6-27B-Q4_K_M.gguf reads general.architecture=qwen35 (llama.cpp reuses the 3.5-era arch code for 3.6), carries the "image-text-to-text" tag, and has qwen35.rope.dimension_sections=[11,11,10,0] - MRoPE sectioning implemented natively in src/models/qwen35.cpp on our build, not requiring a separate VL arch. unsloth/Qwen3.6-27B-GGUF (source repo for our download) publishes mmproj-F16.gguf (927,607,360 bytes), not staged. Runtime: LLM_ARCH_QWEN35 is registered in our build (commit 3653e6d), and clip.cpp implements PROJECTOR_TYPE_QWEN3VL generically on the LLM side. Verdict: mmproj-download-away, not text-only - and notably this is a DENSE vision-capable candidate, worth weighing against the MoE Qwen3.5/3.6 vision candidates on that basis alone.

Muse Glimmer 30B (vmlinux ROCmFPX quant)

muse-glimmer-30b · ~29.6B dense (+ ~1.8B vision encoder) · dense · ROCmFP4 (Q4_0_ROCMFP4_STRIX, fork-only - no Q4_K_M exists)

Surfaced via r/StrixHalo share link pointing at a third-party ROCmFPX quant of Meta's Muse-Glimmer-30B. The operator's framing was that the PARENT model scores well at agentic benchmarks - verified true on meta-models/Muse-Glimmer-30B's own card: SWE-Bench Verified 76.0, SWE-Bench Pro 51.2, MCP Atlas 75.5, DeepSearch QA 74.6, τ3-Banking 23.5 - a real, Meta-published result for the full-precision base model. Worth listing on that basis even though it is not this quant's own number (see note).

in gate since 2026-08-11 (2d)
gate history (5)
  • 2026-08-10 · listed

    Listed on the strength of the parent model's own published agentic benchmark table (SWE-Bench Verified 76.0, MCP Atlas 75.5, SWE-Bench Pro 51.2), confirmed on meta-models/Muse-Glimmer-30B's model card. The Reddit link itself was login-walled and added nothing beyond what the HF repo already documents.

  • 2026-08-10 · listed → acquired

    Downloaded Muse-Glimmer-30B-ROCmFP4.gguf (15,210,123,424 bytes) to aihydra for staging. Acquisition costs nothing while the GPU is busy, even though this specific artifact cannot run yet - see blocker.

  • 2026-08-10 · acquired → blocked

    The downloaded file uses ROCmFPX's experimental Q4_0_ROCMFP4_STRIX tensor type, which the repo's own README states "will not load in stock llama.cpp." It needs charlie12345/ROCmFPX (a fork) with ROCmFPX-Muse-Glimmer.patch applied to a pinned commit - same class of hazard as the ROCmFP4 precedent already carried for Laguna. The Muse-Glimmer architecture ITSELF is merged into stock llama.cpp (ggml-org/llama.cpp#26841), so ordinary K-quant GGUFs of this model (Meta's own official repo, or Unsloth's) would run today - but neither publishes a plain Q4_K_M, so a stock-compatible download would still not be quant-matched against our field. Staying blocked until either a literal Q4_K_M appears or the fork question is revisited.

  • 2026-08-10 · blocked → blocked

    REASSESSED on review: the block applies to the ROCmFPX ARTIFACT, not the model. The muse-glimmer architecture is merged in mainline llama.cpp (PR #26841), so the model IS screenable on our stock build via a standard quant - Unsloth publishes Muse-Glimmer-30B-UD-Q4_K_XL, the same quant family as our gpt-oss artifact. Downloading it as the screening artifact; the FPX file stays staged for a later performance-lever experiment (FPX-vs-standard on identical weights), pending the fork-build decision. Note the fork patch is MODEL-SPECIFIC per its own docs, so it does not amortise across other candidates - the earlier hope that one fork build would unblock several is dead.

  • 2026-08-11 · blocked → screened

    Screened on min-62bf73d (its minimum build, anchor-calibrated: +2.8%/+0.9% vs fleet baseline, identical tau2 capability). FIT ok 12s, guard 3/4 — coherence and toolcall pass, NEEDLE FAIL at depth 8000 (empty content ~5.4k tokens). Smoke: timed out with 2/5 done, the completed tasks taking median 108 turns vs the 122B's 15 on the same harness. VERDICT LEANS IMPLEMENTATION, NOT MODEL: an independent r/LocalLLaMA run on day-1 llama.cpp support measured 3/3 needle retrieval at depths up to 832K tokens (YaRN-stretched), while our failure sits at 8k INSIDE the native 131K window; our build is the literal merge commit of the arch support; and same-week upstream bugs exist in exactly this arch's attention-metadata handling (#26894 open, #26873 open). Discriminators run: no drafter in our setup (not the #26894 spec-decode class), no rope overrides (not a config accident), and our GGUF's sliding_window_pattern is the SCALAR form (dodges the known array crash) — making our symptom UNREPORTED upstream, possibly ROCm/gfx1151-specific (all filed issues are CUDA). RE-SCREEN TRIGGER: bump the min build once #26894/#26873 fixes land and re-run the guard + smoke; an upstream issue dossier for our symptom is prepared for the operator to file in their own words.

  • note

    Parent vs. artifact: this is NOT a finetune riding on a parent's benchmark reputation - vmlinux's repo is a third-party REQUANTIZATION of the exact same weights Meta published (base_model_relation: quantized, pointing at meta-models/Muse-Glimmer-30B). The agentic benchmark table belongs to that full-precision release. Meta's OWN quants (their official GGUF repo's K-Quant-Dynamic and K-Quant-17GB) disclose a measured 0.2% / 1.0% average degradation across 15 benchmarks - a real quant-specific number. vmlinux's ROCmFPX quant has no equivalent: its README's "Validation" section is a 16-token single-turn smoke test plus short throughput runs, explicitly captioned "not a formal benchmark." So the size of the benchmark's gap to reality is known for Meta's quants and unknown for this one. Stock-support answer: NO for this artifact. ROCmFPX predates upstream Muse support and needs a patch porting it in; the patch, base commit (00d54526e...), and upstream Muse commit (62bf73d25c...) are all named in the repo and independently verifiable (the Muse commit is real and merged upstream as ggml-org/llama.cpp#26841). The fork requirement is about the FP4/FP8 TENSOR FORMAT, not the model architecture - Muse Glimmer itself loads fine on stock llama.cpp via ordinary quants from other repos. Quant-equivalence hazard: no plain Q4_K_M exists anywhere sighted - not in this repo (only Q4_0_ROCMFP4_STRIX/_COHERENT and Q8_0_ROCMFPX), not in Meta's own GGUF repo (custom "kquant-17gb"/"kquant-dynamic" mixes, unnamed bpw), not in Unsloth's repo (Unsloth Dynamic UD-Q4_K_XL, a mixed-precision scheme, not plain Q4_K_M). Any of these could be screened on stock llama.cpp, but none would be a controlled quant-matched comparison against a field that is uniformly Q4_K_M - and the artifact actually staged here (ROCmFP4, 4.36 bpw dual-scale FP4) is a fourth, still-different method. Reddit source: the share link resolves to r/StrixHalo comments/1vknppz/ ("vmlinux/Muse-Glimmer-30B-ROCmFPX-GGUF | Hugging Face") but both the share link and the resolved permalink are login-walled for automated fetches, as anticipated. Web search for independent commentary, criticism, or additional tok/s reports on that specific thread turned up nothing beyond what the HF repo's own README/BUILD_RESULTS.md already state, so the Reddit layer added no information the source repo didn't already have. Strix Halo numbers, two different sources: vmlinux's own smoke test (gfx1151, ROCm, this exact ROCmFP4 quant) - 113.7 tok/s prompt / 14.9 tok/s decode alone, 28.3 tok/s decode with DFlash 6-token speculative window (2.07x, 35.8% draft acceptance). Separately, AMD's own blog for the OFFICIAL (non-ROCmFPX) release quotes up to 24 tok/s on Ryzen AI Max+ 395 and up to 53 tok/s on a Radeon AI PRO R9700 with DFlash - a different quant, different exact hardware, not directly comparable to vmlinux's number. Vision audit (2026-08-10): the parent DOES ship a real vision path - merged PR #26841 adds a text tower, perception encoder, and DFlash drafter together, and unsloth/Muse-Glimmer-30B-GGUF (the same repo the staged UD-Q4_K_XL came from) publishes matching mmproj files: mmproj-kquant.gguf (1,400,328,928 bytes), mmproj-Muse-Glimmer-30B-Q8_0.gguf (2,051,685,088 bytes), and a BF16 (3,849,173,728 bytes) - none staged. But the runtime claim two entries up needs a CORRECTION, not just an addition: PR #26841 merged as commit 62bf73d25c53b8161f8a22894d4f90c4aebbd7d0 on 2026-08-10 (today); our deployed build is pinned to commit 3653e6d (2026-08-07 22:35:52+0200), three days earlier, and 62bf73d is confirmed NOT an ancestor of our build via `git merge-base --is-ancestor`. Checked directly against the source tree on aihydra: "muse-glimmer"/"glimmer" appears nowhere in src/llama-arch.cpp on our checkout. So "the Muse-Glimmer architecture ITSELF is merged into stock llama.cpp... ordinary K-quant GGUFs... would run today" is upstream-true but NOT true of our deployed binary - the K-quant GGUF would currently fail to load at all, text or vision, until the build advances past 62bf73d. Blocked for a different, narrower reason than previously recorded (build staleness, not the ROCmFPX/quant-equivalence issues above) - update the rebuild step before the next screening attempt.

Nemotron-Cascade 2

cascade2-30b · 30B / 3B active · moe · Q4_K_M

Gold-medal competitive-programming reasoning at a tiny active-parameter count - the reasoning-per-active-parameter outlier of the field.

in gate since 2026-08-10 (3d)
gate history (3)
  • 2026-06-20 · listed

    Listed as the reasoning specialist spot-check - gold-medal competitive-programming performance at 3B active is an outlier worth confirming or refuting.

  • 2026-08-10 · listed → acquired

    Q4_K_M downloading from bartowski/nvidia_Nemotron-Cascade-2-30B-A3B-GGUF.

  • 2026-08-10 · acquired → screened

    Screened: FIT ok (16s), all four guards pass including needle at 8000. Smoke 0.20 with 3 of 5 tasks CUT at the 200-step limit — a non-termination tendency on agentic multi-turn work, consistent with a reasoning specialist outside its domain. Needle pass + agentic flail is the opposite profile to LFM2/Gemma.

  • note

    Vision audit (2026-08-10): parent Nemotron-Cascade-2-30B-A3B is text-only by NVIDIA's own model card - no image input, no vision variant anywhere in the line. Staged artifact (nvidia_Nemotron-Cascade-2-30B-A3B-Q4_K_M.gguf, arch nemotron_h_moe) carries no vision tensors and none exist to add. Verdict: text-only, not a vision-lane candidate.

Gemma 4 26B-A4B

gemma4-26b · 26B / 3.8B active · moe · Q4_K_M

The efficient fast-daily candidate: ~85 t/s at near-31B quality. If it holds up it is a far cheaper reflex tier than anything currently deployed.

in gate since 2026-08-10 (3d)
gate history (3)
  • 2026-06-20 · listed

    Speed/quality ratio is the best claimed in the field for a reflex tier.

  • 2026-08-10 · listed → acquired

    UD-Q4_K_M, 16 GB, downloaded. Tool-call format needs checking - flagged at listing time as the known risk.

  • 2026-08-10 · acquired → screened

    Screened: FIT ok (13s), coherence/toolcall/isolation pass, needle retrieval FAILED at depth 8000. Smoke 0.75 over 4 scored tasks (one task did not score — unexamined). Reflex-tier speed claim not yet measured; the needle failure caps usable depth pending a retrieval-depth sweep.

  • note

    Vision audit (2026-08-10): parent is multimodal (text+image, all Gemma 4 sizes) - confirmed on our exact staged artifact too: the downloaded gemma-4-26B-A4B-it-UD-Q4_K_M.gguf reads general.architecture=gemma4 and carries the "image-text-to-text" tag. Our llama.cpp build (commit 3653e6d) registers LLM_ARCH_GEMMA4 and clip.cpp implements PROJECTOR_TYPE_GEMMA4V for this size class - runtime support is real, not aspirational. What's missing is the mmproj file: unsloth/gemma-4-26B-A4B-it-GGUF (the repo we pulled the main weights from) publishes mmproj-F16.gguf (1,193,058,784 bytes) alongside it, and we have not staged it. One download away from vision-ready. Caution: a live upstream issue (#21402) reports a SIGABRT in clip_model_loader::load_tensors when loading the Gemma4 mmproj on CUDA - unverified whether that reproduces on our ROCm/Vulkan build, so budget a debug pass rather than assume a clean load on first try.

GLM-4.7-Flash

glm-47-flash · fits-tier · moe · Q4_K_M

Strong agentic and coding lineage with ~40 t/s reported on ROCm; the GLM line has a track record on tool use that none of our incumbents share.

in gate since 2026-08-10 (3d)
gate history (3)
  • 2026-06-20 · listed

    Agentic/coding lineage worth testing against the incumbents.

  • 2026-08-10 · listed → acquired

    Q4_K_M downloading, matched to the field's quant.

  • 2026-08-10 · acquired → screened

    Screened: FIT ok (16s), coherence/toolcall/isolation pass, needle FAILED at 8000. Smoke completed ZERO tasks in the full 60-minute budget — no scores at all. That pattern (guard toolcall passes, real agentic loop produces nothing) smells like a chat-template or tool-format mismatch with the harness rather than pure model weakness; needs diagnosis before any verdict. Not rejected — a harness mismatch is not measured evidence against the model.

  • note

    Vision audit (2026-08-10): parent is text-only - Zhipu's vision line is the separate GLM-4.5V model; GLM-4.7-Flash is agentic/coding-focused text. Confirmed on our staged artifact too: GLM-4.7-Flash-Q4_K_M.gguf reads general.architecture=deepseek2 (GLM's MoE reuses llama.cpp's DeepSeek-V2/V3 arch code) with no vision tags present. Verdict: text-only, not a vision-lane candidate.

Laguna S 2.1 (poolside)

laguna-s-21 · 118B / ~8B active · hybrid · Q4_K_M

Terminal-Bench 70.2% against Qwen's 41.6% - the largest claimed agentic gap of any candidate. Hybrid global+SWA MoE with 36/48 SWA layers means a very small KV, and Q4_K_M at 75.2 GB fits with MORE headroom than the current 122B.

in gate since 2026-08-10 (3d)
gate history (5)
  • 2026-07-21 · listed

    Released 2026-07-21 with a Terminal-Bench number far above our incumbent, at a footprint that fits.

  • 2026-08-10 · listed → acquired

    Q4_K_M downloading. Acquisition costs nothing while the GPU is busy, even though it cannot run yet.

  • 2026-08-10 · acquired → blocked

    Needs llama.cpp #25165 for upstream support; fork-only until it merges. Its SWA class also shares the #25913 checkpoint problem, where our sidecar fix applies. Available on OpenRouter at $0.10/$0.20 per M for feel-testing in the meantime.

  • 2026-08-10 · blocked → blocked

    Blocker RESOLVED UPSTREAM but not yet in our deployed binary: PR #25165 merged 2026-08-01 as 1f66c3ce ("Add support for Laguna XS.2 & M.1"), verified NOT an ancestor of our stock 3653e6d build. Under the new minimum-version build policy the requirement is now ON MAIN, not fork-only — and the min-62bf73d binary being built for Muse Glimmer necessarily contains it. Two open questions before un-gating: whether the merged PR's arch coverage includes the S 2.1 variant (title names XS.2 and M.1; our GGUF header reads arch=laguna), and the #25913-class SWA checkpoint behaviour. Tonight's screen on stock will record the expected unknown-architecture failure as evidence; a FIT probe on the new binary decides the variant question.

  • 2026-08-10 · blocked → screened

    SCREENED ON STOCK — the blocker was already resolved in the deployed binary. FIT ok (61s load), ALL FOUR guard checks pass (needle retrieval at depth 8000 included), smoke 0.80 over 5/5 tasks in 7 minutes with nothing cut. This empirically corrects the previous entry's ancestry check, which failed on an unfetched commit and was misread as "support absent": PR #25165's arch support IS in build 3653e6d, and it covers the S 2.1 variant. Strongest first screen of the five new candidates. Smoke is a smoke test, not a ranking (noise ~0.40, clm-0036).

  • note

    Vision audit (2026-08-10): parent is text-only - poolside's own materials describe Laguna S 2.1 as a coding/agentic model with no vision capability in the line. Moot for this candidate while the #25165 architecture blocker stands anyway, but worth recording so a future revisit doesn't have to re-derive it: no vision path exists to unlock even once the arch lands upstream.

LFM2-24B-A2B (Liquid)

lfm2-24b · 24B / 2B active · moe · Q4_K_M

A speed outlier with a huge prefill advantage. It is the field's 'is fast enough also good enough?' probe - the cleanest test of whether our latency problem is worth trading quality for.

in gate since 2026-08-10 (3d)
gate history (3)
  • 2026-06-20 · listed

    The speed-versus-quality probe for the whole field.

  • 2026-08-10 · listed → acquired

    Q4_K_M downloading from LiquidAI/LFM2-24B-A2B-GGUF.

  • 2026-08-10 · acquired → screened

    Screened: FIT ok (12s), coherence/toolcall/isolation pass, needle retrieval FAILED at depth 8000 — the gpt-oss-class long-context cliff. Smoke 0.80 over 5/5 in 7 minutes. Fast and agentically plausible at short context; the needle failure caps its usable depth pending a retrieval-depth sweep.

  • note

    Vision audit (2026-08-10): parent (base LFM2-24B-A2B) is text-only. Liquid AI ships vision as a SEPARATE family, LFM2-VL (LFM2.5-VL-1.6B/450M, LFM2-VL-3B), built on smaller backbones - not as a variant of the 24B text model staged here. Verdict: text-only, not a vision-lane candidate.

benched (4)

Nemotron-3-Super-120B-A12B

nemotron3-super · 120B / ~12B active · hybrid · Q4_K_M

Purpose-built for agentic reasoning and tool use, which is exactly Warden's workload. Hybrid Mamba-2 + LatentMoE + MTP promises cheaper long-context KV.

in gate since 2026-08-09 (4d) · model page
gate history (4)
  • 2026-06-20 · listed

    Only model in the field designed for the workload Warden actually runs.

  • 2026-06-27 · listed → acquired

    Forced to IQ4_XS: Q4_K hit 77 GiB at 4K and OOMed by 16K on the 96 GB box.

  • 2026-06-27 · acquired → screened

    Phase A: slowest at ~28s, 100% tool-call gate. MTP REJECTED - Mamba is incompatible with draft-mtp, a deliverability finding not a config issue.

  • 2026-08-09 · screened → benched

    Ran the tau2 arms. NOT quant-equivalent to peers - still IQ4_XS while everything else is Q4_K_M, so every Nemotron-vs-peer comparison to date is confounded. aihydra's 104 GiB GTT lifts the constraint that forced it; Q4_K_M re-run is outstanding.

    evidence: clm-0035 clm-0036

  • note

    Vision audit (2026-08-10): parent is text-only, confirmed directly - Nemotron-3-Super-120B-A12B's own documentation states it supports text input only, no image/vision modality anywhere in the line. Staged artifact (arch nemotron_h_moe) carries no vision tensors. Verdict: text-only, not a vision-lane candidate.

Qwen3.6-35B-A3B

qwen36-35b · 35B / 3B active · moe · Q4_K_M

The true incumbent: what Warden actually ran, with best-in-class tool-calling at ~96-100 t/s. The status-quo baseline every challenger must beat.

in gate since 2026-08-09 (4d) · model page
gate history (4)
  • 2026-06-20 · listed

    Listed as the status-quo incumbent: whatever replaces it has to beat a model with proven best-in-class tool-calling in live use.

  • 2026-06-27 · listed → acquired

    Already in production on the Mac mini, so it was staged on the new box rather than acquired - the incumbent needs no justification to be present.

  • 2026-06-27 · acquired → screened

    Phase A: ~11s time-to-correct, 100% tool-call gate.

  • 2026-08-09 · screened → benched

    Throughput, FA, ngram speculation and tau2 arms run. Speculation is a dead end on this model (no MTP head).

    evidence: clm-0026

  • note

    Vision audit (2026-08-10): parent IS vision-capable - natively multimodal per Qwen's own materials, processing images/documents/video as a core architectural capability, not a bolt-on. Confirmed on our staged artifact: Qwen3.6-35B-A3B-UD-Q4_K_XL.gguf reads general.architecture=qwen35moe, carries the "image-text-to-text" tag, and has qwen35moe.rope.dimension_sections=[11,11,10,0] - the same MRoPE sectioning implemented natively in src/models/qwen35moe.cpp. unsloth/Qwen3.6-35B-A3B-GGUF (our source repo) publishes mmproj-F16.gguf (899,283,680 bytes), not staged. Runtime: LLM_ARCH_QWEN35MOE is registered in our build (commit 3653e6d), and clip.cpp's PROJECTOR_TYPE_QWEN3VL path is arch-agnostic on the LLM side. Verdict: mmproj-download-away, not text-only - this is the direct MoE sibling to qwen35-122b's vision path and shares the true incumbent's own architecture, so a vision test here would double as an incumbent-vs-challenger comparison rather than a one-off.

gpt-oss-120b

gpt-oss-120b · 117B / ~5B active · moe · Q4_K_M

The migration's planned daily-driver upgrade and the speed bar for the whole field at ~55 t/s Vulkan.

in gate since 2026-08-08 (5d) · model page
gate history (4)
  • 2026-06-20 · listed

    Listed as the planned daily-driver upgrade and as the speed control the rest of the field is measured against.

  • 2026-06-27 · listed → acquired

    Downloaded at Q4_K_M for the Phase 3C field, matched to the incumbent's quant so the comparison would be clean.

  • 2026-06-27 · acquired → screened

    Phase A: fastest in the field at ~5s time-to-correct, 100% tool-call gate.

  • 2026-08-08 · screened → benched

    Throughput rows recorded. NOT cleared for use: returns verbatim answers from previous requests at --parallel 1, reproduced 6/6, and scored 0.00 on tau2 thinking-off. Kept at bench rather than rejected because the staleness mechanism is undiagnosed and may be a harness fault rather than the model.

    evidence: clm-0025

  • note

    Vision audit (2026-08-10): parent is text-only by OpenAI's own design and documentation - gpt-oss ships with no vision encoder in the line, and OpenAI explicitly points multimodal use cases at their hosted API instead. Staged artifact (arch gpt-oss) carries no vision tensors. Verdict: text-only, not a vision-lane candidate - and separately still not cleared for use per the staleness finding above.

Qwen3.5-122B-A10B (MTP)

qwen35-122b · 122B / 10B active · moe · Q4_K_M

Production incumbent and the quality ceiling that still fits. 262K context; MTP measured adopt-worthy at +50% decode with a ~10% cold-path prefill tax.

in gate since 2026-08-08 (5d) · model page
gate history (4)
  • 2026-06-20 · listed

    Selected as the quality-ceiling candidate that fits 96GB at Q4.

  • 2026-06-27 · listed → acquired

    Downloaded UD-Q4_K_M for the Phase 3C field.

  • 2026-06-27 · acquired → screened

    Phase A agentic eval: ~11s time-to-correct, 100% tool-call gate. Matches the 35B despite 3.5x the params, so latency is reasoning-bound not decode-bound.

  • 2026-08-08 · screened → benched

    Fully characterised on performance: depth to 204.8k, KV quant both ways, FA both ways, speculation curve, cache trace, grammar ceiling.

    evidence: clm-0022 clm-0027

  • note

    Vision audit (2026-08-10): parent IS vision-capable, and not as an afterthought - Qwen bills Qwen3.5-122B-A10B as natively multimodal (early-fusion trained on text+image+video tokens), and this is confirmed directly on our staged artifact: Qwen3.5-122B-A10B-UD-Q4_K_M reads general.architecture=qwen35moe, carries the "image-text-to-text" tag, and has qwen35moe.rope.dimension_sections=[11,11,10,0] - the same MRoPE sectioning llama.cpp's dedicated qwen3vl/qwen3vlmoe arches use for spatial image positions, implemented natively inside src/models/qwen35moe.cpp on our build rather than requiring a separate VL arch tag. unsloth's Qwen3.5-122B-A10B-GGUF (the repo we pulled the main weights from) publishes a matching mmproj-F16.gguf (908,724,960 bytes) that we have not staged. Runtime: LLM_ARCH_QWEN35MOE is registered in our build (commit 3653e6d) and clip.cpp implements PROJECTOR_TYPE_QWEN3VL, which is arch-agnostic on the LLM side. Verdict: mmproj-download-away, not text-only - this is the strongest vision-lane candidate already on the box, and it is our current production incumbent. Caveat: an open upstream issue (#21268, CUDA-only in its report) describes CLIP-graph operator gaps with this exact model+mmproj pairing causing OOM rather than a clean load - unverified on ROCm/Vulkan, so budget a debug pass before treating this as a guaranteed win.

Out of the pipeline — every exit carries its evidence

blocked (0)

none

rejected (0)

none

generated from the candidates collection · a gate change without a reason cannot be recorded (schema-enforced) · a rejection without evidence fails the build · cards sitting over 30 days in acquired/screened carry the amber dot — work-queue honesty without publishing a work queue