Candidates
Every model under consideration, and the recorded reason for every gate transition. The standing rule this board enforces: nothing leaves the pipeline without a measurement — vendor claims and role assumptions are admissible reasons to list a candidate and never grounds to drop one, so rejected cards show their evidence. Expand a card for its full gate history. NPU-lane cards run a different artifact at a different quantisation on a different runtime — their numbers never transfer to or from the GPU lane (4 on the board).
listed (1)
LLaDA 2.2-flash (inclusionAI)
The first credible open diffusion LLM at scale: generation by parallel iterative denoising instead of autoregression, plus a Levenshtein-editing self-correction pass. That shifts the decode bottleneck from memory bandwidth toward compute, which is precisely the trade this bandwidth-starved silicon wants — and its vendor-conceded capability gap makes it a strong test case for Wh-per-correct-answer as the metric that should decide whether "10x faster" means anything.
gate history (1)
- 2026-08-15 · listed
Listed on release-week signal (r/AIDeveloperNews thread 1vk21p9). Vendor claims: fully open 100B MoE diffusion model, agent-oriented, 700+ TPS, "Levenshtein Editing" self-correcting decode. A thread commenter quotes the model card at SWE-bench Verified 49.28 vs 78.8 for Qwen3.6-27B — unverified secondhand, but if directionally right the speed is bought with a large capability gap, which is exactly the Wh-per-correct question this lab exists to measure. NOT acquired: three gates must open first. (1) Upstream support — llama.cpp master today lists only the older llada/llada-moe arches; nothing matching a 2.2-flash 100B MoE, and no open PR found. (2) A servable screen path — llama.cpp diffusion runs via a separate CLI with no server endpoint, so GUARD/SMOKE/tau2 have nothing to attach to; either server support lands or a defensible CLI-based screen variant gets written into the protocol first. (3) A protocol surface definition — "decode t/s" is not well-defined for parallel refinement (tokens finalised per denoising step is not comparable to autoregressive decode), so diffusion results need their own surface class before any number can sit next to the existing matrix. Re-check trigger: llama.cpp gaining a LLaDA-2.x arch, tracked in the lab watchlist alongside the bailingmoe3/maple support watches.
- note
Paradigm note for future screening: diffusion decode quality and speed both depend on the step count / block size chosen at inference time — a configuration axis autoregressive models do not have. Any eventual screen must fix and record those parameters per cell, or the numbers are not reproducible. Modality: text; no vision component claimed.
acquired (0)
none at this gate
screened (15)
LLaDA2.2-flash (inclusionAI / Ant Group)
The field's first block-diffusion model. It tests whether a scaled non-autoregressive (diffusion) LLM can act as a viable tau2 agent on portable AMD hardware, and what the diffusion paradigm costs relative to the autoregressive field — a question no autoregressive candidate can answer.
gate history (2)
- 2026-08-22 · listed
A genuinely novel offshoot: inclusionAI LLaDA2.2-flash, a 100B-A13B block-diffusion MoE with Levenshtein editing, is the first non-autoregressive agent model considered for the field. It needs a diffusion runtime (diffuse-cpp) with GPU offload and an inter-step KV cache before it can be screened at all.
- 2026-08-22 · listed → screened
Screened on aihydra via the headbouyJB/diffuse-cpp fork (GPU-resident KV cache + cross-turn reuse + full Levenshtein editing). tau2 airline 5-task screen (task_ids 0-4): 5/5 solved, mean reward 1.000, real multi-turn tool use, all tasks converged in 8-10 turns (run-0553). Energy-joined (eng-0226): 12.42 Wh per correct answer. A screen, not the full 26-task arm.
GLM-4.7-Flash
Strong agentic and coding lineage with ~40 t/s reported on ROCm; the GLM line has a track record on tool use that none of our incumbents share.
gate history (4)
- 2026-06-20 · listed
Agentic/coding lineage worth testing against the incumbents.
- 2026-08-10 · listed → acquired
Q4_K_M downloading, matched to the field's quant.
- 2026-08-10 · acquired → screened
Screened: FIT ok (16s), coherence/toolcall/isolation pass, needle FAILED at 8000. Smoke completed ZERO tasks in the full 60-minute budget — no scores at all. That pattern (guard toolcall passes, real agentic loop produces nothing) smells like a chat-template or tool-format mismatch with the harness rather than pure model weakness; needs diagnosis before any verdict. Not rejected — a harness mismatch is not measured evidence against the model.
- 2026-08-21 · screened → screened
HG-004 Phase-1 boundary diagnostic captured one real raw llama-server HTTP 200 response through the pinned installed Tau2/native-template task-1 path. The response was reasoning-only with empty assistant content, no tool_calls, finish_reason length, and Tau2 Phase-2 was not launched. This is non-scoring, non-capability diagnostic evidence only: no tau2 score, throughput metric, production recommendation, or GLM accept/reject verdict is admitted.
- note
Vision audit (2026-08-10): parent is text-only - Zhipu's vision line is the separate GLM-4.5V model; GLM-4.7-Flash is agentic/coding-focused text. Confirmed on our staged artifact too: GLM-4.7-Flash-Q4_K_M.gguf reads general.architecture=deepseek2 (GLM's MoE reuses llama.cpp's DeepSeek-V2/V3 arch code) with no vision tags present. Verdict: text-only, not a vision-lane candidate.
Ornith-1.5-35B
Ornith-1.5 (ornith-ai, MIT, released ~2026-08-19) is the successor to Ornith-1.0-35B, same qwen35moe architecture family (hybrid linear/full attention, head_dim 256) and 35B/A3B scale, shipped as Q4_K_M official GGUF. Listed to measure the version jump at matched quant class (Q4_K_M for 1.5 vs UD-Q4_K_XL for 1.0's comparability arm — same family, not identical quant type) and to bound the screening bar before any full- bench or energy join commitment.
gate history (3)
- 2026-08-21 · listed
Ornith-1.5 is the direct successor to Ornith-1.0-35B on the same qwen35moe architecture, published as MIT-licensed GGUF by the vendor. HF API confirms the repo exists. Screened at Q4_K_M (official) — same quant class as the Ornith-1.0 UD-Q4_K_XL comparability arm, though not identical quant type.
- 2026-08-21 · listed → acquired
Downloaded to aihydra ~/models/ornith-15-35b/ and sha256-verified against the sidecar: Q4_K_M (official repo) 21,713,462,848 bytes, sha256 ca6ea26329c88b78ffd90a85163be2e746c2fafd1024f56db47e499f117f9a7f.
- 2026-08-21 · acquired → screened
HO-011 reviewer-admitted screening at Q4_K_M, ROCm0 3653e6d, c32768, f16 KV, --parallel 1. FIT: gtt=20843 MiB (comfortable). GUARD: 4/4 (coherence, tool call, needle at 8000 tokens, isolation skipped under -np 1). SMOKE: 5-task tau2 airline, openrouter/anthropic/ claude-haiku-4.5 simulator, temp 0, seed 42 — 4/5, mean_reward 0.800, 74 tool-call messages, 0 empty assistant turns, 0 empty-argument tool calls, 0 infra errors. Throughput at d262144: 142.38 t/s prefill / 20.99 t/s decode (N=5 fresh-process reps, stddev 0.264 / 0.010). Full-26 tau2 immediately after a fresh 4/4 guard: 8/25 evaluated, mean_reward 0.320, 1 infra error (task 20, JSON parse failure after 4 retries), 5 TOO_MANY_ERRORS terminations, 2 MAX_STEPS terminations, 8 context-overflow errors in the server log from conversation buildup exceeding c=32768. This is a sharp regression from Ornith-1.0-35B UD-Q4_K_XL (22/26, 0.846) — a version-replacement step-back on the standard tau2 airline suite. Energy: not joined; run-meta windows recorded for later batch join.
evidence: clm-0100
- note
VERSION-REPLACEMENT STEP-BACK: Ornith-1.5 scores dramatically lower on the standard tau2 airline suite than its predecessor Ornith-1.0-35B (8/25 mean 0.320 vs 22/26 mean 0.846). The quant types differ (Q4_K_M vs UD-Q4_K_XL) but both are Q4-class, and the gap is far outside any quant-effect envelope seen in this lab — this is a model-quality regression, not a quant artefact. The context-overflow errors (8 total) undercut the already-low score further: a wider serving context might have recovered some tasks but would not close the gap. Do not promote this candidate to full bench without a separate card and explicit justification — the screening evidence says no. MTP/NextN: the server log shows unused blk.40.nextn tensors, matching the Qwen3.5-MoE MTP-architecture pattern seen on Ornith-1.0. No MTP arm was activated for this screen; the MTP/EOS-cliff watch from Ornith-1.0 (clm-0055) applies unchanged to this candidate. ⚠ MTP HEAD CAVEAT (2026-08-22): community report (r/StrixHalo) confirms Ornith-1.5 ships with a broken MTP head — nextn_predict_layers=1 is declared but the head tensors produce corrupted generations when activated. Our 8/25 result was measured with MTP INACTIVE and is not revised by this finding. Two implications: (1) any run that does activate MTP on the current artifact will fail, and (2) if the vendor publishes a corrected GGUF, the capability picture could change and re-screening is warranted. Track via watchlist.
LFM2-24B-A2B (Liquid)
A speed outlier with a huge prefill advantage. It is the field's 'is fast enough also good enough?' probe - the cleanest test of whether our latency problem is worth trading quality for.
gate history (4)
- 2026-06-20 · listed
The speed-versus-quality probe for the whole field.
- 2026-08-10 · listed → acquired
Q4_K_M downloading from LiquidAI/LFM2-24B-A2B-GGUF.
- 2026-08-10 · acquired → screened
Screened: FIT ok (12s), coherence/toolcall/isolation pass, needle retrieval FAILED at depth 8000 — the gpt-oss-class long-context cliff. Smoke 0.80 over 5/5 in 7 minutes. Fast and agentically plausible at short context; the needle failure caps its usable depth pending a retrieval-depth sweep.
- 2026-08-19 · screened → screened
HG-003's corrected tokenizer-measured protocol failed at the minimum depth-2048 early-placement cell. The bounded stop skipped late placement and all performance work; safe depth is below the practical 4096-token promotion threshold, so no promotion or performance claim is admitted.
- note
Vision audit (2026-08-10): parent (base LFM2-24B-A2B) is text-only. Liquid AI ships vision as a SEPARATE family, LFM2-VL (LFM2.5-VL-1.6B/450M, LFM2-VL-3B), built on smaller backbones - not as a variant of the 24B text model staged here. Verdict: text-only, not a vision-lane candidate. HG-003 retrieval promotion is now closed: after two runner-invalid attempts, the corrected scientific run failed the minimum 2048-token early-placement cell. This narrows the retrieval envelope below the practical 4096-token promotion threshold; it does not support a broader model-quality claim.
NVIDIA Nemotron 3.5 Lightning 30B-A3B
NVIDIA's own model card confirms a hybrid Mamba-2 + Attention + MoE "execution layer" design (128 routed experts, 6 active + 1 shared, native MTP baked into the checkpoint, 1M token context) pitched at exactly Warden's workload: long-running agent loops doing tool calling, state management and result validation. It is the direct small-footprint sibling of the already-benched nemotron3-super, and llama.cpp gained Nemotron MTP support (#26725, merged 2026-08-10) after nemotron3-super's "MTP incompatible with Mamba" finding was recorded - that finding may now be stale for this whole family. An official Q4_K_M GGUF exists at 23.69 GiB, no fork required, and NVIDIA's own benchmark table gives a real (if middling) agentic-axis result: SWE-bench Verified 51.56, Terminal-Bench 2.1 24.58, tau3-bench-Banking 9.28 - all below our incumbent Qwen3.6-35B-A3B on NVIDIA's own comparison, but a genuine measurement all the same.
gate history (5)
- 2026-08-12 · listed
Surfaced via an operator-shared r/AIDeveloperNews link ("NVIDIA has launched Nemotron 3.5 Lightning") plus a follow-up r/StrixHalo post on a community ROCmFP4 requant with hardware-matched Strix Halo numbers. Independently verified: a real NVIDIA release (nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-{BF16,NVFP4}, OpenMDW-1.1 licence, released 2026-08-11, commercial use permitted), with an official ggml-org GGUF repo publishing a stock Q4_K_M (25,430,738,944 bytes / 23.69 GiB) alongside BF16/NVFP4/Q8_0 - no fork needed for the standard quant. Architecture (nemotron_h_moe: hybrid Mamba-2 + Attention + MoE) has been in mainline llama.cpp since PR #20411 (2026-03-11), long before any of our deployed builds, and native MTP support for the Nemotron family landed via #26725 merged 2026-08-10T08:25Z - BEFORE our 62bf73d build (2026-08-10T11:07Z), so that build should load this model with MTP working. Our older 3653e6d build (2026-08-07) predates the MTP commit and would load the model (base arch is old enough) but without the MTP speedup. DFlash-specific draft-model support (#26905) merged 2026-08-11T13:16Z, after 62bf73d, so DFlash itself is NOT confirmed available on either resolvable deployed build - 77a9a66eb did not resolve against ggml-org/llama.cpp mainline history via the GitHub API and its date is unconfirmed; treat its DFlash/MTP support as unverified until the operator confirms its lineage. NVIDIA's own published benchmark table places this model BELOW our incumbent Qwen3.6-35B-A3B on nearly every agentic metric (SWE-bench Verified 51.56 vs 70.12, SWE-bench Multilingual 39.33 vs 63.40, Terminal-Bench 2.1 24.58 vs 44.38, tau3-bench-Banking 9.28 vs 10.52, BrowseComp 36.97 vs 48.74, PinchBench 85.37 vs 88.07) - listing on the strength of a real, vendor-published agentic measurement and clean upstream architecture support, not a capability-leadership claim.
- 2026-08-13 · listed → acquired
Official ggml-org Q4_K_M downloaded to aihydra (25,430,738,944 bytes, exactly the listed size) and sha256-verified against the HF LFS oid (6110e2e2e6cd324e6ee69ddced5a6b34fad6c94ca9827222a1e420fb92e3c90b). Mirrored to the NAS gguf-library.
- 2026-08-13 · acquired → screened
Screened same night on build min-62bf73d (62bf73d25) — the newest protocol-registered build that post-dates Nemotron MTP #26725 — under the house FIT/GUARD/SMOKE tier (Aug-12 protocol, pinned haiku-4.5 simulator). FIT ok, loaded in 8s at 32K. GUARD 4/4 PASS including needle at depth 8000 (tok=5370 'chartreuse-viper-88'). SMOKE 5-task tau2 airline: completed=5/5, mean=0.00, cut_at_max_steps=3, 7 min wall — the cascade2-30b profile (needle pass + agentic flail with non-termination), and directionally consistent with NVIDIA's own below-incumbent agentic table. MTP caveat: 62bf73d loads the model but the server logged the layer-52 MTP head as unused ("model has unused tensor blk.52.nextn.enorm.weight ... -- ignoring", likewise hnorm/eh_proj/ shared_head_norm), so this screen ran the base decode path with NO MTP engagement; the "does the small sibling's MTP work now" question from listing is still open and needs whatever draft/spec flag or newer build actually wires the nextn head. Smoke is a cliff-detector, not a ranking (noise ~0.40, clm-0036). Run-meta: screening/nemotron35-lightning-30b 2026-08-13T03:06:34Z->03:14:38Z (484s, rc=0).
- 2026-08-19 · screened → screened
HG-002 supersedes the inactive-MTP caveat from the initial screen: llama.cpp 7077abb explicitly engaged the checkpoint's NextN head on ROCm. N1 was the best bounded arm at 75.5813 tok/s, 1.1569x the same-session plain control, with a 4/4 guard and 15 clean varied sanity prompts. This earns only the conditional five-task smoke; the candidate remains screened and is not promoted. N3 regressed to 0.9433x and is rejected; no full tau2 is admitted.
evidence: clm-0087
- 2026-08-19 · screened → screened
HG-002 Phase C closed without promotion. The corrected matched plain arm reproduced 4/5, mean 0.80. N1 scored 1.0 on its first four tasks but its fifth response was an empty AssistantMessage classified as infrastructure error, leaving the arm invalid with no aggregate mean. The measured Phase A throughput lift is real but insufficient to displace plain for capability.
evidence: clm-0088
- note
Marketing-claim caveat: the r/AIDeveloperNews post (linking a non-canonical promo site, aideveloper44.com, not nvidia.com) claims the model was "trained specifically for popular agent harnesses (like OpenClaw and Hermes Agent)". This is NOT corroborated anywhere in NVIDIA's own model card, which describes only generic "agentic, tool-use" training data (synthetic CLI/web/coding agent corpora from gpt-oss-120b, Qwen3-Coder-480B, GLM-4.7-Flash) with no named-harness claim. Treat the OpenClaw mention as unverified vendor-adjacent marketing copy, not a primary-source fact - flagging because the coincidence with our own deployed harness is exactly the kind of detail that should be checked rather than taken at face value. Community Strix Halo numbers (r/StrixHalo, MrWidmore888, hardware: Framework Ryzen AI Max 395+, gfx1151, 128 GB unified memory - same hardware class as aihydra): these are measured on julianmb's ROCmFP4 community requant, NOT the official Q4_K_M, and are NOT directly inheritable as a screening result for that reason (same pattern as muse-glimmer-30b). Vulkan0 backend, FlashAttention, q8_0 KV cache, full unified-memory offload: STRIX_LEAN preset (Q4_0_ROCMFP4_STRIX_LEAN, ~4.38 bpw, 15.73 GiB) - pp512 1,299.7 t/s, tg128 85.6 t/s, wikitext-2 perplexity 5.9936±0.0358; FAST preset (~4.25 bpw, 15.66 GiB) - pp512 1,310.5 t/s, tg128 86.0 t/s; COHERENT preset (agentic/coding tuned, ~4.70 bpw, 16.74 GiB) - pp512 1,290.4 t/s, tg128 81.6 t/s. Poster notes Vulkan beat ROCm by ~21% on prompt processing on this box. These ROCmFP4 tensor types are NOT stock llama.cpp - loading them needs charlie12345/ROCmFPX (independently confirmed real on GitHub, same fork already carried as a blocker for muse-glimmer-30b) via julianmb's shoutout in the post. The official stock Q4_K_M path does not need this fork at all, so the fork question only applies if chasing this specific community quant rather than NVIDIA's own artifact. Fit: official Q4_K_M is 23.69 GiB - comfortable in either the 84 or 104 GiB GTT budget with large headroom for KV/context even before touching the 1M-token window (compare nemotron3-super, which was forced to IQ4_XS and still OOMed by 16K on the smaller box). Same architecture family as nemotron3-super (nemotron_h_moe) but roughly 4x smaller total / 4x smaller active params - this is a genuine niche question, not a clean duplicate: nemotron3-super occupies the large-agentic-reasoning slot (120B/12B active, already benched), while this candidate would sit in the reflex/executor tier alongside qwen38-27b and qwen36-27b-mtp - except MoE rather than dense, and with MTP genuinely native to the checkpoint rather than bolted on. Worth screening against BOTH neighbours: nemotron3-super (does the small sibling's MTP actually work now that #26725 is upstream, re-testing the "Mamba rejects MTP" finding on a Q4_K_M-clean artifact) and Qwen3.6-35B-A3B (the true incumbent this would need to beat on tool-calling, per NVIDIA's own numbers it currently does not).
Muse Glimmer 30B (vmlinux ROCmFPX quant)
Surfaced via r/StrixHalo share link pointing at a third-party ROCmFPX quant of Meta's Muse-Glimmer-30B. The operator's framing was that the PARENT model scores well at agentic benchmarks - verified true on meta-models/Muse-Glimmer-30B's own card: SWE-Bench Verified 76.0, SWE-Bench Pro 51.2, MCP Atlas 75.5, DeepSearch QA 74.6, τ3-Banking 23.5 - a real, Meta-published result for the full-precision base model. Worth listing on that basis even though it is not this quant's own number (see note).
gate history (6)
- 2026-08-10 · listed
Listed on the strength of the parent model's own published agentic benchmark table (SWE-Bench Verified 76.0, MCP Atlas 75.5, SWE-Bench Pro 51.2), confirmed on meta-models/Muse-Glimmer-30B's model card. The Reddit link itself was login-walled and added nothing beyond what the HF repo already documents.
- 2026-08-10 · listed → acquired
Downloaded Muse-Glimmer-30B-ROCmFP4.gguf (15,210,123,424 bytes) to aihydra for staging. Acquisition costs nothing while the GPU is busy, even though this specific artifact cannot run yet - see blocker.
- 2026-08-10 · acquired → blocked
The downloaded file uses ROCmFPX's experimental Q4_0_ROCMFP4_STRIX tensor type, which the repo's own README states "will not load in stock llama.cpp." It needs charlie12345/ROCmFPX (a fork) with ROCmFPX-Muse-Glimmer.patch applied to a pinned commit - same class of hazard as the ROCmFP4 precedent already carried for Laguna. The Muse-Glimmer architecture ITSELF is merged into stock llama.cpp (ggml-org/llama.cpp#26841), so ordinary K-quant GGUFs of this model (Meta's own official repo, or Unsloth's) would run today - but neither publishes a plain Q4_K_M, so a stock-compatible download would still not be quant-matched against our field. Staying blocked until either a literal Q4_K_M appears or the fork question is revisited.
- 2026-08-10 · blocked → blocked
REASSESSED on review: the block applies to the ROCmFPX ARTIFACT, not the model. The muse-glimmer architecture is merged in mainline llama.cpp (PR #26841), so the model IS screenable on our stock build via a standard quant - Unsloth publishes Muse-Glimmer-30B-UD-Q4_K_XL, the same quant family as our gpt-oss artifact. Downloading it as the screening artifact; the FPX file stays staged for a later performance-lever experiment (FPX-vs-standard on identical weights), pending the fork-build decision. Note the fork patch is MODEL-SPECIFIC per its own docs, so it does not amortise across other candidates - the earlier hope that one fork build would unblock several is dead.
- 2026-08-11 · blocked → screened
Screened on min-62bf73d (its minimum build, anchor-calibrated: +2.8%/+0.9% vs fleet baseline, identical tau2 capability). FIT ok 12s, guard 3/4 — coherence and toolcall pass, NEEDLE FAIL at depth 8000 (empty content ~5.4k tokens). Smoke: timed out with 2/5 done, the completed tasks taking median 108 turns vs the 122B's 15 on the same harness. VERDICT LEANS IMPLEMENTATION, NOT MODEL: an independent r/LocalLLaMA run on day-1 llama.cpp support measured 3/3 needle retrieval at depths up to 832K tokens (YaRN-stretched), while our failure sits at 8k INSIDE the native 131K window; our build is the literal merge commit of the arch support; and same-week upstream bugs exist in exactly this arch's attention-metadata handling (#26894 open, #26873 open). Discriminators run: no drafter in our setup (not the #26894 spec-decode class), no rope overrides (not a config accident), and our GGUF's sliding_window_pattern is the SCALAR form (dodges the known array crash) — making our symptom UNREPORTED upstream, possibly ROCm/gfx1151-specific (all filed issues are CUDA). RE-SCREEN TRIGGER: bump the min build once #26894/#26873 fixes land and re-run the guard + smoke; an upstream issue dossier for our symptom is prepared for the operator to file in their own words.
- 2026-08-15 · screened → screened
RE-SCREENED on Nathanw1014/strix-halo-llamacpp v0.6.2 (baf0025de portable tarball, checksum-verified against the GitHub release digest before use; two upstream cherry-picks on 3be50ccc/v0.6.1: Muse Glimmer support #26841 and a tool-call-after-EOM parsing fix #26879; registered in protocol.json). UD-Q4_K_XL, f16 KV, no drafter (baseline; v0.6.2's DFlash MTP support is unexercised here — a follow-up throughput experiment). FIT: loads clean in 15s, gtt_used 15,369 MiB (small model, no fit concern). One bench cell (pp1024/tg256, d0, 3 reps): pp~293.5 t/s / tg~14.29 t/s, driver-checked RADV STRIX_HALO (not CPU fallback) — broadly consistent with the release's own published pp512 375.6 / tg128 13.21 at d0, same driver/compiler lineage. GUARD 3/4 — coherence and toolcall pass, NEEDLE FAIL at depth 8000, REPRODUCES IDENTICALLY on this dedicated fork build. This narrows the 2026-08-11 "leans implementation, not model" hypothesis: that reading rested partly on our own build being a stale literal-merge-commit checkout, but v0.6.2 is Nathan's OWN purpose-built release for this exact model and shows the same symptom at the same depth. Does not confirm a genuine upstream/model defect either — whether v0.6.2's lineage carries the #26894/#26873 fixes referenced in the 2026-08-11 entry is unverified here — but it rules out "our checkout specifically is stale" as a sufficient explanation on its own. SMOKE, BY CONTRAST, IMPROVED DRAMATICALLY: 5/5 @ 1.000, VALID under the SMOKE gate (43 tool-call messages, 0 empty turns, 0 infrastructure errors), tasks completing in 105-331s with 2-17 tool calls each — against the 2026-08-11 screen's timeout at 2/5 done, median 108 turns. Whatever immaturity was driving the agentic-execution failure on min-62bf73d appears resolved on this build; the needle-retrieval cliff is a SEPARATE, still-open capability question that a clean SMOKE run does not paper over. Gate stays screened, not benched: protocol §9 does not let a guard failure be rescued by a good number elsewhere, and long-context retrieval remains broken regardless of how well tool-use now performs. RE-SCREEN TRIGGER, updated: confirm whether a future build fixes the needle regression specifically (still unresolved); the DFlash drafter throughput experiment is now unblocked and outstanding.
- note
Parent vs. artifact: this is NOT a finetune riding on a parent's benchmark reputation - vmlinux's repo is a third-party REQUANTIZATION of the exact same weights Meta published (base_model_relation: quantized, pointing at meta-models/Muse-Glimmer-30B). The agentic benchmark table belongs to that full-precision release. Meta's OWN quants (their official GGUF repo's K-Quant-Dynamic and K-Quant-17GB) disclose a measured 0.2% / 1.0% average degradation across 15 benchmarks - a real quant-specific number. vmlinux's ROCmFPX quant has no equivalent: its README's "Validation" section is a 16-token single-turn smoke test plus short throughput runs, explicitly captioned "not a formal benchmark." So the size of the benchmark's gap to reality is known for Meta's quants and unknown for this one. Stock-support answer: NO for this artifact. ROCmFPX predates upstream Muse support and needs a patch porting it in; the patch, base commit (00d54526e...), and upstream Muse commit (62bf73d25c...) are all named in the repo and independently verifiable (the Muse commit is real and merged upstream as ggml-org/llama.cpp#26841). The fork requirement is about the FP4/FP8 TENSOR FORMAT, not the model architecture - Muse Glimmer itself loads fine on stock llama.cpp via ordinary quants from other repos. Quant-equivalence hazard: no plain Q4_K_M exists anywhere sighted - not in this repo (only Q4_0_ROCMFP4_STRIX/_COHERENT and Q8_0_ROCMFPX), not in Meta's own GGUF repo (custom "kquant-17gb"/"kquant-dynamic" mixes, unnamed bpw), not in Unsloth's repo (Unsloth Dynamic UD-Q4_K_XL, a mixed-precision scheme, not plain Q4_K_M). Any of these could be screened on stock llama.cpp, but none would be a controlled quant-matched comparison against a field that is uniformly Q4_K_M - and the artifact actually staged here (ROCmFP4, 4.36 bpw dual-scale FP4) is a fourth, still-different method. Reddit source: the share link resolves to r/StrixHalo comments/1vknppz/ ("vmlinux/Muse-Glimmer-30B-ROCmFPX-GGUF | Hugging Face") but both the share link and the resolved permalink are login-walled for automated fetches, as anticipated. Web search for independent commentary, criticism, or additional tok/s reports on that specific thread turned up nothing beyond what the HF repo's own README/BUILD_RESULTS.md already state, so the Reddit layer added no information the source repo didn't already have. Strix Halo numbers, two different sources: vmlinux's own smoke test (gfx1151, ROCm, this exact ROCmFP4 quant) - 113.7 tok/s prompt / 14.9 tok/s decode alone, 28.3 tok/s decode with DFlash 6-token speculative window (2.07x, 35.8% draft acceptance). Separately, AMD's own blog for the OFFICIAL (non-ROCmFPX) release quotes up to 24 tok/s on Ryzen AI Max+ 395 and up to 53 tok/s on a Radeon AI PRO R9700 with DFlash - a different quant, different exact hardware, not directly comparable to vmlinux's number. Vision audit (2026-08-10): the parent DOES ship a real vision path - merged PR #26841 adds a text tower, perception encoder, and DFlash drafter together, and unsloth/Muse-Glimmer-30B-GGUF (the same repo the staged UD-Q4_K_XL came from) publishes matching mmproj files: mmproj-kquant.gguf (1,400,328,928 bytes), mmproj-Muse-Glimmer-30B-Q8_0.gguf (2,051,685,088 bytes), and a BF16 (3,849,173,728 bytes) - none staged. But the runtime claim two entries up needs a CORRECTION, not just an addition: PR #26841 merged as commit 62bf73d25c53b8161f8a22894d4f90c4aebbd7d0 on 2026-08-10 (today); our deployed build is pinned to commit 3653e6d (2026-08-07 22:35:52+0200), three days earlier, and 62bf73d is confirmed NOT an ancestor of our build via `git merge-base --is-ancestor`. Checked directly against the source tree on aihydra: "muse-glimmer"/"glimmer" appears nowhere in src/llama-arch.cpp on our checkout. So "the Muse-Glimmer architecture ITSELF is merged into stock llama.cpp... ordinary K-quant GGUFs... would run today" is upstream-true but NOT true of our deployed binary - the K-quant GGUF would currently fail to load at all, text or vision, until the build advances past 62bf73d. Blocked for a different, narrower reason than previously recorded (build staleness, not the ROCmFPX/quant-equivalence issues above) - update the rebuild step before the next screening attempt.
EmbeddingGemma (NPU)
Semantic triage for the continuous escalation loop. Embeddings let the loop compare an incoming item against memory before deciding whether to escalate, rather than judging on surface features alone - the difference between "this mentions a deadline" and "this contradicts something we already recorded".
gate history (3)
- 2026-08-10 · listed
Listed as the highest-value NPU workload we have: a real recurring production job that does not need the big model and currently steals slot time from one that does.
- 2026-08-10 · listed → listed
RATIONALE CORRECTED - the reason above is false. qmd's re-embedding runs on the MAC MINI and never competed for the inference box's GPU slot at all; the operator caught it. I asserted a contention I had not checked, which is the same failure as dropping a candidate on an unverified claim, just aimed at keeping one instead of cutting one. The candidate stays listed on the corrected rationale in why_listed above.
- 2026-08-14 · listed → screened
FIRST NPU-LANE SCREEN. surface=npu, backend=fastflowlm, FLM v1.0.1, model tag pulled `embed-gemma:300m` (FastFlowLM/Embedding-Gemma-300M-NPU2, catalog quantization_level "none", default_context_length 2048, footprint 0.62 GB). NPU-lane result; nothing here transfers to or from a GPU row. FIT ok, but ONLY AS A SIDECAR, and the discovery is worth more than the timing. `flm serve embed-gemma:300m` does not serve this model: it prints [ERROR] Unsupported model family or non-llm: "embed-gemma" and then SILENTLY LOADS Llama-3.2-1B-NPU2 instead - a model nobody asked for and which is not even in the installed list. The server comes up healthy, /v1/models answers, and /v1/embeddings returns a bare `null`. Nothing in the HTTP surface says the wrong model is resident. The correct form is a sidecar on a chat model: `flm serve <chat-model> --embed 1`, which logs "Embedding mode enabled: reserving additional 300MB". Screened that way with qwen3:0.6b as host - server ready 2.65s, power_state D3hot -> D0, "NPU Locked!" x6, one lock per embedding call. GUARD N/A, with reason, not FAIL. All four house probes (coherence, native tool call, needle at depth, cross-request isolation) presuppose a chat-completions generative model; there is no meaningful adaptation of them to an embedding endpoint. SMOKE N/A, with reason: tau2 is an agentic tool-calling suite and this model exposes no chat endpoint to run it against. FUNCTIONAL CHECK instead, gated on: embeddings produced, stable dimension, sane similarity ORDERING (ordering rather than an absolute score, because no calibrated threshold exists for this model on this backend). PASS. Dimension 768, identical across all inputs; L2 norm 0.996975, i.e. returned unit-normalised; mean request latency 0.177s over 5 embeds. Ordering held with room to spare: cos(anchor, paraphrase) 0.8000 > cos(anchor, same-topic) 0.7552 > cos(anchor, unrelated) 0.7005. Two properties measured and REPORTED rather than gated. (1) Not bit-deterministic: the same input embedded twice returns cos 0.99966432, not an identical vector - close enough for retrieval, but dedup-by-vector-equality in a vector store will not work against this backend. (2) The `model` field in the request body is IGNORED: a deliberately bogus id, "not-a-real-model:999b", returns cos 0.99966432 against the anchor - exactly the self-similarity baseline, so it is being served by the same loaded model. Combined with /v1/models returning FLM's entire catalog rather than what is resident, the endpoint cannot name its own subject, and any future record must take the served model from the serve invocation, never from the request or the response. Run-meta (suite npu-screening, host aihydra, contention:false, GPU idle throughout): fit 2026-08-14T06:05:54Z->06:05:59Z (5s). An earlier window, 06:02:37Z->06:03:23Z (46s, rc=1), is the failed direct-serve attempt described above and is kept deliberately as the evidence for the silent-substitution finding.
- note
NPU runs are QUEUED BEHIND GPU work, never concurrent. The NPU shares the same LPDDR5X pool and memory bandwidth as the GPU on Strix Halo, so a concurrent NPU run would both contend for bandwidth and perturb any GPU number measured alongside it - which would quietly corrupt the GPU series rather than merely slowing it. ⚠ amd_iommu=off is worth +5-12% on the GPU and DISABLES the NPU entirely. Those are mutually exclusive strategies, so the second lane has a standing cost that has to be counted against whatever it buys.
LFM2 (Liquid, NPU)
The only family in FLM's catalog that also sits in our GPU field, which makes it the one candidate that could answer whether NPU-versus-GPU is a fair trade at all - same family, both lanes, measured separately.
gate history (2)
- 2026-08-10 · listed
Listed for cross-lane comparability of the FAMILY, explicitly not for metric transfer. The GPU entry is LFM2-24B-A2B at Q4_K_M; the NPU artifact will be a different size at FLM's own quantisation, so the two produce independent numbers that may be compared but never merged.
- 2026-08-14 · listed → screened
FIRST NPU-LANE SCREEN. surface=npu, backend=fastflowlm, FLM v1.0.1, model tag pulled `lfm2:2.6b` (FastFlowLM/LFM2-2.6B-NPU2, Q4_0, catalog default_context_length 32768, footprint 1.8 GB). NPU-lane numbers, NOT comparable to the GPU row for LFM2-24B-A2B. ARTIFACT CHOICE, stated because it limits the cross-lane comparison this candidate was listed for: FLM's catalog carries no MoE LFM2 at all (lfm2:1.2b, lfm2:2.6b, lfm2-trans:2.6b, lfm2.5-it:1.2b, lfm2.5-tk:1.2b - all small and non-MoE). The largest available was taken. Its config.json reads model_type lfm2 with conv_L_cache/conv_dim alongside GQA (32 heads / 8 KV heads, 30 layers) and no MoE fields, so arch is corrected from moe to hybrid: this is the dense hybrid conv+attention line, not the 24B-A2B. The comparison the listing wanted is therefore FAMILY-level only, and a size difference sits inside it. FIT ok, WITH device evidence, and the fastest load of the four: server ready 1.68s, first (cold) inference 1.35s. power_state D3hot -> D0, "NPU Locked!" in the serve log. GUARD 3/4 at stock budgets. coherence PASS; needle at depth 8000 PASS at the STOCK 128-token budget (tok=5367, 'chartreuse-viper-88') - a non-thinking model needs no budget adaptation; isolation PASS (skipped at --parallel 1); toolcall FAIL, verbatim: "no tool call - check --jinja, the template, and the token budget". The model answered in prose that it "does not have access to external tools". ROOT CAUSE, measured rather than assumed: FastFlowLM SILENTLY DROPS the tools array for this manifest. Controlled probe, same request with and without tools - lfm2:2.6b prompt_tokens 19 vs 19 (delta 0, the schema is never rendered into the prompt); qwen3-tk:4b on the same server 19 vs 159 (delta +140). HTTP 200 and no warning in either the response or the server log. The FLM catalog labels qwen3-tk:4b "tool-calling" and labels lfm2 nothing, which turns out to be load- bearing. SMOKE 5-task tau2 airline (same pin: haiku-4.5 simulator, seed 42, max-steps 200, max-concurrency 1) returned completed=5/5, mean=1.00 in 372s - and that result is INVALID, not a pass. Reading the transcripts: ZERO assistant messages carry a tool call in any of the five. Every reward decomposes to DB 1.0 (the database is untouched, which matches the reference end-state for these tasks) plus COMMUNICATE 1.0 (vacuous - "No communicate_info to evaluate"). A model that cannot act scores 1.00 by never acting, above the 0.60 of the NPU model that actually calls tools and above every GPU model screened on this box. Recorded as the R1 rule requires: scores propose, transcripts decide. The honest screening verdict is the failure mode, stated as the tier asks - no tool-call support on this backend. Throughput, FLM self-reported, NPU lane only: prefill 570.2 t/s at a 682-token prompt, decode 29.5 t/s, TTFT 0.74-1.20s - about 1.4x the 4B thinking model's prefill and 1.6x its decode, as expected for less than two-thirds the parameters with no reasoning tokens. Run-meta (suite npu-screening, host aihydra, contention:false, GPU idle throughout): fit 2026-08-14T05:54:25Z->05:54:29Z (4s); guard 05:54:29Z->05:54:44Z (15s); smoke 05:54:44Z->06:00:56Z (372s); throughput 06:07:04Z->06:07:10Z (6s).
- note
NPU runs are QUEUED BEHIND GPU work, never concurrent. The NPU shares the same LPDDR5X pool and memory bandwidth as the GPU on Strix Halo, so a concurrent NPU run would both contend for bandwidth and perturb any GPU number measured alongside it - which would quietly corrupt the GPU series rather than merely slowing it. ⚠ amd_iommu=off is worth +5-12% on the GPU and DISABLES the NPU entirely. Those are mutually exclusive strategies, so the second lane has a standing cost that has to be counted against whatever it buys. ⚠ SMOKE INVALID (protocol gate added 2026-08-14 as a direct result of this screen): the 1.00 mean is VACUOUS - zero tool calls in all five transcripts. FastFlowLM silently drops the tools array for this manifest (measured: prompt_tokens 19 with and without tools, a 0-token delta, against +140 for qwen3-tk on the same server; HTTP 200, no warning). Every reward came from an untouched database plus a COMMUNICATE point, i.e. the score rewards never acting. It must NOT be compared with any other model's SMOKE, on any surface. The model's tool capability on this backend is UNMEASURED, not perfect.
Qwen3-4B-Thinking-2507 (NPU)
The triage brain for the continuous escalation loop. FLM's validated headline model, and thinking-capable at 4B, which makes it the actual escalate-or-hold decision maker rather than merely a proof the lane runs. Bring-up and first real candidate are the same model here, which is convenient but incidental.
gate history (3)
- 2026-08-10 · listed
Listed as the NPU lane's bring-up model: if this does not run, nothing else on the lane matters. Deliberately chosen for low risk rather than interest.
- 2026-08-10 · listed → listed
REFRAMED against the actual design intent. The original list was bring-up shaped - does the lane work - with use-case reasons retrofitted. The standing intent is a CONTINUOUS TRIAGE LOOP watching inputs and escalating to the big model, and this is the one candidate that is genuinely the loop's decision maker rather than a component near it. Its evaluation is therefore escalation precision/recall over sustained operation, NOT tau2 task success.
evidence: clm-0034
- 2026-08-14 · listed → screened
FIRST NPU-LANE SCREEN. surface=npu, backend=fastflowlm, FLM v1.0.1, model tag pulled `qwen3-tk:4b` (FastFlowLM/Qwen3-4B-Thinking-2507-NPU2, Q4_1, catalog default_context_length 32768, max_prefill_len 4096, footprint 3.1 GB). Every number below is NPU-lane and is NOT comparable to any GPU row - different artifact, different quantisation, different runtime. FIT ok, WITH device evidence. Server ready in 2.71s; first (cold) inference 4.39s - two numbers, not one, because FLM binds the port before it touches the NPU and loads lazily at the first request. Evidence: power_state on /sys/class/accel/accel0/device/power_state moved D3hot (runtime_status suspended) -> D0 during that first request, "NPU Locked!" appeared in the flm serve log, and /api/npu/status reported npu_available:false with active_requests:1 while generating. GUARD 3/4 at the STOCK house budgets. coherence PASS ('the quick brown fox'); toolcall PASS - native tool call with real arguments, get_booking {"reference":"ABC123"}, 0/1 empty; isolation PASS (skipped by design at --parallel 1); needle at depth 8000 FAIL. The needle failure is a HARNESS limit, not retrieval loss: this checkpoint has think_toggleable=false, so reasoning is always on, and the house probe's 128-token needle budget is spent before a single answer token is emitted. A supplementary re-run at 16x budget retrieved 'chartreuse-viper-88' cleanly at tok=5363, so retrieval at 8000 tokens is intact. The stock 3/4 stays the counted verdict; the supplementary pass is recorded but not counted. SMOKE 5-task tau2 airline, pinned simulator openrouter/anthropic/claude-haiku-4.5, seed 42, 1 trial, max-steps 200, max-concurrency 1: completed=5/5, mean=0.60, per-task 1.0/0.0/1.0/0.0/1.0, zero cut at max_steps (all user_stop), 3214s wall (54 min). 23 assistant messages carried tool calls across the five transcripts - this model does the agentic work rather than talking around it. Smoke is a cliff-detector, not a ranking (noise ~0.40, clm-0036). Throughput, FLM self-reported, for the NPU lane only: prefill 408.2 t/s at a 680-token prompt, decode 18.6-19.3 t/s, TTFT 1.02-1.67s. Prefix caching is real on this backend - the serve log shows "Use cached prompt! Matched 37 out of 39 messages" with checkpoint restore mid-conversation. Run-meta (suite npu-screening, host aihydra, contention:false, GPU idle throughout - use_pct 0 and zero llama-server/llama-bench/llama-swap processes at both ends of every phase): fit 2026-08-14T04:58:29Z->04:58:37Z (8s); guard 04:58:37Z->04:59:53Z (76s); smoke 04:59:53Z->05:53:27Z (3214s); throughput 05:53:56Z->05:54:15Z (19s).
- note
NPU runs are QUEUED BEHIND GPU work, never concurrent. The NPU shares the same LPDDR5X pool and memory bandwidth as the GPU on Strix Halo, so a concurrent NPU run would both contend for bandwidth and perturb any GPU number measured alongside it - which would quietly corrupt the GPU series rather than merely slowing it. ⚠ amd_iommu=off is worth +5-12% on the GPU and DISABLES the NPU entirely. Those are mutually exclusive strategies, so the second lane has a standing cost that has to be counted against whatever it buys.
Whisper (NPU)
STT already runs on wardenmac via whisper.cpp and is confirmed working end to end. Moving it to the NPU is the cheapest possible test of whether the lane can take a real production workload off an already-loaded machine.
gate history (3)
- 2026-08-10 · listed
Listed because there is an existing working baseline to compare against, which is rare - most NPU numbers would have nothing to be judged relative to.
- 2026-08-10 · listed → listed
Role clarified for the escalation loop: Whisper is a genuine CONTINUOUS INPUT SOURCE, which is what the loop consumes, not a model the loop chooses between. It tests sustained residency and non-interference - can something stay resident on the NPU for hours without perturbing concurrent GPU work - which is a loop requirement no one-shot benchmark covers.
- 2026-08-14 · listed → screened
FIRST NPU-LANE SCREEN. surface=npu, backend=fastflowlm, FLM v1.0.1, model tag pulled `whisper-v3:turbo` (FastFlowLM/Whisper-V3-Turbo-NPU2, Q4_1, catalog default_context_length 448, footprint 0.62 GB). NPU-lane result; not comparable to the wardenmac whisper.cpp baseline either, which is a different runtime on different silicon - that comparison needs its own matched run. FIT ok, as a SIDECAR. Like the embedding model, this cannot be served on its own; it loads alongside a chat model with `flm serve <chat-model> --asr 1`, which logs "ASR mode enabled: reserving additional 1GB of memory". Screened with qwen3:0.6b as host - server ready 1.71s, power_state D3hot -> D0, "NPU Locked!" x2, one lock per transcription. GUARD N/A, with reason, not FAIL - the four house probes all presuppose a chat-completions generative model and a speech model has no needle to retrieve. SMOKE N/A, with reason - tau2 is an agentic tool-calling suite. FUNCTIONAL CHECK instead: does the ASR path run on the NPU and return plausible text, and how fast. PASS on both files available on the box. llama.cpp's tools/mtmd/test-2.mp3, 17.45s of audio, transcribed in 3.77s wall = 4.6x realtime, returning a clean paragraph of narration about the New York Times front page of 21 July 1969. tau2-bench's blog/audio/clean/task_59.wav, 225.80s of audio, transcribed in 30.86s = 7.3x realtime, returning a coherent customer-service utterance with two order numbers repeated consistently and a full postal address. Audio lengths are FLM's own, from the serve log. ACCURACY IS BLOCKED-WITH-REASON, and deliberately not estimated. No reference transcript for either file exists on this box and no TTS is installed to synthesise known text, so WER cannot be computed. What is screened here is that the path runs on the NPU, at what speed, and that the output is coherent English - not how accurate it is. A WER number needs a reference corpus staged first. Run-meta (suite npu-screening, host aihydra, contention:false, GPU idle throughout): fit 2026-08-14T06:06:11Z->06:06:49Z (38s).
- note
NPU runs are QUEUED BEHIND GPU work, never concurrent. The NPU shares the same LPDDR5X pool and memory bandwidth as the GPU on Strix Halo, so a concurrent NPU run would both contend for bandwidth and perturb any GPU number measured alongside it - which would quietly corrupt the GPU series rather than merely slowing it. ⚠ amd_iommu=off is worth +5-12% on the GPU and DISABLES the NPU entirely. Those are mutually exclusive strategies, so the second lane has a standing cost that has to be counted against whatever it buys.
Gemma-4-12B
Runs as a working ORCHESTRATOR in domdoss/Warden, which demonstrates that routing and brief-writing have a far lower capability bar than task execution. If a 12B can route, the reflex tier is much cheaper than we assumed.
gate history (3)
- 2026-08-09 · listed
Evidence from a comparable project that a 12B suffices for the orchestrator role.
evidence: clm-0034
- 2026-08-10 · listed → listed
I had excluded this from the field on the grounds that orchestrators should not be scored on tau2. That was my own role label, not a measurement, and the operator correctly rejected it. Re-listed. The right response to 'tau2 will not characterise routing' is to BUILD a routing benchmark, not to skip the model.
evidence: clm-0034
- 2026-08-13 · listed → screened
Downloaded overnight and sha256-verified (companion mmproj-F16.gguf fetched alongside; this screen is TEXT-ONLY per the house tier, projector deliberately not loaded). Screened on aihydra, build min-62bf73d (62bf73d25) — the newest protocol-registered build — under the Aug-12 protocol (pinned haiku-4.5 simulator). FIT ok, loaded in 8s at 32K. GUARD 3/4, NOT clean: coherence PASS, toolcall PASS, isolation PASS (skipped under --parallel 1), needle FAIL — NEEDLE LOST at depth 8000 (tok=5364): in place of the needle the model returned a run of '<unused49>' control-token spam. SMOKE 5-task tau2 airline: rc=124 — the 3600s wall cap expired with 0/5 tasks completed, task 1 still running at 3,600s; no results file, so mean=n/a, cut count n/a. The server log shows a single runaway generation at n_decoded=13,525 and climbing when killed — a degenerate-output profile consistent with the needle spam, exactly the gpt-oss-class defect GUARD exists to catch. Caveat a screen cannot resolve: whether this is the checkpoint, the Q4_K_M quant, or a chat-template/tokenizer gap for the 12B's "unified encoder-free" variant in current llama.cpp. The orchestrator-role question from listing is untouched either way — a model that spams control tokens at depth cannot route. Smoke is a cliff-detector, not a ranking (noise ~0.40, clm-0036). Run-meta: screening/gemma4-12b 2026-08-13T06:52:41Z->07:53:06Z (3625s, rc=124).
- note
Vision audit (2026-08-10): parent IS vision-capable - Gemma 4 ships multimodal across all five sizes, and the 12B specifically uses a distinct "unified encoder-free" architecture (raw image/audio patches projected directly, no separate CLIP tower) rather than the standard vision path used by the 26B-A4B/31B sizes. Our llama.cpp build (commit 3653e6d) already registers both paths - clip.cpp defines PROJECTOR_TYPE_GEMMA4UV/GEMMA4UA for this unified 12B variant alongside GEMMA4V/GEMMA4A for the rest of the line. Not staged, so no mmproj question applies yet; if acquired, confirm the specific HF repo publishes a companion mmproj (the 26B-A4B unsloth repo does - see gemma4-26b) before assuming vision comes along for free.
GLM-4.5-Air
Appeared in an earlier chat-tier group config alongside gpt-oss-120b and qwen-fast, so it was once considered deployable - but it was never benchmarked and never explained.
gate history (2)
- 2026-08-10 · listed
Surfaced by search rather than recall. Listed to be resolved: either it supersedes GLM-4.7-Flash for our purposes or it does not, and right now nothing on record says which.
- 2026-08-13 · listed → screened
Downloaded overnight and sha256-verified (2-part Q4_K_M GGUF, 68 GiB; server pointed at part 00001). Screened on aihydra, build min-62bf73d (62bf73d25), Aug-12 protocol (pinned haiku-4.5 simulator). FIT ok, loaded in 45s at 32K — the fit question is settled, it runs with headroom. GUARD 3/4, NOT clean: coherence PASS, toolcall PASS, isolation PASS (skipped under --parallel 1), needle FAIL — NEEDLE LOST at depth 8000 (tok=5363): the model returned a run of '?' characters in place of the needle. SMOKE 5-task tau2 airline: rc=124 — the 3600s wall cap expired with 0/5 completed, task 1 still running; no results file, mean=n/a, cut count n/a. Server log shows a runaway generation at n_decoded=9,500 and climbing when killed. IMPORTANT shared-profile caveat: this is the SAME failure signature gemma4-12b produced an hour earlier on the same build (needle lost to junk-token spam + smoke runaway) while qwen36-27b-mtp and nemotron35-lightning-30b screened CLEAN on this identical build/harness last night — so the harness demonstrably can pass, but two unrelated archs failing identically in one session leaves model-defect-vs-build-interaction (FA/rope/template on gemma4+glm4moe paths) genuinely undecided. Do not treat this screen as a capability verdict on the checkpoint without a cross-build repro; do not treat it as harness noise either — the runaway is real and measured. The supersedes-GLM-4.7-Flash question from listing stays open. Smoke is a cliff-detector, not a ranking (noise ~0.40, clm-0036). Run-meta: screening/glm-45-air 2026-08-13T07:53:25Z->08:54:32Z (3667s, rc=124).
- note
Vision audit (2026-08-10): parent is text-only. Zhipu ships vision separately as GLM-4.5V; GLM-4.5-Air itself is the agentic/tool-use text line. Not staged, so no artifact to check, but the parent-level answer alone settles it: text-only, not a vision-lane candidate.
Llama 4 Scout
The in-range Llama 4 at ~61 GB Q4 with a large context window - a fits-tier representative of the Llama lineage.
gate history (3)
- 2026-06-20 · listed
Marked optional at listing: ~half gpt-oss speed at 17B active, and not clearly better than the gpt-oss control in its size class.
- 2026-08-10 · listed → listed
Still listed. 'Optional' was a priority call, not a rejection, and it stays in the field until something measured removes it.
- 2026-08-13 · listed → screened
Downloaded overnight and sha256-verified (2-part Q4_K_M GGUF, ~61 GiB; server pointed at part 00001; the vision-note mmproj question remains untouched — text-only screen). Screened on aihydra, build min-62bf73d (62bf73d25), Aug-12 protocol (pinned haiku-4.5 simulator). FIT ok, loaded in 40s at 32K. GUARD 3/4: coherence PASS, needle PASS at depth 8000 (tok=5361 'chartreuse-viper-88'), isolation PASS (skipped under --parallel 1), toolcall FAIL — probe raised HTTPError 500. SMOKE 5-task tau2 airline: rc=0 in 22 min but completed=0/5, mean=n/a, cut_at_max_steps=0 — every task died on the same server-side rejection. Verbatim error, repeated throughout the server log: {"error":{"code":500,"message":"The model produced output that does not match the expected peg-native format","type":"server_error"}} — and the common_chat_peg_parse warning shows WHAT was rejected: 'unparsed peg-native output: [sophia_silva_7557, get_user_details(user_id="sophia_silva_7557")]', i.e. well-formed Llama-4 pythonic tool-call syntax the build's peg-native grammar will not accept. Distinct profile from the gemma4-12b/glm-45-air runaway-generation failures the same night: here decode is sane and needle retrieval works — the failure is confined to tool-call format handling, which reads as a chat-template/parser dialect mismatch in the build (fixable harness-side) rather than model incapability. Do not score the model on this smoke; re-screen once a registered build parses Llama-4 pythonic calls. Smoke is a cliff-detector, not a ranking (noise ~0.40, clm-0036). Run-meta: screening/llama4-scout 2026-08-13T08:54:34Z->09:17:41Z (1387s, rc=0).
- note
Vision audit (2026-08-10): parent IS vision-capable, and natively so - Llama 4 uses early fusion (image patches tokenized alongside text from pretraining, no bolt-on encoder), and Scout specifically is documented for visual recognition, image reasoning, and captioning. Our llama.cpp build (commit 3653e6d) already carries real runtime support: LLM_ARCH_LLAMA4 is registered and clip.cpp implements PROJECTOR_TYPE_LLAMA4 end-to-end. Not staged yet (still the ~61GB Q4 download from listing), so the mmproj question doesn't apply until then - but unlike the Qwen/Gemma/Muse-Glimmer cases, this one is runtime-verified-ready pending acquisition rather than a hazard to route around. When it comes off the list: confirm whichever HF repo we pull from actually publishes a companion mmproj before assuming it ships bundled with the quant.
DeepGrove Maple-Preview
A ternary 20.2B at just 5.31 GB, MIT-licensed. That footprint would make it viable as a Mac-mini second lane where nothing else in the field fits at all.
gate history (3)
- 2026-07-30 · listed
Ternary quantisation at a footprint that opens a second lane on hardware we already own.
evidence: clm-0018
- 2026-08-10 · listed → listed
I had excluded this because its own model card concedes underperformance on agentic benchmarks. That is the vendor's claim about their own model, which is exactly the kind of thing this harness exists to test rather than accept. Exclusion withdrawn. Its undocumented 'on-device weight adaptation (dreaming)' remains unverified at source and is a separate question from whether the model is any good.
evidence: clm-0018
- 2026-08-13 · listed → screened
Downloaded overnight and sha256-verified (5.9 GB single artifact; quant field corrected to the shipped TQ2_0-head-Q4_K — DeepGrove publishes no Q4_K_M, the ternary singleton is the only artifact). Screening attempted on aihydra, build min-62bf73d (62bf73d25), Aug-12 protocol; the off-list quant required a recorded PROTOCOL_OVERRIDE of the quant gate (screening-only, entering no same-quant comparison — override text lands on the run-meta line). FIT FAILED after 4s and the verbatim error IS the screening result (laguna precedent): "llama_model_load: error loading model: unknown model architecture: 'maple'". Secondary probe on 77a9a66eb, the only by-date-newer registered build, fails with the identical line — the arch is unsupported on every registered general-purpose binary, not just the screening default. No GUARD, no SMOKE. The Mac-mini second-lane question from listing stays open, but it is now measurably blocked on upstream llama.cpp 'maple' arch support rather than on anything we control. Run-meta: screening/maple-preview 2026-08-13T07:53:08Z->07:53:16Z (8s, rc=1).
- note
Vision audit (2026-08-10): parent is text-only. DeepGrove's own materials describe Maple-Preview as a natively-trained ternary reasoning model with no vision component; the model card's acknowledged agentic weaknesses (already noted above) are about tool use, not modality. Verdict: text-only, not a vision-lane candidate.
Nemotron-Cascade 2
Gold-medal competitive-programming reasoning at a tiny active-parameter count - the reasoning-per-active-parameter outlier of the field.
gate history (3)
- 2026-06-20 · listed
Listed as the reasoning specialist spot-check - gold-medal competitive-programming performance at 3B active is an outlier worth confirming or refuting.
- 2026-08-10 · listed → acquired
Q4_K_M downloading from bartowski/nvidia_Nemotron-Cascade-2-30B-A3B-GGUF.
- 2026-08-10 · acquired → screened
Screened: FIT ok (16s), all four guards pass including needle at 8000. Smoke 0.20 with 3 of 5 tasks CUT at the 200-step limit — a non-termination tendency on agentic multi-turn work, consistent with a reasoning specialist outside its domain. Needle pass + agentic flail is the opposite profile to LFM2/Gemma.
- note
Vision audit (2026-08-10): parent Nemotron-Cascade-2-30B-A3B is text-only by NVIDIA's own model card - no image input, no vision variant anywhere in the line. Staged artifact (nvidia_Nemotron-Cascade-2-30B-A3B-Q4_K_M.gguf, arch nemotron_h_moe) carries no vision tensors and none exist to add. Verdict: text-only, not a vision-lane candidate.
benched (13)
Qwen3.8-27B
Successor to the Qwen3.6-27B-MTP line, a dense architecture — the case where MTP's larger dense speed-up (1.4-2x, against 1.15-1.25x for MoE) applies. If the 27B line holds its claimed coding strength - the 3.6 generation claimed 77.2 SWE-bench at 27B - then a same-size successor is the highest-leverage candidate we have: reflex-tier footprint with executor-tier capability.
gate history (19)
- 2026-08-10 · listed
Listed on the operator's expectation of an imminent release. Provenance is honest about itself: this is an expectation, not a vendor benchmark and not a measurement, and it is a fine reason to LIST - listing costs nothing and commits us to nothing.
- 2026-08-10 · listed → listed
Confirmed the line EXISTS but is not yet runnable for us. HF carries huginnfork/Qwen3.8-27B-FP8, huginnfork/Qwen3.8-27B-NVFP4A16 and neroued/Qwen3.8-27B-nvfp4-NInfer, so the 27B weights are out - but the only Qwen3.8 GGUFs are 4B and 1.2B distills. No 27B GGUF means nothing we can run on llama.cpp yet. A daily two-tier watch is now live on aihydra (ANY = line exists, early warning; GGUF = actionable) writing to ~/.model-watch.NEW, so the first 27B GGUF surfaces without anyone having to remember to check.
- 2026-08-14 · listed → listed
RELEASED and CONFIRMED runnable. Official Qwen/Qwen3.8-27B is live, apache-2.0 (the 2.4T flagship's move to a custom licence did NOT carry to the 27B), and its config declares architectures ["Qwen3_5ForConditionalGeneration"], model_type qwen3_5 - already registered in mainline llama.cpp including the vision path, so NO new architecture support is needed. This is not the Ling situation. Predicted from a field-by-field match against Qwen3.5-27B's config before release and confirmed exactly on the day. vision_config present; tie_word_embeddings false. unsloth shipped 25 GGUF quants plus mmproj-BF16/mmproj-F16 on day zero (Dynamic V3.0 preview quants, day-zero access from Qwen). Acquired to the ladon library under the new library-first rule: Q8_0 (primary screen, removes the quant confound), UD-Q4_K_XL (comparability arm, matches the quant class the Qwen3.6-35B incumbent is benched at), and mmproj-F16.
- 2026-08-15 · listed → screened
Screen tier passed on release day, on-box. FIT: both quants load on both backends at every probed context up to the full declared 262,144 (12/12 load probes clean, phase A) with headroom to spare — clm-0056. SMOKE: the 5-task tau2 airline subset under the pinned simulator scored 1.000 (5/5, 21 tool-call messages, 0 empty assistant turns — VALID under the smoke gate), meeting the registered >=0.80 prediction and beating the incumbent qwen36-27b-mtp's 0.80 on the identical tasks (clm-0057). Degeneracy guard in place of the house 4-item capability guard: 25/25 clean full-length generations on the serving config at n_max=3 (run-0267) — necessary because the screen ALSO found spec-draft-n-max>=4 corrupts generations outright (clm-0055), so the guard is load-bearing. Screen and bench ran as one overnight programme; the bench entry below carries the full results.
- 2026-08-15 · screened → benched
DAY-ONE SCREEN COMPLETE, all four phases, overnight on release day. MEMORY (phase A): KV measured byte-exact at 64.00 KiB/token — the 16-of-64 hybrid-attention allocation working as designed; the community 8.1 GiB-at-32k figure was 64-layer arithmetic, actual is 2.00 GiB; full-context (262,144) footprints Q8_0 42.2 GiB / UD-Q4_K_XL 32.6 GiB — fit is a non-issue (clm-0056). BACKEND x DEPTH (phase B, both quants, median-of-3): decode is backend-independent (within 5% everywhere) but Vulkan prefill collapses with depth (0.55x/0.48x ROCm at d32768) and Vulkan LOSES THE DEVICE at d131072 in 2/2 reps on both quants (vk::DeviceLostError, amdgpu ring reset) while ROCm completes every cell — the d0 backend pick reversed at depth; ROCm is the serving backend for this model (clm-0054). MTP (phases C): draft-mtp peaks at n_max=3 over clean cells (2.23-2.33x floor, acceptance ~0.62 on both backends) and is BROKEN at n_max>=4 — the EOS-cliff discovery: after ~1.4k generated tokens of varied-prompt session volume, generations terminate at 1 token with im_end, the target's own pre-sampling distribution driven to ~92% on EOG; ignore_eos restores text and draft accounting; 1-token generations report as 1,000,000 t/s so unfiltered sweeps flatter the broken cells; no upstream issue matches the signature; evidence bundle staged for filing (clm-0055). TAU2 (phase D, ROCm Q8_0 n_max=3, reasoning_effort=medium pinned+proven, max_tokens 4096, pinned haiku simulator, seed 42): SMOKE 1.000 (5/5, 21 tool calls, VALID) — the registered >=0.80 prediction MET, beating the incumbent qwen36-27b-mtp's 0.80 on the identical tasks; FULL 0.577 over effective n=26 of 50 (wall-bound cut at rc=124, all scored, 165 tool calls, 0 empty) with the capped-run caveat (clm-0057). Records: cfg-0062..cfg-0069, run-0245..run-0271.
evidence: clm-0054 clm-0055 clm-0056 clm-0057 run-0271 run-0267 run-0268
- 2026-08-16 · benched → benched
Read-only audit of a community fork/artifact promoted for THIS model (r/StrixHalo 1vpiwz0, github.com/julianmb/q38rocm) -- no execution of their code, gate unchanged, added as evidence rather than a new measurement. The fork (charlie12345/ROCmFPX, based on official llama.cpp b9438) genuinely implements a --spec-mtp-strict-qwen exact- verification mode for qwen35/qwen35moe draft-mtp, but the repo's own launcher (run_server.sh) does not enable it by default and runs spec-draft-n-max=6 (README's own "Deep Spec" row goes to 7) -- both past the n_max>=3 hard ceiling this candidate's own screen established (clm-0055's EOS-cliff), on a config the fork's own source code admits may diverge from no-spec decoding when strict mode is off. Their FP4 GGUF artifact (julianmb/Qwen-3.8-27B-ROCmFP4-FAST-GGUF) was downloaded, sha256-verified, and empirically fails to load on stock 3653e6d: verbatim "tensor 'output.weight' has invalid ggml type 101. should be in [0, 43)" -- confirms fork-only, matching the same charlie12345/ROCmFPX dependency already on record for nemotron35-lightning-30b and muse-glimmer-30b. Per task protocol ("if the artifact only works on their fork, record that as the screening result and stop"), no GUARD/SMOKE was attempted. Full writeup: clm-0069.
evidence: clm-0069
- 2026-08-17 · benched → benched
Depth-ladder backfill (protocol extension, same date): ROCm Q8_0 extended past the original matrix's d131072 ceiling to d204800 (61.00 pp / 5.449 tg t/s) and d262144 -- this candidate's real declared context maximum (qwen35.context_length in the GGUF header, confirmed 262144, not a YaRN projection) -- at 51.23 pp / 5.025 tg t/s. Both cells are a single fresh-process rep (N=1, house convention for cells past d131072: the question at this depth is survival and coarse throughput, not variance). New config cfg-0103 minted for this session (identical fingerprint to cfg-0062: build 3653e6d, ROCm, Q8_0, f16 KV, -fa 1, --load-mode none -- only depth differs, not a fingerprint field), guard_waived with the same text cfg-0062's own cells carry (no capability guard exists for a bare llama-bench throughput sweep on this backend/build/KV/load-mode combination). Monotonic with the existing series across all five depths, 0 through 262144 (pp: 226.2/159.98/83.44/61.00/51.23 t/s; tg: 7.814/7.287/6.109/5.449/5.025 t/s) -- no cliff, no scatter anomaly, ROCm's 131072-depth lead now measured rather than extrapolated. Vulkan NOT re-attempted at either new depth: it was already device-lost on both allowed attempts at d131072 (clm-0054), below this backfill round's bar of "previously survived its prior max tested depth" for adding a new Vulkan cell. Wall time: d204800 ~37 min, d262144 ~98 min (each pp/tg sub-test independently repays the depth-fill cost, so the deepest cell in this round's whole priority list took the most wall-clock of any single cell run this session). Energy: not joined before session release (recorded explicitly per run, not left silent) -- HA history retention ~10 days, join before 2026-08-27.
- 2026-08-19 · benched → benched
Community-only DFlash2 triage from live llama.cpp PR #27342. The useful signal is not a new measured result and does not change this candidate's measured serving pick: DFlash2 is a separate speculation path whose current public evidence is hardware- and workload-dependent. One Strix Halo Vulkan report has DFlash2 Q8_0 roughly tied with MTP on prose and ahead on code, but adjacent comments report a severe Intel B70 multi-agent collapse, a V100 multimodal/M-RoPE draft-context failure, n_max=7 regression on V100, RTX 3090 near-parity to modest gain, and Blackwell scaling better at higher parallelism. Future aihydra work, if any, should be a matched DFlash2-vs-MTP-vs-plain lab check with the house guard and varied prompts, not inherited from these comments.
evidence: clm-0090
- 2026-08-20 · benched → benched
Upstream #25618 prompt-sensitivity addendum added as community report evidence, not a new HaloBench run. The independent Strix Halo Vulkan reproduction confirms the exact F16-V PASS on the original prose prompt, but the same target/draft artifacts, runtime/server/template/completion path and greedy reporter sampler produced 0/5 exact parity on five other prompts; a separate fixed-sampler replay of those prompts also stayed 0/5 with unchanged first mismatch locations. Local reading: one exact trajectory is real evidence for that trajectory only; future MTP invariance/losslessness dossiers need varied prompts before claiming a general baseline/MTP invariance property.
- 2026-08-20 · benched → benched
HO-004 matched DFlash2-vs-stock-MTP-vs-plain lab check reached only a negative Stage A diagnostic. On the reviewed v0.6.5 portable Vulkan runtime, stock MTP n_max=3 and DFlash2 Q8_0 width 4 both showed log-backed draft activation, but plain, stock-MTP, and DFlash2 arms all failed the varied-prompt sentinel gate on code/toolish prompts (6/15 bad each). No guard, Stage B performance, throughput, production recommendation, quality equivalence, community-number inheritance, or stock MTP n_max>=4 safety claim is admitted. Stage A wall-meter windows were joined as Wh total only for the rejected diagnostic arms.
evidence: clm-0098 run-0541 run-0542 run-0543 eng-0222 eng-0223 eng-0224
- 2026-08-22 · benched → benched
Upstream #25618 follow-up (RDNA4 R9700/gfx1201) added as community report evidence, not a new HaloBench run. The 2026-08-22 snick525 comment (con-0006/clm-0104) isolates the spec-path divergence as DRAFTER- INDEPENDENT on a Qwen3.8-27B Q6_K_XL target: the built-in MTP head and an external DFlash2 drafter (PR #27342) both diverge from --spec-type none at the same first-diff byte 204 (code prompt) at Q4/Q8/BF16 drafter quant, lossless on a reasoning prompt, and DFlash2 Q4/Q8/BF16 are byte-identical with identical acceptance. It also shows KV-cache quantization alone moves greedy output (reasoning prompt lossless under f16 KV diverges at byte 155 under q8_0 KV; --spec-type none differs at byte 561), so q8_0 KV is never output-preserving. For this candidate (a #25618 target) the note is that the drift is not a drafter fixable quantity and is compounded by q8_0 KV; recorded for MTP-correctness provenance only, and on RDNA4 not our ROCm/Vulkan surface, so it does not change any measured serving result.
- 2026-08-23 · benched → benched
HO-004-DIAG v0.6.10 null-update: the v0.5.x Stage-A sentinel empty-content defect (clm-0098, 6/15 on code/toolish) is RESOLVED on v0.6.10 (2586f6edd, Vulkan/RADV). The plain/control arm now passes Stage-A sentinel (0 empty-assistant/tool-call/single-token-EOS, code-valid 3/3, tool-correct 3/3) plus the 4/4 house guard, so the plain-vs-DFlash2-vs- stock-MTP comparison is legal again. Both spec arms (DFlash2 Q8_0 n_max=3, stock-MTP n_max=3) also pass sentinel+guard with log-backed draft activation, but the G2 varied-prompt invariance check is NEGATIVE for both — neither is byte-invariant to plain control (2/5 prompt classes) — so neither spec arm is promoted. Capability/throughput unchanged; this adds an empty-content-defect-resolution check only, no DFlash2-vs-MTP or production claim.
- 2026-08-23 · benched → benched
Upstream MTP-correctness corroboration added as community evidence, not a new HaloBench run. PR #27210 (clm-0115/con-0011) shows a high fixed n_max CEILING carries a fixed ~2.6% cost independent of adaptive logic, and that adaptive depth-changing does not slow CUDA-graph decode (RTX 5090) — the standing shallow/adaptive pair, not a deep static or deep blind-sweep n_max, remains the target. This is external confirmation of this record's own n_max<=3 hard ceiling rationale (clm-0055) without changing it. Issue #27122 (clm-0114/con-0010) independently confirms the spec-path fragility on CUDA multi-GPU --split-mode tensor (6/6 hard crashes on a 4x-RTX-A4000 build while the 2-GPU AllReduce path is stable; LLAMA_GRAPH_REUSE_DISABLE=1 workaround). Both are upstream context only on platforms other than our single-GPU ROCm/Vulkan surface, so no serving result or n_max guidance changes on this record.
- 2026-08-23 · benched → benched
HK-RERUN-REASONOFF delta row (cfg-0172, clm-0118): reasoning-OFF (-rea off) rerun on the SAME Kairic Edge TheRock build eliminates the empty-assistant defect — tau2 full-26 now 26/26 completed, reward 0.423 (11/26 passed, 0 empty turns), vs the reasoning-ON run that aborted at 60% (disposition-only, clm-0116). Matched served-path spec sweep: --kairic-edge on/off ratio ~1.0 (tg 13.1-13.2 t/s flat at every rung 0→204800, both arms 10/10) — the speculative flag buys NO decode throughput here; it drew ~23% less wall energy for identical tokens (17.1 vs 22.2 Wh, one observation). Previously-crashed d32768 stable in both spec states. Blocked claims unchanged: no reasoning-ON reward or throughput comparison. Records: cfg-0172, run-0618..run-0620, eng-0267..eng-0269, clm-0118.
- 2026-08-23 · benched → benched
HK-001-REASONING-MATRIX delta row (cfg-0173, clm-0119): 5-cell tau2 smoke matrix isolating the reasoning lever on the same Kairic Edge build. Reasoning-ON cells (R1-R4, unlimited or capped budgets) all score 0.8 (4/5) vs reasoning-OFF R0 at 0.6 (3/5) — same tasks 0-4, seed 42, haiku judge; budget capping at 1024/2048 preserves quality. 5/5 cells rc=0, zero infra errors / empty turns; guard PASS on R0, waived R1-R4 (binary invariant). Cost: reasoning-ON ~2x wall clock and energy (~84-92 Wh vs 43.2 Wh per smoke cell). Reviewer pick for the full-26 confirmation run: R2-B8k (max_tokens=8192, reasoning=on, budget=-1). Smoke-scale only — no full-26 claim made. Records: cfg-0173, run-0621..run-0625, eng-0270..eng-0274, clm-0119.
- 2026-08-24 · benched → benched
COMPARABILITY ANNOTATION + v1.2 RUNTIME GATE (HK-V12-ANNOTATE, per hbreviewer verdict t_b05473c9; no new measurement). The vendor released Kairic Edge runtime tag kairic-edge-qwen38-27b-v1.2 @ 205a3e5f: v1.1's native IU4 M65 verifier could flip a greedy target token on low-margin reproduced cases; v1.2 defaults to strict-compact verify (unsafe native only via KAIRIC_UNSAFE_NATIVE_M65_VERIFY=1). GGUF + .pfs sidecars are byte-unchanged, so all recorded THROUGHPUT figures stand; recorded tau2 OUTPUT CORRECTNESS on clm-0116 / clm-0118 / clm-0119 carries a comparability caveat (annotations added to each claim). WATCHLIST GATE: all future Kairic Edge box runs MUST use runtime tag kairic-edge-qwen38-27b-v1.2 @ 205a3e5f — a run specifying any other tag fails preflight and must not dispatch. Capability-comparability re-runs (full-26 confirmation at R2-B8k per t_8869ab61) are justified on v1.2 but stay gated until the overnight aihydra reservation is lifted. Records: con-0012 (v1.2 mandate), clm-0116/clm-0118/clm-0119 (annotations).
- 2026-08-24 · benched → benched
AUTONOMOUS TUNING HANDOFF AUDIT — bounded screening evidence ingested without changing the measured production selection. The retained served-path samples, five-task smoke, and cache-reuse trace lack artifact identities, a standard guard, paired floor/served-path throughput cells, and energy joins. The multi-slot report lacks raw stdout, and the Kairic comparison lacks the paired raw capability, footprint, runtime, and dual-throughput evidence. These are explicit evidence gaps, not negative model results or a production promotion.
- 2026-09-06 · benched → benched
WORKER-CONFIG CAMPAIGN (2026-09, aihydra) — a full HaloBench on a DIFFERENT serving profile from the 2026-08 screen's cfg-0066, aimed at the dense-worker role (capability + throughput beside a faster interactive model), not a re-measurement of the Q8_0 pick. Config cfg-0183: UD-Q4_K_XL, the pr27311 leak-fix build (c530ea79c), self-speculative draft-mtp n_max 2 (NO -md; draft context against the target's own nextn head), --parallel 4, reasoning_effort low, greedy, via the slotpin proxy. CAPABILITY: tau2 airline 0-25 = 91.7% (22/24 scored, 0 empty turns over 106 min, 2 cloud-user-sim infra errors excluded) — NOT a clean A/B against run-0271's 0.577 (four fingerprint fields differ at once; leading hypothesis is the max_tokens cap), clm-0129. THROUGHPUT: served single-stream decode 19.2/17.0/14.7/12.5 t/s at 32k/65k/131k/205k (prefill 239/141/74/50, MTP accept 0.76-0.84), and ~12.8 t/s/slot at mean concurrency 3.6 = ~46 t/s aggregate under real 4-slot load, clm-0130. BINARY BUILD FINDING: the same flags on fresh master c5a5535e went EMPTY within ~4 min of 4-slot load (#25992), while pr27311 held 0 EMPTY for 106 min (run-0660) — the earlier config pinned --parallel 1 to dodge exactly this, so pr27311 (con-0018) is what unlocks a multi-slot dense worker. ENERGY: 15.18 Wh/correct (eng-0278) vs the single-slot Q8_0 run's 38.29, driven by multi-slot amortisation + more-correct, clm-0131. Records: cfg-0183, cfg-0184, run-0658..run-0664, eng-0278, eng-0279, con-0018, clm-0129..clm-0131.
evidence: clm-0129 clm-0130 clm-0131 run-0658 run-0660 cfg-0183 con-0018
- 2026-09-06 · benched → benched
THIRD-PARTY STACK SCREEN (pwilkin/ilintar + halo-box), gate unchanged. Reproduced the pwilkin Strix Halo build from source at its exact pins (custom ROCr + retained-PM4 HIP from rocm-systems ilintar-experiments 78d1160, + pwilkin/llama.cpp strix-halo d3b5cc4) and screened it against the worker (cfg-0183) on the SAME model. It faithfully reproduced the author's headline (26.9 t/s vs claimed 26.256, 31.5k prompt, DFlash2 width 3). Full matrix: CAPABILITY tau2 0-25 = 83.3% (20/24, 0 empty over 66 min) vs worker 91.7% -> worker +8pts (clm-0132); THROUGHPUT single-stream 27.98/23.18/22.12/ 17.07 t/s @32k/64k/131k/205k = +37-50% vs worker, but multi-slot aggregate 47.3 t/s @4 only ties the worker's ~46 (poor 1.3x uneven scaling from DFlash2 draft contention), clm-0133; ENERGY 10.26 Wh/correct vs 15.18 (clm-0134). ISOLATION: ran the worker's own UD-Q4_K_XL + self-spec MTP on the retained-PM4 runtime — only +4%, lossless (clm-0135) — so the ~45% lead is the IQ4_XS quant + DFlash2 (the capability trade), NOT the runtime. Verdict: worker cfg-0183 stands as the capability-first pick; the speed/quality trade is real and not circumventable by the runtime. PROVENANCE: pwilkin's llama.cpp changes are upstream PRs (e.g. #27311/con-0018); retained-PM4 alone has no upstream PR and per the author is not expected to mainline (con-0019). HALO-BOX fork also screened: merged build ~= pwilkin on our models; its 65.6 headline is an unbacked README claim on an unmerged branch (no artifact). Records: cfg-0185..0187, run-0665..0673, eng-0280, con-0019, clm-0132..0135.
evidence: clm-0132 clm-0133 clm-0134 clm-0135 run-0665 run-0668 con-0019
- note
Watch it against the incumbent it would replace rather than in isolation: Qwen3.6-27B-MTP is already on disk with zero rows, so screening the two together costs barely more than screening either alone and answers the generation question directly. ⚠ Do not inherit the 3.6 generation's MTP assumption. MTP support is per-model and has to be confirmed in the GGUF rather than assumed from the line - Ling-3.0-flash ships its MTP head INACTIVE (clm-0015), which would silently distort any decode comparison, and Nemotron's MTP turned out incompatible with draft-mtp entirely because of Mamba. Two precedents is enough to make this a check, not a footnote. Vision audit (2026-08-10): Qwen positions the whole 3.8 generation, including this 27B, as multimodal-capable per its own announcement materials - consistent with the 3.5/3.6 pattern already confirmed vision-capable on this box (see the qwen35-122b, qwen36-27b-mtp and qwen36-35b notes). But that is a parent-line claim only: as recorded above, no 27B GGUF conversion exists anywhere yet, vision or otherwise, so there is no artifact to check an mmproj against and nothing for llama.cpp to load. When a 27B conversion surfaces on the existing watch, check for a companion mmproj in the same repo before assuming one exists. SCREENING PARAMETERS (vendor-published, 2026-08-14 - use these, do not use harness defaults). Hybrid thinking model with DIFFERENT sampling per mode: thinking temp 1.0, top_p 0.95, top_k 20, min_p 0.0, presence_penalty 0.0 non-thinking temp 0.7, top_p 0.80, top_k 20, min_p 0.0, presence_penalty 1.5 The non-thinking presence_penalty of 1.5 is unusually high and is the vendor's explicit anti-repetition recommendation. reasoning_effort defaults to xhigh and MUST be pinned to medium via chat_template_kwargs - llama.cpp still discards the OpenAI reasoning_effort field pending PR #26941, and we measured xhigh making a 5-task tau2 smoke unfinishable in 90 minutes on another model. UNSLOTH TEMPLATE PATCHES: the unsloth GGUFs carry vendor-modified chat templates - "Developer Role Support for agentic tools like Codex" and "improved parsing of nested objects to make tool calls succeed more". This is the same pattern found on the Qwen3.6-35B GGUF ("Unsloth fixes - developer role, tool calling"), where the baked template differed materially from upstream and silently DROPPED mid-dialogue system messages. So: diff the baked template against the official one before drawing any template conclusion, and do not swap templates without measuring - the tool-calling patch is plausibly load-bearing for tau2 scores. QUANT QUALITY CURVE (vendor, top-1 token agreement with BF16 vs GGUF size): UD-Q2_XXS ~82.5%, UD-IQ3_XXS ~90%, UD-Q3_K_XL ~92.5%, UD-Q4_K_XL ~95%, UD-Q5_K_XL ~96.5%, Q6_K ~97%, Q8_0 ~98%, UD-Q8_K_XL ~98.5%. The knee sits around Q4-Q5. This is the vendor's own metric on their own quants, not ours, but it is the stated basis for screening at Q8_0 (primary) with UD-Q4_K_XL as the comparability arm. REGISTERED PREDICTION (2026-08-14, before any measurement here): the vendor's table claims large agentic gains over Qwen3.6-27B - agentic terminal coding 73.0 vs 63.4, SWE-bench Pro 61.7 vs 53.5, DeepSWE 42.2 vs 13.3, QwenSWEBench 79.0 vs 49.3, long-horizon office work 70.7 vs 61.0. Our screened qwen36-27b-mtp scored 0.80 on the 5-task tau2 smoke (joint-strongest in the field). PREDICTION: Qwen3.8-27B meets or beats that on the same smoke. Recorded per the predictions-tested-in-public principle; the vendor numbers are vendor-claim provenance and are NOT evidence for our gate. VENDOR BENCHMARK PROVENANCE (full tables, 2026-08-14) - read the footnotes, they matter. Qwen3.8-27B's LARGEST margins over Qwen3.6-27B sit on benchmarks Qwen designed and runs itself: QwenSWEBench 79.0 vs 49.3 ("in-house coding benchmark", avg@3, 8h timeout) and CoWorkBench 70.7 vs 61.0 ("in-house cowork benchmark"); on the VL side RecreationBench 47.1 vs 29.8 is likewise "an in-house, long-horizon application-recreation benchmark". On third-party benchmarks the margins are real but smaller (SWE-bench Pro 61.7 vs 53.5, IFBench 79.5 vs 69.1, LiveCodeBench v6 90.3 vs 83.9), and Opus 4.6 Max still leads on several (agentic terminal coding 78.2 vs 73.0, NL2Repo 47.6 vs 42.3, GPQA 91.3 vs 89.2, HLE 40.0 vs 30.8). Also note SWE-bench Pro was run with "problematic tasks corrected and all baseline models re-evaluated on the refined benchmark" - vendor-modified, not the stock benchmark. None of this is an accusation; it is the ordinary reason vendor tables are provenance vendor-claim in this record. It does sharpen why an independent agentic measurement is worth taking: EVERY headline claim here is agentic/tool-use, which is exactly the axis tau2 measures, so our screen is a direct independent check on the claim class where the vendor's own margins are largest. VISION CLAIMS (mmproj-F16 acquired, not yet screened - the NPU/vision tiers do not currently define a vision screen): computer use OSWorld-Verified 84.3 vs Opus 4.6 Max 72.7; browser use WebArena-Verified 64.8; mobile AndroidWorld 81.9. Several VL rows are reported twice, "Without CI" and "With CI" (code interpreter), and the gap is large - BabyVision 65.7 -> 85.6, MathVision 90.0 -> 94.6 - so any vision comparison must state which setting it used or it is not comparable. Vision2Web scores are "judged by gpt-5.4"; HLE is "judged by GPT-4o" - LLM-judged rows, not deterministic scoring. ⚠ FIRST STRIX HALO THROUGHPUT (community, 2026-08-14, r/StrixHalo 1vobzvd, uncanny_instinct/thefrontierlab.ai - the same reporter as the nine-model KV sweep). Vulkan on Nathanw1014's strix-halo-vulkan fork (build 10283, b7b85da9), Bosgame M5, 128 GB, ~124 GiB GPU window, amd_iommu=off, FA on, batch 2048/ubatch 512, pp512/tg128 llama-bench medians, 2 reps, <0.8% drift: UD-Q5_K_XL (20 GB) f16 KV: 347.0 pp / 10.5 tg @d0 · 252.4 pp / 9.4 tg @d32768 Q8_0 (29 GB) f16 KV: 340.3 pp / 7.5 tg @d0 · 249.5 pp / 6.9 tg @d32768 KV quant buys 4-6% decode, costs ~2% prefill (small because only 16 of 64 layers carry a conventional cache - the 3:1 linear-to-full attention ratio). His same-box Qwen3.6-35B-A3B Q5_K_M anchor (779 pp / 47.4 tg @32k) sits within ~12% of our own 698/50 for that model, so his box and ours are performance twins and these rows transfer with a light correction. ⇒ EXPECTATION FOR OUR SCREEN (recorded before measuring): roughly 225-250 pp / 10-11 tg at d32768 for UD-Q5_K_XL, i.e. about 3x worse prefill and 5x worse DECODE than the Qwen3.6-35B-A3B incumbent. The cause is architectural and not a tuning artifact: this is a DENSE 27B reading every weight per token, against the incumbent's ~3B active MoE. The reporter's own memory-cliff measurement corroborates it physically (227 GB/s past LLC / 20.2 GB = ~11 t/s, matching his 10.5). ⇒ THIS CHANGES THE CANDIDATE'S PREMISE. It was listed as "reflex-tier footprint with executor-tier capability". The footprint claim survives (17-29 GB fits trivially); the REFLEX-TIER claim does not - ~10 t/s decode is executor-tier latency at best, and far slower than the incumbent it was proposed to complement. Whatever capability the screen finds must be weighed against that, and the registered tau2 prediction above remains a CAPABILITY prediction only - it says nothing about speed and is not refuted by these numbers. Comparability gaps in the community figures (why our own run is still worth taking): community fork not stock; amd_iommu=off worth +34-38% dense prefill by his own A/B (and mutually exclusive with our NPU lane); Vulkan only - NO public ROCm figures for this model exist on this silicon; mmap status unstated (llama-bench defaults to mmap, and we established today that matters); pp512-at-depth is not a real 32k prefill; 2 reps against his usual tighter drift; no raw JSON published yet; and explicitly NO quality or agentic testing - his own closing line is that he has no idea yet whether 3.8 is better than 3.6 at anything. ⇒ OPEN UPSTREAM QUESTION worth instrumenting on our run: he reports an 8.1 GiB KV footprint at 32k for this dense 27B. That matches allocating a full attention cache across ALL 64 layers (~8.6 GiB) rather than the 16 that actually carry one (~2.1 GiB). If llama.cpp is not exploiting the hybrid layout, roughly 6 GiB is being wasted at 32k and it scales with depth - measure our own KV allocation before accepting the fit maths. ⇒ MTP: both GGUFs carry MTP tensors (mtp_num_hidden_layers 1); the reporter has not tested speculative decode. His prior work found llama.cpp's old draft default of 16 cost up to 75% of generation throughput on this family and that the optimum walks left with depth. Screen without MTP first, per the existing note. RECOMMENDED CONFIG FROM MEASUREMENT (2026-08-15, the qwen38-screen result — this is the configuration the screen's own numbers select, recorded as cfg-0066): backend ROCm (Vulkan prefill collapses at depth and loses the device at d131072 — clm-0054) quant Q8_0 (the designated primary screening quant; the capability result was measured on it) spec draft-mtp, --spec-draft-n-max 3 — HARD ceiling 3 until the EOS bug is understood (n_max>=4 corrupts generations in live sessions, clm-0055); MTP is mandatory for usable decode (2.33x floor) and capped for correctness kv f16, -c 32768 (KV costs 64 KiB/token — context is nearly free, clm-0056; 32768 is the measured serving point) thinking reasoning_effort=medium via --chat-template-kwargs (llama.cpp discards the request field pending PR 26941; xhigh is unfinishable at this decode speed) with -n 4096 as the output cap misc --parallel 1 (issue 25992) --load-mode none (clm-0053) --jinja --reasoning-format deepseek --slots REFLEX-PREMISE VERDICT (2026-08-15): the screen split the original premise in two and settled both halves. The CAPABILITY case is made — the registered smoke prediction was met at 1.000 vs the incumbent's 0.80 on identical tasks, so "executor-tier capability at this size" survived its pre-registered test. The REFLEX-TIER case is dead: ~18 t/s decode WITH speculation (~7.8 t/s floor) against the ~50 t/s MoE incumbent is executor-tier latency, exactly as the community throughput rows predicted before our screen. What this model earned is a capability slot, not a reflex slot — where it fits in the tier map is an adoption decision that the speed number, not the score, will drive.
Qwen3.8-Flash-Next
Qwen4-architecture preview (released 2026-08-26): a 125B MoE with a large N-gram/PLE embedding-table component and hybrid QSA (linear/full) attention, designed to fit a single 128 GB Strix Halo with the tables held off-GPU. Listed as the highest-interest new candidate for the fleet's strong- orchestrator role — reflex-tier active parameters (~6B) with frontier-tier agentic capability, if it holds up on tau2.
gate history (8)
- 2026-08-26 · listed
Qwen4-arch preview released. Runtime support is a draft PR (#27742) on the Unsloth qwen4exp fork; official llama.cpp master lacks qwen4exp. GGUF quants published by Unsloth. Listed pending an admitted runtime + a clean build.
- 2026-08-26 · listed → acquired
UD-Q4_K_XL (the competence-validated Q4 class) staged to the Ladon library and byte/sha256-verified: 4 shards totaling 111,334,654,784 bytes (shard-1 sha256 4448186216…), matching the current HF revision exactly. Then copied to aihydra local NVMe for benchmarking.
- 2026-08-27 · acquired → screened
Screen tier on-box, Vulkan r3 build 250b6144. FIT: UD-Q4_K_XL loads on Vulkan from c8192 up to c262144; GTT-resident 76.7 GiB with ~43 GiB free for KV (the ~27 GiB N-gram/PLE table is placed on CPU automatically, so it never occupies GTT). GUARD: a direct tool-call gate passed (valid get_weather call) plus a competence battery (arithmetic 17*23=391, instruction-following, a correct sum-of-squares one-liner) all completing cleanly (finish=stop) — the prior "no competence" was purely a max_tokens=16 truncation, not a model fault. SMOKE: 2-task tau2 airline subset (tasks 0-1, seed 42, claude-haiku-4.5 simulator) both reward 1.0 before committing to the full-26. This is the tool-call+smoke screen, not the formal 4/4 needle guard at depth 8000 (disclosed deviation).
- 2026-08-27 · screened → benched
Built the Unsloth qwen4exp fork at HEAD 250b6144 (Vulkan/RADV; the ROCm mmap-load path stalls on gfx1151, so Vulkan is the measured backend). Tool-calling validated, then FULL tau2 airline 26 (cfg-0177, seed 42, claude-haiku-4.5 simulator): mean_reward 0.9231, 24/26 at reward 1.0 — LEADS the board over Ornith-1.0-35B UD-Q4_K_XL (22/26, 0.846). 184 tool calls, 0 empty-argument calls, 24.7 messages/task. Wall-metered energy joined same-day (eng-0275): 9.90 Wh per correct answer, ~4x cheaper than the ho003 stock-f16 control (42.06). Footprint: GTT 76.7 GiB (the 26.8 GiB N-gram table is placed on CPU automatically, never in GTT), ~43 GiB free for KV, KV ~48 MiB/1000 tokens. Performance: prefill 277-313 tok/s at 5-8K, decode ~22 tok/s short / 8.5 at 131K, TTFT ~2.9 s. Caveats (load-bearing): single trial N=1, reasoning_effort=low (not default xhigh), draft/WIP runtime, GGUF LICENSE 404 + unpinned conversion-base provenance = LAB-ONLY no redistribution, MTP unavailable (no head layers in the GGUF), 262K Vulkan workgroup-count crash (cap -c <= 262140).
evidence: clm-0124
- 2026-08-28 · benched → acquired
Independent review retained every first-look record but held its current leader, energy, production and role-fit interpretation. The candidate returns to acquired while an intended/default-reasoning protocol screen, formal guard and repeated full capability arm are prepared. This is an evidence-preserving supersession step, not deletion of the earlier run.
evidence: clm-0125
- 2026-08-29 · acquired → acquired
The separate KingJones R2 full-STRIX ROCmFP4 artifact/runtime campaign stopped before model execution. It verified the exact tracked 1,017-byte rotate-bits header, clean-applied the admitted receipt-only capture patch, then failed the targeted llama-server build because sha256.c could not resolve that header include. clm-0126 records this terminal negative and the explicit absence of guard, Tau2, performance, depth, cache and energy results. It neither advances nor rejects the model and does not merge with the earlier Unsloth UD-Q4_K_XL/Vulkan series.
evidence: clm-0126
- 2026-08-30 · acquired → benched
The separate KingJones full-STRIX ROCmFP4 campaign completed on the exact HIP/gfx1151 runtime and was independently admitted BOUNDED/PARTIAL. It contributes strict 8k-budget Tau2 all-attempt coverage, corrected served QSA prompt processing, nonzero-output generation rows at selected depths, separate prefix-reuse traces and bounded whole-file/memory receipts. A separate post-hoc wall-meter join recovers the full Tau2 window while explicitly not claiming campaign-time metering compliance. Exactly two original Tau2 calls reached the inherited 8,192-token ceiling; a later two-task 16k recovery remains held for missing lossless transport evidence and contributes no mixed-budget score. Incomplete per-PLE mincore, two engine CV trigger rows, zero-byte generation attempts and the full-window HTTP 400 remain explicit exclusions. This is a bench gate, not a production or role-fit recommendation.
- 2026-09-02 · benched → benched
Third serving series benched, distinct from both the Unsloth UD-Q4_K_XL first-look and the KingJones full-STRIX arm: the agention ROCmFP4-FAST-v2-ple16 imatrix quant (87.06 GiB @ 4.23 bpw, per-head n-gram/PLE, fully GPU-resident) on the agention/Laurent Vulkan fork (cfg-0182, LaurentZuijdwijk/llama.cpp vulkan/qwen4exp-rocmfpx commit 5e085d12, build b10809; branch tip verified, quant byte + sha256 verified). Served single-slot with adaptive MTP: ~34 tok/s decode on the live tau2 workload; a tokenizer-exact served depth sweep on source code holds 16-26 tok/s across 32-200K (rising with MTP acceptance, run-0654..run-0657), prefill 298 -> 127; warm agent-turn TTFT ~0.6 s; loads in ~60 s, fits fully on GPU (GTT ~95-106 GiB), no host-RAM thrash or OOM. tau2 airline full-26 = 21/25 (0.84, run-0653) at 12.74 Wh per correct answer (eng-0277). N=1 and at the model's DEFAULT reasoning effort (as deployed), NOT reasoning-matched to the UD-Q4_K_XL reasoning=low baseline, so no cross-series capability or energy ranking is drawn. Adopted into fleet production serving for its speed and GPU-resident fit, pending a matched reasoning=low re-run. This series does not transfer identity or results from the Unsloth or KingJones arms.
evidence: clm-0128
- note
BENCHED WITH THREE SEPARATE SERIES. The historical Unsloth UD-Q4_K_XL/Vulkan first look remains visible under clm-0125's held headline boundary. The KingJones full-STRIX ROCmFP4/HIP series is current only within clm-0127's bounded admission: it does not transfer identity or results from the Unsloth arm and does not establish a production recommendation, complete PLE series, unrestricted engine floor, or publisher-equivalent result. A third series, the agention ROCmFP4-FAST-v2-ple16/Vulkan quant (clm-0128), is the one now in fleet production serving — single-slot with adaptive MTP, 21/25 (0.84) on tau2 at 12.74 Wh/correct — admitted N=1 and default-reasoning (not reasoning-matched to the Unsloth arm), and likewise non-transferable to the other two series.
Qwen3.6-35B-A3B
The true incumbent: what Warden actually ran, with best-in-class tool-calling at ~96-100 t/s. The status-quo baseline every challenger must beat.
gate history (8)
- 2026-06-20 · listed
Listed as the status-quo incumbent: whatever replaces it has to beat a model with proven best-in-class tool-calling in live use.
- 2026-06-27 · listed → acquired
Already in production on the Mac mini, so it was staged on the new box rather than acquired - the incumbent needs no justification to be present.
- 2026-06-27 · acquired → screened
Phase A: ~11s time-to-correct, 100% tool-call gate.
- 2026-08-09 · screened → benched
Throughput, FA, ngram speculation and tau2 arms run. Speculation is a dead end on this model (no MTP head).
evidence: clm-0026
- 2026-08-20 · benched → benched
HO-009 admitted a separate genuine MTP artifact, Qwen3.6-35B-A3B-MTP UD-Q4_K_M, with qwen35moe.nextn_predict_layers=1. Reviewer-admitted evidence proves MTP activation/safety/performance for n_max=2 through d204800 only; it does not inherit capability from this incumbent artifact, does not admit n_max >=4, and drops n2 d262144 as an HTTP 400 context-fit refusal with no substitution.
evidence: clm-0097
- 2026-08-21 · benched → benched
HU-004 cleared the long-prompt guard for the same genuine MTP artifact, Qwen3.6-35B-A3B-MTP UD-Q4_K_M. The raw llama-server guard passed 20/20 cells across plain (cfg-0141) and native-MTP n2 (cfg-0142) arms at the ~8K/16K/17.5K/20K/40K prompt-token brackets with natural and native tools fixtures; no silent-empty-success, no empty-tool-arguments, no explicit context refusal, no HTTP error, no timeout, no server crash. This clears ggml-org/llama.cpp#27442 for this exact ROCm0/gfx1151 f16-KV n2 fingerprint (context 204800). Guard evidence only; it is not a capability, throughput, or production result, and does not inherit capability from the incumbent UD-Q4_K_XL artifact.
evidence: clm-0101
- 2026-08-22 · benched → benched
Fork v0.6.8/v0.6.9/v0.6.10 upstream context added (community report evidence, not a new HaloBench run and not a gate change). The Nathanw1014/strix-halo-llamacpp v0.6.9 rollback-exactness tradeoff (con-0007/clm-0105) named THIS model's architecture as the affected hybrid GDN target: on Vulkan the fork reverted token-exact MTP rollback (full sequence-state checkpoints) after a deterministic hang, returning to snapshot-plane restore that "can diverge slightly from a no-draft run after a rejected draft". That tradeoff was REVERSED in v0.6.10 (con-0009/clm-0107): the hang's root cause was the server re-verifying replayed draft tokens after a checkpoint restore, fixed by 9c5d899 with full-checkpoint MTP rollback re-applied (f25eefe), so rollback is again token-exact on the current stable fork. Recorded for MTP-correctness provenance only — fork-vendor release-note statements on gfx1151 Vulkan, not the BENCHMARKS.md protocol, so they inherit into no HaloBench number. The MTP activation/guard results on this candidate (clm-0097, clm-0101) were measured on upstream ggml-org/llama.cpp 7077abb on ROCm0/gfx1151, which never used the fork's snapshot-plane path, so the tradeoff never directly applied to them and v0.6.10 changes no comparison boundary; the house varied-prompt/raw-token MTP-invariance requirement stands intact (upstream #25618/#26750 arms).
- 2026-08-23 · benched → benched
HO-013 extended the capability matrix on the genuine Qwen3.6-35B-A3B-MTP UD-Q4_K_M artifact to MTP n_max=4 and the reasoning-ON lane (build 2586f6edd, v0.6.10 fork/Vulkan, f16/f16 KV, c32768). Reviewer-admitted: n4-off full-26 is 17/26 (mean 0.6538), statistically identical to n2-off (17/26, p=1.000) and not significantly worse than plain (22/26, p=0.199) — n_max=4 is FLAT, does NOT narrow the 5-correct deficit to plain, does NOT degrade/collapse (EOS-cliff sentinel PASS 0/8). The reasoning-ON path is BLOCKED, not degraded: the plain-rea-ON content sentinel FAILED G6 with empty assistant content (reasoning_content_len=1846, assistant_content_len =0) — the same empty-content defect lineage that forced reasoning OFF in HO-009-AB, re-emerging on this fork under -rea on. Cell-2 n2-ON tau2 skipped and gated Cell-3 n4-ON cancelled. Scope single-model/single-build MTP n2|n4 × reasoning on|off; no cross-model/KV-quant/DFlash2/n>4 claim. Plain mode remains the capability winner on this stack.
- note
Vision audit (2026-08-10): parent IS vision-capable - natively multimodal per Qwen's own materials, processing images/documents/video as a core architectural capability, not a bolt-on. Confirmed on our staged artifact: Qwen3.6-35B-A3B-UD-Q4_K_XL.gguf reads general.architecture=qwen35moe, carries the "image-text-to-text" tag, and has qwen35moe.rope.dimension_sections=[11,11,10,0] - the same MRoPE sectioning implemented natively in src/models/qwen35moe.cpp. unsloth/Qwen3.6-35B-A3B-GGUF (our source repo) publishes mmproj-F16.gguf (899,283,680 bytes), not staged. Runtime: LLM_ARCH_QWEN35MOE is registered in our build (commit 3653e6d), and clip.cpp's PROJECTOR_TYPE_QWEN3VL path is arch-agnostic on the LLM side. Verdict: mmproj-download-away, not text-only - this is the direct MoE sibling to qwen35-122b's vision path and shares the true incumbent's own architecture, so a vision test here would double as an incumbent-vs-challenger comparison rather than a one-off.
DeepSeek-V4-Flash-0731
A community report on this model independently confirmed our mmap/GTT double-residency finding and demonstrated a 120 GiB GTT ceiling in production - so the report is credible on details we can check. A second, richer community measurement (2026-08-12) puts it at 23-24 t/s steady-state decode on our exact silicon with a speculative drafter, and the community is openly asking a quality question our rig can answer with data.
gate history (8)
- 2026-08-09 · listed
Community report worth testing directly.
evidence: clm-0031
- 2026-08-10 · listed → listed
Partly MIS-FRAMED as a model addition. Its most consequential claim - Vulkan beating ROCm on 3 of 4 cells, including 55% faster decode at depth - is about the BACKEND, and is testable on models we already hold, more cheaply and more decisively. Acquire it as a model too, but do not let acquisition block the Vulkan test. Fit against 104 GiB GTT is unverified.
evidence: clm-0031
- 2026-08-12 · listed → listed
Identity now concrete via the r/LocalLLaMA Strix Halo guide thread: DeepSeek V4 Flash 0731, deepseek4 arch, 256 experts/6 active + 1 shared, MIT licence, 1M native context. Unsloth UD-IQ3_XXS is ~98 GB (a commenter says 104 GB, unanswered) - tight but plausible against the 104 GiB GTT budget; the thread's own numbers were produced on 126976 MiB GTT with a 4 GB BIOS carve-out, so fit at OUR boot params still needs a probe before any download decision. Author's steady-state on this exact silicon: 23-24 t/s decode with an 11 GB bf16 DSpark drafter (1.46x over plain decode), ~236 t/s prefill, on Nathan's Vulkan fork - headline 27 t/s is a peak-window cherry by his own caveat. KV note for our protocol: his KL-divergence data on q8_0 KV shows a fat tail (99.9%-ile KLD 64x the mean) and he recommends f16 KV for agentic work - consistent with clm-0035's direction. Open community question our tau2 rig could settle: antirez Q2-Q4 quants vs Unsloth UD-IQ3_XXS quality (currently argued on vibes in both directions). Backend decider spun out as coverage.md item 0; this candidate stays acquisition-gated on the fit probe.
evidence: clm-0031
- 2026-08-15 · listed → screened
Acquisition gate SETTLES: fit probe passed. Stock ROCm 3653e6d, UD-IQ3_XXS, f16 KV (per the COMMUNITY HAZARD note below - always), -c 32768, --load-mode none: loaded clean in 63s, gtt_used 100,422 MiB against the 122,880 MiB boot budget (~22.5 GiB headroom, comfortably fits). House GUARD 4/4, no cliff detected (coherence, toolcall, needle@8000, isolation-skipped at --parallel 1). SMOKE (5-task tau2 airline, pinned haiku-4.5 simulator, seed 42, max_tokens 4096): 5/5 @ 1.000, VALID under the SMOKE gate (48 tool-call messages, 0 empty turns, 0 infrastructure errors) - a clean pass on the first try, no retries needed. One llama-bench cell (pp1024/tg256, f16 KV, 3 reps): d0 pp~139 t/s / tg~15.3 t/s; d32768 pp~81 t/s / tg~12.1 t/s (no drafter run - this is the plain-decode floor the community's 23-24 t/s DSpark-drafter figure sits above; ~15.3 t/s here divided by their claimed 1.46x drafter speedup lands within noise of their own plain-decode baseline). Numbers pending full ingest into run/config records.
- 2026-08-16 · screened → benched
Full bench complete. Throughput matrix (both backends, f16 KV mandatory, d0 through d262144): ROCm swept every cell clean (140.80/15.30 t/s at d0 down to 21.52/6.86 t/s at d262144 pp/tg); Vulkan completed only d0 (134.73/12.44 t/s, ROCm ahead on both phases) and lost the GPU device on both allowed attempts at every depth from 32,768 tokens on (rc=134, vk::DeviceLostError, kernel-evidenced amdgpu ring resets) - ROCm is the required backend at any useful depth [clm-0059]. KV-cost-per-token measured via a two-point GTT delta at --parallel 1: ~7.13 KiB/token, about a ninth of a conventional hybrid-attention model's - this candidate's full 1,048,576-token native context projects to fit the 120 GiB GTT window with ~16 GiB to spare, so depth was never a fit risk here [clm-0060]. Full standard 26-task tau2 airline run (ROCm, f16 KV, c32768, no speculation - this candidate has no native draft/MTP head in the stock build): completed in full (rc=0, not a wall-bound cut) at 0.846 mean reward (22/26), 272 tool-call messages, 0 empty turns - VALID; the first five tasks reproduced the screen's 5/5 = 1.000 exactly on the same seed [clm-0061]. Energy: 24.29 Wh per correct answer whole-session (0.74 pence at 30.3 p/kWh), joined same-session immediately after completion [clm-0061]. Model page published with page: true.
- 2026-08-17 · benched → benched
Part 2(d) stretch cell: ROCm d524288 (2x this candidate's previously-tested matrix ceiling of d262144), ONE attempt per the job brief, budget allowed (>=1h remained when started). RESULT: TIMEOUT, not a completion or a crash. The cell ran the full 7200s (2h) cap under active, healthy compute the whole time (process stayed in running state, 200-220% CPU, GTT usage settled at 108,454,141,952 bytes = 101.0 GiB -- comfortably inside the 120 GiB GTT budget, so this was never a fit/OOM problem) and llama-bench never finished even the pp1024 prefill measurement at this depth before being killed by the timeout wrapper (rc=124). No vk::DeviceLostError-class signature, no crash, no dmesg hazard -- this is a pure wall-clock-cost finding: prefilling to 524,288 tokens on this 97 GiB IQ3_XXS model, at this box's house ROCm flags, costs MORE than 2 hours of compute, an order of magnitude past the ~15-20 min this candidate's own d262144 cell took in the original matrix. No run record created (no metrics were produced to ingest) -- the timeout itself is the result, per the job brief's own worked precedent for a "2-attempt cap + dmesg if device-lost" class outcome; here the outcome class is timeout, not device-loss, and is recorded the same way: explicitly, not silently. Run-meta: depth-backfill-part2/deepseek-v4-flash rocm-d524288-stretch 2026-08-17T13:36:38Z -> 15:36:38Z (7200s, rc=124).
- 2026-08-18 · benched → benched
Follow-up on the carried Strix Halo Vulkan fork at baf6360be changes the backend-stability finding without changing the production recommendation yet. The same UD-IQ3_XXS artifact and f16 KV passed the house 4-item guard and completed clean throughput cells at d0, d32768, d131072, d204800 and d262144, with no device loss. This supersedes the earlier stock-build-only statement that Vulkan could not survive beyond d0, but the fork bundles multiple changes and received only a binary guard rather than the full capability suite, so the result is recorded as fork-specific evidence.
evidence: clm-0084
- 2026-08-20 · benched → benched
Narrow HO-001 v0.6.6 follow-up on Nathanw1014's portable Vulkan payload (source 7b6c6133/build 10569) admitted matching f16/f16, q8_0/q8_0 and q4_0/q4_0 K/V cache guard+throughput cells through d262144, with reviewer-admitted sparse-path proof for q8_0 and q4_0. This is fork-specific sparse-path/performance evidence only: no full capability result, no q8_0/q4_0 quality equivalence, no production recommendation, no stock llama.cpp attribution, and no wide speculation/verify claim.
evidence: clm-0096
- note
The thread's quantised-drafter crash (invalid token = -1 on a Q2K draft) is a llama.cpp bug with a claimed GGUF-header fix ("Confirmed works", no patch link) - if real it cuts the draft footprint 11 GB -> ~6.5 GB, which matters at this model's size. Drafter must otherwise be bf16. COMMUNITY HAZARD (thefrontierlab.ai field report, 2026-08-13, r/StrixHalo 1vn9kn4): run f16 KV ALWAYS on this model. Quantising the K cache trips an open upstream llama.cpp issue - an incoherence rotation that disables the sparse-attention dispatch (dense fallback) and corrupts output on CPU/CUDA. The reporter could not reproduce the corruption on Vulkan (sparse buffers still allocate) but still measured ~-13 to -14% on BOTH prefill and decode from KV quant - opposite sign to every GQA model in his nine-model sweep. His build refused mixed cache types, so K-only isolation is untested. Unverified here; screen with f16 KV only.
Ornith-1.0-35B
Ornith-1.0 (ornith-ai / deepreinforce-ai, MIT, released ~2026-06-21) declares arch qwen35moe -- the SAME architecture family as our production 122B (hybrid linear/full attention, head_dim 256, eos id 248046) and as the already-screened qwen38-27b dense sibling, so it is both an active-param scaling point on our own production architecture (A3B-class MoE vs the 122B's larger MoE and qwen38-27b's dense shape) and a heavily-discussed Qwen-family alternative in its own right. It sits at the same MTP-variant / EOS-cliff intersection this lab already holds evidence on: identical tokenizer/eos token to the model clm-0055's EOS-cliff was characterised on, so any MTP arm on this model inherits that open question rather than starting from zero. Also carries a vision file in the library from day one via unsloth's mirror (see caveat below -- the vision claim is NOT independently confirmed).
gate history (9)
- 2026-08-16 · listed
Listed on the architecture-match rationale above (qwen35moe, same family as production 122B and qwen38-27b) plus the model's own MIT-licensed, vendor-published agentic-coding benchmark table (Terminal-Bench 2.1 64.2, SWE-bench Verified 75.6, both ahead of Qwen3.6-35B on the vendor's own numbers) -- vendor-claim provenance, not evidence, per house policy, but a real published measurement all the same. HF API confirmed both target repos exist and carry the intended artifacts before acquisition began: official ornith-ai/Ornith-1.0-35B-GGUF ships Q4_K_M/Q5_K_M/Q6_K/Q8_0/bf16 with NO mmproj; unsloth/Ornith-1.0-35B-GGUF ships the full UD-quant ladder plus mmproj-{BF16,F16,F32}.gguf. gguf.total (34,660,610,688, ~34.66B params) and gguf.architecture ("qwen35moe") match on both repos.
- 2026-08-16 · listed → acquired
Downloaded to aihydra ~/models/ornith-35b/ and sha256-verified exact against the HF LFS oid on all three files: Q8_0 (primary screening quant, OFFICIAL repo) 36,903,138,880 bytes, sha256 cbc992bca07901c1a51f33e65e6fc5d687de179c852a772dfd15e4c3261dbf5c; UD-Q4_K_XL (comparability arm, unsloth -- matches the quant class protocol.json quants.expected already carries) 22,324,804,000 bytes, sha256 67081ae4a1a291bd6c72834094ea056332cb3cb5fa15e88536ec7f233a475b71; mmproj-F16 (vision file, unsloth -- the official repo has none) 899,283,680 bytes, sha256 217ee3ae58ef7b1f743341a0037c9da50aa6ae06e051ccb49c517132b0ac2bf6. Mirrored to the NAS gguf-library (/volume1/Models/gguf-library/ornith-35b/) via the house scp -O -l 245760 pattern; MANIFEST.tsv updated there.
- 2026-08-16 · acquired → screened
House standard screen, stock 3653e6d ROCm, Q8_0, -np 1, c=32768. PROTOCOL_OVERRIDE carried on every run: Q8_0 is not in protocol.json quants.expected (same override rationale as qwen38-27b's cfg-0066 -- removes the quant confound, matches this candidate's own designated primary quant). FIT: loaded in 12s, 35,607 MiB GTT (comfortable against the 122,880 MiB boot window). GUARD: 4/4 (coherence, tool call, needle at 8000 tokens tok=5368 'chartreuse-viper-88', isolation skipped under -np 1). SMOKE: 5-task tau2 airline, pinned openrouter/anthropic/claude- haiku-4.5 simulator, temp 0, seed 42 -- 5/5, mean_reward 1.000, 21 tool-call messages, 0 empty assistant turns, 0 infra errors (VALID per the smoke gate; run-0330, run-0331, clm-0068). A joint-strongest smoke result on this box to date, directionally consistent with the vendor's own agentic-coding table. Not yet run: throughput/depth matrix, any MTP/speculative arm, vision-path test. energy: null on both runs -- run-meta.jsonl has the timestamp pair for a follow-up batch join before the ~10-day HA retention window closes.
- 2026-08-16 · screened → benched
Full bench per house pattern: throughput matrix both backends (ROCm 3653e6d / Vulkan 3653e6d6d), depths 0/32768/131072, pp1024/tg256, N=3 fresh-process reps <=32K and N=1 at 131072. Vulkan wins decode at every depth and prefill at d0/d32768 (ROCm only +2.6% ahead on prefill at d131072) -- and, notably, Vulkan did NOT lose the GPU device at d131072, breaking the device-loss precedent this silicon showed on laguna-s-21 and deepseek-v4-flash at comparable depth (clm-0070). Two-point GTT KV probe reproduced bit-for-bit across two independent probe pairs: 20.00 KiB/token (f16 KV), reconciled to within 0.04 GiB against the GGUF's own model_size from first principles (clm-0071). GGUF header parsed directly (no numpy/gguf-py available on-box; a minimal pure-Python GGUF v3 reader was written): zero nextn/mtp/eagle/medusa/draft tensors among all 733 -- no MTP/speculation arm run, plain decode throughout, matching the EOS-cliff watch this candidate's own screen flagged as moot for this reason. House guard re-run against the Vulkan tau2 serving config (cfg-0095, run-0338, 4/4, since a backend swap is itself the `binary` lever the guard exists to catch) before the full 26-task tau2 airline run: 0.8846 mean reward (23/26), 202 tool-call messages, 0 empty turns, 0 infra errors -- VALID, and the NEW LEADER among this programme's full-26-task tau2 comparisons (run-0345, clm-0072). Energy: 6.4834 Wh per correct answer -- ALSO the lowest of any full-26-task comparison on this box, a joint capability+energy lead this programme has not seen before (eng-0134, clm-0072). Screen-arm energy backfilled in the same session (eng-0132, eng-0133), closing the gap the screen's own gate entry flagged. PROTOCOL_OVERRIDE carried on every matrix/guard/tau2 run (Q8_0 not in protocol.json quants.expected, same rationale as the screen). Not run: MTP arm (none exists in this GGUF), vision-path test (candidate record's own open caveat), cross-quant comparison against the UD-Q4_K_XL comparability arm.
evidence: clm-0070 clm-0071 clm-0072 run-0332 run-0333 run-0334 run-0335 run-0336 run-0337 run-0338 run-0339 run-0340 run-0341 run-0342 run-0343 run-0344 run-0345
- 2026-08-17 · benched → benched
Depth-ladder backfill (protocol extension, same date): both backends extended past the original matrix's d131072 ceiling. ROCm (cfg-0104, guard reused from run-0330): d204800 173.66 pp / 22.23 tg t/s, d262144 -- this candidate's real declared context maximum (qwen35moe.context_length in the GGUF header, confirmed 262144 via direct header parse, not a YaRN projection) -- 141.96 pp / 19.37 tg t/s. Vulkan (cfg-0105, guard reused from run-0338): d204800 123.51 pp / 26.53 tg t/s -- SURVIVED cleanly, no vk::DeviceLostError, no GPU wedge, extending this candidate's already-unusual streak (the one full-bench candidate on this board where Vulkan broke the general device-loss precedent at d131072) one step further. Vulkan d262144 NOT attempted this session (time-boxed; a natural next cell for this streak). All four new cells single fresh-process reps (N=1, house convention past d131072). Both backends' full five-depth series are monotonic, no cliff, no scatter anomaly: ROCm pp 865.28/525.19/244.84/173.66/141.96 t/s, Vulkan pp 1063.66/679.22/238.54/123.51/(not run) t/s; ROCm tg 47.66/39.75/27.48/22.23/19.37 t/s, Vulkan tg 55.70/46.22/32.37/26.53/(not run) t/s -- Vulkan keeps its decode lead over ROCm at every depth measured, including the two new ones. Energy: not joined before session release (recorded explicitly per run, not left silent) -- HA history retention ~10 days, join before 2026-08-27.
evidence: run-0377 run-0378 run-0379 run-0380 run-0381 run-0382
- 2026-08-17 · benched → benched
Same-day follow-up: Vulkan d262144 (cfg-0106, guard reused from run-0338) -- the cell the prior entry flagged as time-boxed out -- run to completion: 84.24 pp / 23.31 tg t/s. SURVIVED cleanly, no vk::DeviceLostError, no GPU wedge, taking this candidate's Vulkan clean streak all the way to its model-max context (262144). This is the ONLY full-bench candidate on this board where Vulkan has completed its whole five-depth ladder (0/32768/131072/204800/ 262144) without a single device-loss incident. Complete six-depth series, both backends: ROCm pp 865.28/525.19/244.84/173.66/141.96 t/s (no d0 row -- ROCm's own series starts at the original matrix), Vulkan pp 1063.66/679.22/238.54/123.51/84.24 t/s; ROCm tg 47.66/39.75/27.48/22.23/19.37 t/s, Vulkan tg 55.70/46.22/32.37/26.53/23.31 t/s -- Vulkan's decode lead over ROCm holds at every depth measured, including the new model-max cell (23.31 vs 19.37 t/s, +20.3%). Energy joined live this session: eng-0152, 59.668 Wh over the 22.7-minute cell (mean 157.94 W, delta 147.84 W over the 10.1 W idle floor) -- notably lower mean draw than either ROCm deep cell in this session (eng-0149/eng-0150, ~179 W), a same-model cross-backend efficiency contrast worth surfacing per protocol.json's compute-surfaces comparison rule.
- 2026-08-20 · benched → benched
HO-005 reviewer-admitted UD-Q4_K_XL r3 bounded preflight ingested without broadening it into a full-26 capability or Q4/Q8 equivalence claim. The lossy weight-quant arm has its own measured bounded evidence only: artifact identity matched the acquired unsloth UD-Q4_K_XL file (22,324,804,000 bytes, sha256 67081ae4a1a291bd6c72834094ea056332cb3cb5fa15e88536ec7f233a475b71), FIT/header/load at c32768 was healthy, the Q4 Vulkan/f16-KV guard passed 4/4 with schema-visible run id run-0494, the 5-task tau2 airline smoke passed 5/5 at mean_reward 1.000 (run-0495), and one planned d262144 pp1024/tg256 N=1 throughput cell survived cleanly (run-0496/run-0497). Failed r1/r2 custody/wrapper attempts remain non-ingested operational artifacts only. Energy could not be joined by hbrunner because the active profile had no Home Assistant token/URL; each new run carries a structured energy_unjoined_reason with the exact window rather than a silent null.
- 2026-08-20 · benched → benched
HO-005 reviewer-admitted UD-Q4_K_XL full-26 successor ingested as measured Q4_K_XL capability, not Q8 inheritance and not Q4/Q8 equivalence. On the same cfg-0129 artifact/runtime/backend/f16-KV serving fingerprint, a fresh guard immediately before tau2 passed 4/4 (run-0498). The full task-id 0..25 tau2 airline run completed rc=0 at 22/26 with mean_reward 0.8461538461538461, failed task ids 7, 21, 23, 24, 185 tool-call messages, 0 empty-argument tool calls, 0 empty assistant turns, 0 infrastructure errors, 0 max-step cuts, and all termination reasons user_stop (run-0499, clm-0093). hborchestrator joined wall-meter energy after ingestion while HA history was still available: guard eng-0202, tau2 eng-0203; the Q4 full-26 tau2 arm used 146.5378 Wh total and 6.6608 Wh per correct task over its 22 correct tasks. The earlier clm-0092 remains bounded preflight/smoke/throughput history, and the Q4 energy join is not a Q4/Q8 equivalence or energy-ranking claim.
- 2026-08-20 · benched → benched
HO-005 reviewer-admitted UD-Q4_K_XL production-depth performance tuning matrix ingested as guarded performance/energy bookkeeping only, not as a Q4/Q8 capability equivalence or broad production recommendation. Three exact guarded fingerprints are schema-visible: ROCm f16/f16 KV (run-0501 guard, run-0503..run-0506 performance at d204800/d262144 N=1), Vulkan f16/f16 KV (run-0500 guard, run-0507..run-0516 performance at d0/d32768/d131072 N=3 and d204800/d262144 N=1), and Vulkan q8_0/q8_0 KV as a lossy-KV performance probe only (run-0502 guard, run-0517..run-0522 performance at d32768/d204800/d262144 N=1). Q5_K_M and q4_0/q4_0 KV remain excluded; r1/r2 wrapper exits remain operational deviations, not metrics. Energy could not be joined by this hbrunner profile because Home Assistant history credentials were absent; every new run records structured energy_unjoined_reason with exact window(s) rather than silent null energy.
evidence: clm-0095 run-0500 run-0501 run-0502 run-0503 run-0504 run-0505 run-0506 run-0507 run-0508 run-0509 run-0510 run-0511 run-0512 run-0513 run-0514 run-0515 run-0516 run-0517 run-0518 run-0519 run-0520 run-0521 run-0522
- note
VISION CLAIM CAVEAT (see clm-0068 for the full writeup): the official README's own highlights describe Ornith-1.0 purely as an agentic-coding family (9B-Dense / 31B-Dense / 35B-MoE / 397B-MoE) with zero mention of vision or multimodal input anywhere on the page, and the official repo ships no mmproj file at all. unsloth's mmproj-{BF16,F16,F32}.gguf exist and were acquired for the library, but they are plausibly a structural carryover from the Qwen3.5 base checkpoint ("post-trained on top of ... Qwen 3.5", the README's own framing) rather than evidence the coding fine-tune trained or evaluated vision capability. Do not cite "vision-capable" as a confirmed fact for this candidate without an actual vision-path screen against the unsloth mmproj -- which the current NPU/GPU screening tiers do not define (same open gap as qwen38-27b's mmproj-F16). MTP / EOS-CLIFF WATCH: GGUF metadata was not inspected for MTP tensors in this screen. If a later arm finds them, check clm-0055 (qwen38-27b, same architecture family, same tokenizer/eos id per this candidate's own listing rationale) before trusting any spec-draft-n-max >= 4 result -- that is exactly the corruption signature (target logits collapsing to im_end after accumulated session volume) this lab already holds evidence for on a sibling model, and the r/StrixHalo 1vpiwz0 audit (clm-0069, on qwen38-27b's own candidate record) found a THIRD-PARTY fork users are already running past that ceiling (spec-draft-n-max 6-7) without exact verification enabled by default.
Ling-3.0-flash (Ant Group)
A near-ideal CONTROLLED comparison for the 122B: same footprint, half the active parameters. That isolates active-parameter count as a variable, which nothing else in the field does.
gate history (7)
- 2026-07-26 · listed
The only candidate that functions as a controlled experiment rather than another data point.
- 2026-08-10 · listed → listed
Still listed, not acquired. Two hazards to handle first: it needs a non-stock llama.cpp, and its MTP head ships INACTIVE - benchmarking it as shipped would silently rig any decode comparison in Qwen's favour.
evidence: clm-0015
- 2026-08-13 · listed → screened
Downloaded overnight and sha256-verified (2-part Q4_K_M_STOCK GGUF, 70 GiB; server pointed at part 00001). Screening attempted on aihydra, build min-62bf73d (62bf73d25), Aug-12 protocol. FIT FAILED after 4s and the verbatim error IS the screening result (laguna precedent): "llama_model_load: error loading model: unknown model architecture: 'bailingmoe3'". Secondary probe on 77a9a66eb, the only by-date-newer registered build, fails with the identical line — unsupported on every registered general-purpose binary. No GUARD, no SMOKE. This measures what the 2026-08-10 entry predicted ("needs a non-stock llama.cpp"): the controlled active-parameter comparison against the 122B stays the most scientifically interesting question in the field, and it is now concretely blocked on upstream bailingmoe3 support landing in a protocol-registered build. The shipped-INACTIVE MTP head hazard from that entry still stands unmeasured and must be handled before any future decode comparison. Run-meta: screening/ling-30-flash 2026-08-13T07:53:17Z->07:53:25Z (8s, rc=1).
- 2026-08-17 · screened → screened
Re-screened same day the blocker merged: llama.cpp PR #26608 (BailingMoE3 Support) landed upstream 2026-08-17T07:49:49Z. Shallow-cloned current master to a fresh tree (7077abbe14c510cb829c93a1328c2815b5805ebd, 2026-08-17T11:14:33Z, depth 50) with the merge commit (373336672029b12e09f272bc027cc801345a3fd6) verified in-tree, and built BOTH variants from that one clone: build-rocm (-DGGML_HIP=ON -DAMDGPU_TARGETS=gfx1151) and build-vulkan (-DGGML_VULKAN=ON -DGGML_HIP=OFF), both Release. Registered as known build "7077abbe" in protocol.json (repo copy and box copy verified byte-identical, sha256 7e246207...af75adf9e92). Screened plain-decode only, NO speculation: the 2026-08-10 entry's shipped-INACTIVE MTP head hazard still stands, and there is no vendor-documented activation procedure, so enabling speculative decode against an untrained/inactive draft head would risk silently corrupting or no-op'ing output rather than measuring anything real — recorded as the reason for omission, not a silent gap. FIT PROGRESSED past the prior blocker (no longer "unknown model architecture") but FAILED differently on BOTH backends after ~3s each, and the verbatim error IS the result: "llama_model_load: error loading model: check_tensor_dims: tensor 'blk.0.ssm_f_a.weight' not found". Root cause read from the merged source: BailingMoE3's loader (src/models/bailingmoe3.cpp) requires a hybrid linear-attention forget- gate tensor (ssm_f_a) structurally shared with the Kimi Linear / Kimi K3 loaders — i.e. the merged architecture assumes a hybrid SSM/attention layer this box's downloaded Q4_K_M_STOCK GGUF (acquired 2026-08-13, pre- merge) does not contain. Two live hypotheses, neither resolved here: (a) the artifact was converted with an incompatible/older gguf-py tensor mapping and needs re-conversion from the original safetensors against the now-merged convert script, or (b) this checkpoint variant genuinely lacks the hybrid-linear layers the loader now assumes present. Identical failure confirmed on both build-rocm and build-vulkan (backend-agnostic — the check fires in generic tensor-validation code before backend dispatch), so this is not a backend gap. No GUARD, no SMOKE (house pattern: FIT-only result stands). Run-meta: ling-30-flash-screen ling-30-flash fit-rocm-7077abbe-c32768 2026-08-17T11:38:18Z->11:38:21Z (3s, rc=1); Vulkan probe run manually outside runmeta (identical verbatim error, same timeframe). NEXT: re-acquire/re-convert the GGUF from source safetensors against gguf-py's current bailingmoe3 tensor mapping before attempting another screen — the architecture gap is now closed, but the artifact-format gap is open and is a re-acquisition job, not a build job.
- 2026-08-18 · screened → screened
Corrected post-PR-26608 GGUF acquired and screened: bloomer010/Ling-3.0-flash-GGUF `Ling-3.0-flash-Q4_K_M.gguf`, 78,285,287,904 bytes, sha256 bcce6e32799749e8e52a52c161127b989db52430e00fbebc77f4d61ef94e754d, staged first on ladon then copied to aihydra local NVMe and re-verified. Header pre-check and the successful load both close the 2026-08-17 artifact-format blocker: this reference conversion declares `bailingmoe3` and contains the `ssm_f_a` tensors the AtomicChat `Q4_K_M_STOCK` pair lacked. Screened on the same post-merge registered build family as the failed re-screen, build-rocm commit 7077abb, plain decode only (NO speculative/MTP path; the shipped MTP activation question remains a separate hazard, not silently exercised). FIT on ROCm at c32768 passed in 44s with `--load-mode none`, `-fa on`, `--parallel 1`, `--jinja`, no mmap. GUARD 4/4 passed: coherence, native tool call (`get_booking {"reference":"ABC123"}`), needle at depth 8000 (`chartreuse-viper-88` at tok=5381), isolation skipped under parallel 1 per house rule. SMOKE 5-task tau2 airline with pinned Haiku-4.5 simulator completed 5/5 in 238s, mean reward 0.80, 26 tool-call messages, 0 empty assistant messages, 0 infrastructure errors, 0 max-step cuts — VALID under the zero-tool-call gate. Operator artifact note: an immediately prior smoke attempt in the same runroot is explicitly invalid infrastructure noise (OpenRouter key omitted from the environment, all five tasks 401 `infrastructure_error`); the authenticated rerun is the screening result. Runroot: aihydra `~/bench-results/ling-30-flash-corrected-screen-20260818T005734Z`; run-meta: `ling-30-flash-corrected-screen/fit-rocm-7077abb-c32768` 2026-08-18T01:00:06Z->01:00:50Z (44s, rc=0), `guard-rocm-7077abb-c32768` 01:00:50Z->01:01:01Z (11s, rc=0), `smoke-auth-rocm-7077abb-c32768` 01:03:16Z->01:07:14Z (238s, rc=0).
- 2026-08-18 · screened → benched
Full publishable bench completed on the corrected GGUF under the conservative ROCm/plain-decode path. Matrix cells recorded d0 prefill/decode and d32768 prefill/decode with no mmap; the same serving config passed the house guard 4/4 and then completed all 26 tau2 airline tasks with pinned Haiku-4.5 user simulator. Final capability result: 13/26 correct, mean reward 0.500, 170 tool-call messages, 0 empty assistant messages, 0 infrastructure errors, 0 max-step cuts. Wall energy was joined from raw Home Assistant history for the full tau2 window: 69.218 Wh total, 5.3245 Wh per correct answer. No speculative/MTP path was exercised, so the MTP activation question remains an open follow-up rather than part of this verdict.
- 2026-08-19 · benched → benched
HO-002 resolved the shipped-MTP hazard on the corrected GGUF and llama.cpp 7077abb. MTP n_max=1, 2 and 3 each proved real head engagement and nonzero acceptance, passed 15/15 varied sanity generations, and passed the house guard 4/4. Matched median decode ratios versus the same-session plain floor were 1.007x, 0.968x and 0.850x, below the 1.10x promotion threshold. Keep plain decode as production default; activation works but does not help this bounded workload. Two prior attempts stopped before any MTP probe and remain operational history only.
- note
Vision audit (2026-08-10): parent is text-only. Ant Group splits its model families by modality - Ling (text MoE, this candidate), Ring (reasoning), and Ming (multimodal: vision/speech/audio) are three separate lines, not variants of one checkpoint. Ling-3.0-flash carries no vision capability to inherit. Verdict: text-only, not a vision-lane candidate. Screening-relevant (2026-08-22, upstream watch, no gate change): llama.cpp PR #27508 "model: support DSpark for bailingmoe3" merged 2026-08-22T09:19:49Z, and Nathanw1014/strix-halo-llamacpp v0.6.10 cherry-picked it, so DSpark speculative decoding is supported for bailingmoe3 (incl. this candidate). The runtime is complete, but the path is NOT exercisable end-to-end until Ling-3.0-flash draft GGUFs are published (inclusionAI/Ling-3.0-flash-dspark holds model.safetensors only as of 2026-08-22T16:19Z). When the DSpark GGUFs land, re-examine this candidate against the DSpark draft path (a registered fork build with bailingmoe3 DSpark, plus a draft GGUF, would be required) — distinct from the already-benched native MTP n_max 1/2/3, which engaged (clm-0086) but stayed below the 1.10x promotion threshold. Recorded in clm-0107.
Nemotron-3-Super-120B-A12B
Purpose-built for agentic reasoning and tool use, which is exactly Warden's workload. Hybrid Mamba-2 + LatentMoE + MTP promises cheaper long-context KV.
gate history (8)
- 2026-06-20 · listed
Only model in the field designed for the workload Warden actually runs.
- 2026-06-27 · listed → acquired
Forced to IQ4_XS: Q4_K hit 77 GiB at 4K and OOMed by 16K on the 96 GB box.
- 2026-06-27 · acquired → screened
Phase A: slowest at ~28s, 100% tool-call gate. MTP REJECTED - Mamba is incompatible with draft-mtp, a deliverability finding not a config issue.
- 2026-08-09 · screened → benched
Ran the tau2 arms. NOT quant-equivalent to peers - still IQ4_XS while everything else is Q4_K_M, so every Nemotron-vs-peer comparison to date is confounded. aihydra's 104 GiB GTT lifts the constraint that forced it; Q4_K_M re-run is outstanding.
- 2026-08-15 · benched → benched
Q4_K_M re-run DONE - the fair-quant screen this record has flagged outstanding since 2026-08-09. House FIT/GUARD/SMOKE, pinned Haiku-4.5 simulator. FIT ran on BOTH backends in one session (clm-0050's amendment history: this model now runs on ROCm with --load-mode none): ROCm 79,472 MiB GTT / Vulkan 79,159 MiB GTT, both healthy and comfortably inside the 122,880 MiB boot budget, well under the resident-memory ceiling that failed it under mmap (clm-0053). Screened on ROCm (fleet default; no health differentiator between the two backends here). GUARD 4/4, no cliff. SMOKE (5-task tau2 airline, seed 42, max_tokens 4096): 4/5 @ 0.800 mean reward, VALID under the SMOKE gate (37 tool-call messages, 0 empty turns, 0 infrastructure errors) - task 1 failed (0.0), tasks 0/2/3/4 passed. This is the first Nemotron-3-Super capability number measured at the SAME quant as its peers; every prior number (clm-0035, retracted by clm-0036) was IQ4_XS-confounded.
evidence: run-0274
- 2026-08-16 · benched → benched
Full bench complete (throughput matrix + house guard + full 26-task tau2 + energy). Throughput matrix (both backends, f16 KV, d0/d32768 -- this candidate's fit ceiling in this job's scope; the earlier "OOM at d32768" reading, perf-matrix-gtt120 2026-08-14, was a mmap double-residency defect (clm-0053), not a memory fact about this model, confirmed by this session's clean 3/3-rep completion on BOTH backends under --load-mode none): ROCm wins prefill by ~27% at both depths (271.55 vs 213.73 t/s d0, 244.77 vs 192.83 t/s d32768) but Vulkan wins decode by 8-9% (18.20 vs 16.71 t/s d0, 17.74 vs 16.40 t/s d32768) with ZERO device-loss at d32768 -- the FIRST model in this programme where Vulkan decode beats ROCm outright [clm-0063]. KV cost measured at 8.00 KiB/token (two-point GTT delta), close to deepseek-v4-flash's 7.13 and far below qwen38-27b's 64.00 -- this candidate is nowhere near GTT-bound at any depth this job tested [clm-0062]. Because decode dominates a multi-turn agentic conversation, this job served the tau2 arm on Vulkan (cfg-0089) -- a fresh house 4-item guard ran against that exact config first (4/4, no cliff) per protocol's guard-substitutions rule. Full standard 26-task tau2 airline run completed in full (rc=0, not a wall-bound cut) at 0.7692 mean reward (20/26), 263 tool-call messages, 0 empty turns -- VALID, and this is the FIRST tau2 result for this candidate at ANY quant run under the protocol's pinned independent simulator (every prior full-run number, including the IQ4_XS-vs-Q4_K_M quant-flip pair above, ran self-play, which this candidate's own history already flagged as measuring something weaker); the run's first five tasks reproduce the 2026-08-15 screen's 0.800 exactly on the same seed [clm-0064]. Energy: 32.43 Wh per correct answer whole-session (0.98 pence at 30.3 p/kWh), joined same-session [clm-0064]. Model page published with page: true, UD-Q4_K_M as the primary series, UD-IQ4_XS kept as an explicitly separate historical series (never silently merged, per protocol's quant axis rule).
- 2026-08-17 · benched → benched
Depth-ladder extension past the fullbench's d0/d32768 ceiling (Part 2 deep cells, budget-aware N=1 per house convention past d32768). ROCm (cfg-0107): d131072 190.04 pp / 15.53 tg t/s, d204800 163.62 pp / 15.01 tg t/s -- both cells clean, no OOM, no device issue (clm-0053's mmap fix under --load-mode none continues to generalise at this candidate's own deeper cells). Vulkan (cfg-0108): d131072 148.62 pp / 16.62 tg t/s -- SURVIVED cleanly on the FIRST of the house 2-attempt cap, no vk::DeviceLostError, no GPU wedge, no dmesg hazard signature -- adding to this box's #25664 device-loss dataset as a second candidate (after ornith-35b) where Vulkan does not lose the device at this box's usual failure band. Vulkan keeps its decode lead over ROCm at this depth too (16.62 vs 15.53 t/s, +7.0%), consistent with the fullbench's own d0/d32768 finding that Vulkan decode beats ROCm outright on this candidate. d204800 not attempted on Vulkan this session (budget; ROCm's own two cells and the Vulkan attempt already consumed this priority slot's time). Energy joined live for all three cells: eng-0153 (ROCm d131072, 32.569 Wh), eng-0154 (ROCm d204800, 53.548 Wh), eng-0155 (Vulkan d131072, 39.369 Wh) -- mean draw within a few watts across all three (172-177 W), no meaningful backend-power divergence at this candidate's cheap 8.00 KiB/token KV cost.
evidence: run-0385 run-0386 run-0387 run-0388 run-0389 run-0390 eng-0153 eng-0154 eng-0155
- 2026-08-19 · benched → benched
HO-006 filled the previously missing Vulkan d204800 throughput cell under the existing cfg-0108 fingerprint, using the same stock llama.cpp 3653e6d6d Vulkan / UD-Q4_K_M / f16 KV / --load-mode none boundary as the d131072 Vulkan rows. Single admitted llama-bench cell, rc=0, contention=false: 130.8127 t/s prefill and 15.8832 t/s decode. Narrow comparison only: versus cfg-0107 ROCm d204800, Vulkan is slower on prefill and faster on decode; versus cfg-0108 Vulkan d131072, both rates fall with depth. Energy is explicitly unjoined on the run rows pending a Home Assistant counter join; this is throughput evidence, not a capability or production verdict.
- note
Vision audit (2026-08-10): parent is text-only, confirmed directly - Nemotron-3-Super-120B-A12B's own documentation states it supports text input only, no image/vision modality anywhere in the line. Staged artifact (arch nemotron_h_moe) carries no vision tensors. Verdict: text-only, not a vision-lane candidate.
Deep-Thought-Posttrain
A post-trained constant-output model (SmolLM2-360M-Instruct base, fine-tuned to produce elaborate chain-of-thought inside <think> tags and then answer "42" to every query, regardless of content) is a genuine null-hypothesis / negative-control candidate for this lab's own capability harness: every instrument in the house protocol (coherence, tool-calling, retrieval, the SMOKE validity gate, wall-metered Wh-per-correct-answer) assumes a model is at least attempting the task it is asked. Running the full protocol against a model that provably never is tests whether the instruments themselves fail safe or fail silent.
gate history (4)
- 2026-08-17 · listed
Listed on operator direction as a deliberate negative-control candidate for the house benchmark suite, to be run with the same rigour as every capability candidate rather than as a joke entry — the humour, if any, is expected to come from the deadpan measurement, not from the write-up winking at the reader.
- 2026-08-17 · listed → acquired
always42-universal.gguf (726 MB, F16, sha256 fa0a43b12b90ed9e8aa0e3c346b8f0e4e627d971f171a0b5b549fad31ce2d5df) downloaded to aihydra ~/models/deep-thought-posttrain/ and mirrored to the NAS gguf-library (/volume1/Models/gguf-library/deep-thought-posttrain/, remote size verified byte-identical to local). Manifest updated (~/gguf-expected-manifest.txt). GGUF header parsed directly (this job's own from-scratch parser, gguf_probe_dtp.py, no gguf-py/huggingface_hub dependency): general.architecture=llama, 32 blocks, 290 tensors (225 F16 / 65 F32), llama.context_length=8192, embedding_length=960, head_count=15, head_count_kv=5, rope.freq_base=100000 — matches the HF config.json exactly (max_position_embeddings=8192, num_hidden_layers=32, hidden_size=960, num_attention_heads=15, num_key_value_heads=5, rope_theta=100000). Zero nextn/mtp/eagle/medusa/draft tensor names among all 290 — plain decode, no speculation path, consistent with a 361.8M model having no reason to carry one. /v1/models on a live load independently confirms n_params=361,821,120 and size=723,767,040 bytes, both matching the HF card's stated 361.8M / 724 MB safetensors figure.
- 2026-08-17 · acquired → screened
SCREENED on aihydra under the house FIT/GUARD/throughput-matrix tier (build 3653e6d/3653e6d6d, both stock). FIT: loads cleanly on both ROCm and Vulkan at the candidate's full native context (8192 — HF max_position_embeddings, confirmed via GGUF header and a live /v1/models read); no larger depth is meaningful since it exceeds the model's own trained positions. Two-point KV probe with real /sys/class/drm/card0/device/mem_info_gtt_used readings (c=2048: 941.92 MiB delta; c=8192: 1181.98 MiB delta) gives a measured KV cost of exactly 40.00 KiB/token, reproduced from first principles against the GGUF's own attention metadata (2 x 32 layers x 5 KV heads x 64 head_dim x 2 bytes f16). Interactive probing before any scripted run confirmed the model's advertised behavioural contract directly: every prompt tried (a factual question, an instruction to reply with an exact string, a tool-call request) produced a well-formed, on-topic <think> block of invented reasoning followed by a fixed "42" — except the tool-call probe, where it abandoned even that pattern and answered in free text with no tool_calls field at all, HTTP 200, tools array silently ignored. GUARD (bench/guard.sh, depth reduced from the house default 8000 to 6000 — no headroom above that inside an 8192-token native context): 2/4 — coherence PASS (well-formed prose, not garbled; the literal "reply with exactly X" instruction test cannot be diagnostic here since this model is fine-tuned to never comply, so coherence was judged on prose well-formedness instead, documented as a deliberate protocol adaptation, not a lowered bar), native tool call FAIL (no tool_calls structure produced at all), retrieval-at-depth FAIL (needle lost — consistent with a model whose "reasoning" never actually depends on its input), isolation SKIPPED (--parallel 1). Throughput matrix (pp1024/tg256, N=3 fresh-process reps, d0 and d4096 — d32768 does not fit inside this candidate's 8192-token native context and was not attempted): ROCm leads prefill decisively at both depths (16740.8/9276.9 t/s vs Vulkan's 12914.1/6948.3), Vulkan leads decode at both depths (245.2/193.4 t/s vs ROCm's 187.5/164.9) — a real, measured backend split, all cells clean under 3% scatter (max CV 2.56%).
evidence: clm-0077
- 2026-08-17 · screened → benched
FULL 26-task tau2 airline run (run-0372), pinned simulator openrouter/anthropic/claude-haiku-4.5, seed 42, temperature 0, max_tokens 4096, serving config cfg-0101 (ROCm, c=8192) — the same config the guard (run-0363) already found 2/4 with zero native tool calls. RESULT CONFIRMS THE PREDICTION: tool_call_messages = 0 across all 26 simulations (233 assistant messages, 0 empty). Per the SMOKE validity gate (protocol.json comparison_rules), the resulting mean is INVALID — not low, not a fail, INVALID — and that verdict, not a numeric score, is the headline of this bench. Of the 26 requested tasks, 12 terminated as infrastructure_error before any score existed (10 from tau2's own "AssistantMessage must have either content or tool_calls" validation, where the --reasoning-format deepseek split left nothing in the visible content field on some turns; 2 from the accumulated conversation exceeding this candidate's own 8192-token native context — its own verbose invented reasoning crowding out its context window mid-conversation). Of the 14 tasks that did score, 5 landed reward 1.0 — entirely via the airline domain's untouched-database point plus a vacuous COMMUNICATE point, available to any agent that never acts, exactly the mechanism the SMOKE gate (added 2026-08-14 from the npu-lfm2 screen) exists to catch. Wall-metered energy: 22.7285 Wh over the 1496s run window (eng-0146); a Wh-per-correct-answer figure is arithmetically computable (4.5457 Wh, since 5 of 26 tasks scored 1.0 — not the literal zero-correct-answers case) but is published with a hard caveat and excluded from every honest efficiency comparison on this site, because the denominator behind it is exactly the vacuous reward the SMOKE gate flags as invalid, not evidence of a completed task. This candidate's negative-control role is now fully discharged: every house instrument (guard, tool-calling, the SMOKE validity gate, Wh-per-correct-answer) was exercised against a model built never to attempt its task, and each one failed safe — no instrument silently reported a false capability.
- note
Fictional characters Lunkwill and Fook (commissioners of the model card's own fictional "Deep Thought" supercomputer, per Douglas Adams' Hitchhiker's Guide) are named in the source model card as a framing device; they are not real people and are not the lab's own attribution, so their appearance in prose is compatible with the site's no-personal-names editorial rule.
Laguna S 2.1 (poolside)
Terminal-Bench 70.2% against Qwen's 41.6% - the largest claimed agentic gap of any candidate. Hybrid global+SWA MoE with 36/48 SWA layers means a very small KV, and Q4_K_M at 75.2 GB fits with MORE headroom than the current 122B.
gate history (7)
- 2026-07-21 · listed
Released 2026-07-21 with a Terminal-Bench number far above our incumbent, at a footprint that fits.
- 2026-08-10 · listed → acquired
Q4_K_M downloading. Acquisition costs nothing while the GPU is busy, even though it cannot run yet.
- 2026-08-10 · acquired → blocked
Needs llama.cpp #25165 for upstream support; fork-only until it merges. Its SWA class also shares the #25913 checkpoint problem, where our sidecar fix applies. Available on OpenRouter at $0.10/$0.20 per M for feel-testing in the meantime.
- 2026-08-10 · blocked → blocked
Blocker RESOLVED UPSTREAM but not yet in our deployed binary: PR #25165 merged 2026-08-01 as 1f66c3ce ("Add support for Laguna XS.2 & M.1"), verified NOT an ancestor of our stock 3653e6d build. Under the new minimum-version build policy the requirement is now ON MAIN, not fork-only — and the min-62bf73d binary being built for Muse Glimmer necessarily contains it. Two open questions before un-gating: whether the merged PR's arch coverage includes the S 2.1 variant (title names XS.2 and M.1; our GGUF header reads arch=laguna), and the #25913-class SWA checkpoint behaviour. Tonight's screen on stock will record the expected unknown-architecture failure as evidence; a FIT probe on the new binary decides the variant question.
- 2026-08-10 · blocked → screened
SCREENED ON STOCK — the blocker was already resolved in the deployed binary. FIT ok (61s load), ALL FOUR guard checks pass (needle retrieval at depth 8000 included), smoke 0.80 over 5/5 tasks in 7 minutes with nothing cut. This empirically corrects the previous entry's ancestry check, which failed on an unfetched commit and was misread as "support absent": PR #25165's arch support IS in build 3653e6d, and it covers the S 2.1 variant. Strongest first screen of the five new candidates. Smoke is a smoke test, not a ranking (noise ~0.40, clm-0036).
- 2026-08-16 · screened → benched
Full bench complete. Throughput matrix (both backends, f16 KV, d0/d32768/d131072): ROCm swept every cell clean (300.7/23.68 t/s at d0 down to 116.43/7.19 t/s at d131072 pp/tg); Vulkan completed d0 and d32768 (256.9/13.29, 51.6/12.27 t/s, ROCm ahead on both phases at both depths) and lost the GPU device on both allowed attempts at d131072 (rc=134, vk::DeviceLostError, kernel-evidenced amdgpu ring resets, independently re-pulled this session via sudo dmesg -T after finding the job's own dmesg capture file empty from a sudo bug) -- ROCm is the required backend at any useful depth, and unlike nemotron3-super/deepseek-v4-flash (both Vulkan-favoured on this box) it is not a close call here [clm-0065]. KV cost ~48.0 KiB/token from a two-point GTT probe -- published at LOWERED confidence because the underlying sysfs readings were not recoverable from any surviving artifact this session, only from prose (both this session's and an earlier delegate's) [clm-0066]. Full standard 26-task tau2 airline run (ROCm, f16 KV, c32768, no speculation -- this architecture has no MTP/draft-tensor wiring in llama.cpp): completed in full (rc=0, not wall-bound) at 0.6923 mean reward (18/26), 187 tool-call messages, 0 empty turns -- VALID. Energy: 7.2588 Wh per correct answer whole-session (0.2199 pence at 30.3 p/kWh) -- the lowest Wh/correct of any full standard-26-task tau2 comparison on this box to date, screened against every other model page's energy record before claiming it [clm-0067]. Model page published with page: true.
- 2026-08-17 · benched → benched
Depth-ladder extension (Part 2 deep cells): ROCm d204800 (cfg-0109, this candidate's required backend -- Vulkan lost the GPU device at d131072, clm-0065) -- 81.95 pp / 4.91 tg t/s, the longest single throughput cell run this session (~28.5 min), clean completion, no OOM, no device issue. Energy joined live: eng-0156, 81.039 Wh over the cell (mean 170.72 W). d-max (1,048,576, this model's declared context ceiling) was NOT attempted -- recorded as a FIT-PROJECTION REFUSAL per protocol.json's model_max_context_note ("a refused d<model-max> attempt because the fit projection says no is a recorded result, not a silent skip"), not a silent gap. Projection: model weights 96,028,095,488 bytes = 89.41 GiB (llama-bench's own model_size field, cfg-0090) + KV at this candidate's measured ~48.0 KiB/token (lowered-confidence figure, clm-0066) x 1,048,576 tokens = exactly 48.0 GiB (1,048,576 = 1024^2, so the KiB-to-GiB conversion is exact) = 137.41 GiB projected, before any compute- buffer/graph overhead, against this box's 122,880 MiB (120.0 GiB) GTT boot budget (amdgpu.gttsize=122880) -- ~17.4 GiB (14.5%) over budget on weights+KV alone. This is the same refusal class the job brief names for this exact candidate/depth pair.
- note
Vision audit (2026-08-10): parent is text-only - poolside's own materials describe Laguna S 2.1 as a coding/agentic model with no vision capability in the line. Moot for this candidate while the #25165 architecture blocker stands anyway, but worth recording so a future revisit doesn't have to re-derive it: no vision path exists to unlock even once the arch lands upstream.
Qwen3-Coder-Next
Direct request (a friend of the operator's), PRIORITY. HF API confirmed the real identity before acquisition: base model Qwen/Qwen3-Coder-Next (config.json architectures: ["Qwen3NextForCausalLM"], model_type qwen3_next, num_experts 512, num_experts_per_tok 10, num_hidden_layers 48, full_attention_interval 4 -- i.e. 12 of 48 blocks carry full attention (GQA head_dim 256, num_attention_heads 16, num_key_value_heads 2) and the other 36 are linear-attention/SSM (Gated DeltaNet class), same hybrid shape family as this lab's already-benched ornith-35b and the production 122B, just deeper (48 vs 40 layers) and far larger (80B total vs 35B). A coder-tuned Qwen3-Next variant, 256k native context (max_position_embeddings 262144), Apache-2.0. Official GGUF ships at Qwen/Qwen3-Coder-Next-GGUF with a Q8_0 quant that this job's own fit-projection puts comfortably inside the 120 GiB GTT window even at full model-max context (see acquired-gate entry).
gate history (4)
- 2026-08-17 · listed
Direct friend-request, PRIORITY per operator. HF API confirmed both the base model (Qwen/Qwen3-Coder-Next, qwen3_next arch, 512 experts/10 active+1 shared, 48 layers, full_attention_interval 4) and the official GGUF repo (Qwen/Qwen3-Coder-Next-GGUF, Q8_0 shipped as 4 shards, 84,812,055,968 bytes total = 78.99 GiB) exist and carry the intended artifacts. Fit-projected against this box's 120 GiB (128,849,018,880 byte) GTT window before download: per-token KV cost from the config's own full-attention-layer count (12 full- attention blocks x 2 KV heads x 256 head_dim x 2(K+V) x 2 bytes(f16) = 24,576 bytes/token = 24 KiB/token) projects to just 6.0 GiB of KV at this model's own 262,144-token max -- Q8_0 weights (78.99 GiB) + full-model-max KV (6.0 GiB) + overhead is comfortably inside 120 GiB with >30 GiB headroom to spare, so Q8_0 was chosen as the PRIMARY quant per the job brief ("Q8_0 if published and it fits ... a coder model benefits from higher precision") rather than a UD-Q4 fallback. qwen3next architecture confirmed already present in this box's stock fleet-baseline build (llama.cpp commit 3653e6d6d, src/models/qwen3next.cpp exists) -- no fresh build required.
- 2026-08-17 · listed → acquired
Downloaded to aihydra ~/models/qwen3-coder-next-80b/ (official Qwen/Qwen3-Coder-Next-GGUF, Q8_0, 4 shards, 84,812,055,968 bytes total) and sha256-verified exact against HF's own X-Linked-ETag on all 4 shards (shard1 30b7554fc0c846a5dc3ecf585884c77471f73e3da698a8ba4fabd8e7868c6533, shard2 3f96379de5a5c4655cb378710ea571d5e9cc96f260120a44a6477198efcdc27d, shard3 5dd1ce07eaae95ee430331dc9c6f3120ff88e4211ad3a0cceeaa963f25328504, shard4 76730702c630bf76305139165cb85421858604030851dfd64fe96a5e67cda99d). MTP/EOS-CLIFF WATCH resolved: this job's own from-scratch GGUF header parser (gguf_probe2.py) scanned all 807 tensors across all 4 shards (needle scan for nextn/mtp/eagle/medusa/draft) -- ZERO hits on every needle, on every shard. Genuinely no speculation path; matches the config.json listing-gate finding. Architecture independently confirmed from the GGUF's own metadata: qwen3next, 48 blocks, 512 experts/10 active + 1 shared, full_attention_interval=4 (blocks 3/7/11/15/19/23/27/31/35/39/43/47 carry attn_* tensors -- 12 of 48 -- the rest carry ssm_* tensors only, gated-DeltaNet-class linear attention, no growing KV cache). Mirroring to NAS gguf-library (/volume1/Models/gguf-library/ qwen3-coder-next-80b/) via the house scp -O -l 245760 pattern, running concurrently with screening (network-bound, does not contend with GPU work); MANIFEST.tsv entries added.
- 2026-08-17 · acquired → screened
House standard screen, stock 3653e6d ROCm, Q8_0, -np 1, c=32768. PROTOCOL_OVERRIDE carried on every run: Q8_0 is not in protocol.json quants.expected (fit-projected comfortably, same override rationale class as ornith-35b/qwen38-27b). FIT: loaded in 24s, 81,662 MiB GTT (comfortable against the 122,880 MiB boot window, matching the pre-download fit projection almost exactly). GUARD: 4/4 (coherence, tool call, needle at 8000 tokens, isolation skipped under -np 1). SMOKE: 5-task tau2 airline, pinned openrouter/anthropic/claude-haiku-4.5 simulator, temp 0, seed 42 -- 4/5, mean reward 0.800, 39 tool-call messages, 0 empty assistant turns, 0 infra errors (VALID per the SMOKE gate; task 1 the only miss, user_stop after 7 tool calls/38.9s -- no shared pattern with the 4 passes visible from this screen's summary alone). Not yet run: throughput/depth matrix, full 26-task tau2, any MTP/speculative arm (none exists in this GGUF per the acquired-gate parse). energy: null on the smoke run -- to be joined in the same session as the fullbench arms per protocol §11 (HA history retention ~10 days).
- 2026-08-17 · screened → benched
Full bench per house pattern: throughput matrix both backends (ROCm 3653e6d / Vulkan 3653e6d6d), full five-depth ladder (0/32768/131072/204800/262144 model-max) per protocol.json's throughput_ladder, N=3 fresh-process reps <=32K and N=1 past. Vulkan wins decode at every depth and prefill at d0/d32768; ROCm overtakes prefill at the two deepest cells (d204800, d262144) -- a genuine new inversion this programme has not shown on a candidate that ALSO survived Vulkan cleanly through model-max (zero device-loss across all 10 matrix cells, extending ornith-35b's precedent-breaking streak to this programme's largest candidate yet) [clm-0079]. KV cache cost derived from the GGUF's own attention-layer metadata (12 of 48 blocks full-attention): 24 KiB/token (f16 KV), cross-checked against the screen's own FIT probe; projects to ~6 GiB at this candidate's 262,144-token model-max against the 120 GiB GTT window, never a fit risk [clm-0080]. GGUF header parsed directly across all 4 shards (gguf_probe2.py, no numpy/gguf-py dependency): zero nextn/mtp/ eagle/medusa/draft tensors among all 807 -- no MTP/speculation arm run, plain decode throughout [clm-0081]. House guard re-run against the Vulkan tau2 serving config (cfg-0111, run-0395, 4/4, since a backend swap is itself the `binary` lever the guard exists to catch) before the full 26-task tau2 airline run: 0.5385 mean reward (14/26), 192 tool-call messages, 0 empty turns, 0 infra errors -- VALID, and the LOWEST of this programme's full-26-task tau2 comparisons, below qwen38-27b's 0.577 -- expected: this is a coder model scored on tau2's airline domain purely for cross-model comparability, not its home ground (run-0433) [clm-0082]. Energy: 4.7896 Wh per correct answer -- the NEW LOWEST of any full-26-task comparison on this box (ahead of ornith-35b's previous-leading 6.4834 Wh), meaning this candidate holds BOTH the lowest capability score AND the lowest energy-per-correct-answer of this programme's full-bench series simultaneously -- the inverse of ornith-35b's joint-highest result (eng-0168, clm-0082). PROTOCOL_OVERRIDE carried on every matrix/guard/tau2 run (Q8_0 not in protocol.json quants.expected, same rationale as the screen). Guard chain wired from birth on every performance run (both backends) per the house guard-chain convention. Not run: MTP arm (none exists in this GGUF), any coding-domain capability instrument (out of scope for this bench's cross-model comparability protocol).
- note
MTP/EOS-CLIFF WATCH: the base model's config.json carries no nextn/mtp field of any kind (standard Qwen3NextForCausalLM causal-LM config, no draft-head keys) -- provisionally NO speculation path, to be independently confirmed post-download via this lab's own from-scratch GGUF header parser (gguf_probe2.py) per house convention (we were burned once by a mislabeled non-MTP artifact) before any spec-decode arm is attempted or ruled out.
Qwen3.6-27B-MTP
A dense architecture — the case where MTP's larger dense speed-up (1.4-2x, against 1.15-1.25x for MoE) applies. Vendor claims 77.2 SWE-bench at 27B, near Sonnet 4.5 coding. Gated DeltaNet linear attention for cheap long context.
gate history (7)
- 2026-06-20 · listed
Dense architecture + dense-MTP multiplier + a strong vendor coding number.
- 2026-06-27 · listed → acquired
Q4_K_M, 16 GB, on aihydra now.
- 2026-06-27 · acquired → screened
Phase A: 16-22s with MTP active, 100% tool-call gate. MTP works but does not rescue latency - reasoning tokens dominate.
- 2026-08-10 · screened → acquired
REGRESSED to acquired, not dropped. It has zero rows in the aihydra matrix and I lost track of it entirely, describing it to the operator as an accidental find. Its Phase A screen was on the OLD box under a different harness, so it needs re-screening here rather than inheriting a stale pass.
- 2026-08-13 · acquired → screened
RE-SCREENED on aihydra under the house FIT/GUARD/SMOKE tier (build 3653e6d, Aug-12 protocol: pinned simulator openrouter/anthropic/claude-haiku-4.5). FIT ok, loaded in 16s at 32K. GUARD 4/4 PASS — coherence, toolcall (get_booking {"reference":"ABC123"}), needle at depth 8000 (tok=5368 'chartreuse-viper-88' retrieved), isolation (skipped by design at --parallel 1; #25992 only manifests above 1). SMOKE 5-task tau2 airline: completed=5/5, mean=0.80, cut_at_max_steps=0, 11 min wall. Equal to laguna-s21's 0.80 — the joint-strongest smoke on this box — but a smoke test is a cliff-detector, not a ranking (noise ~0.40, clm-0036). Run-meta: screening/qwen36-27b-mtp 2026-08-13T02:54:10Z->03:06:26Z (736s, rc=0).
- 2026-08-16 · screened → benched
FULL BENCH on aihydra (build 3653e6d/3653e6d6d). TWO ARCHITECTURE CLAIMS IN THIS RECORD'S OWN why_listed/name WERE WRONG, caught by this job's own from-scratch GGUF header parse (851 tensors, all metadata) AND confirmed empirically by a live --spec-type draft-mtp load attempt on BOTH backends: (1) NOT dense-uniform-attention -- general.architecture=qwen35, qwen35.full_attention_interval=4, 16 of 64 blocks carry full attention (head_count_kv=4, key/value length 256), the other 48 carry ssm_a/ssm_alpha/ssm_beta/ssm_conv1d/ssm_dt.bias/ssm_norm/ssm_out -- a HYBRID architecture, same shape class as ornith-35b, not the uniform-dense case the why_listed rationale argued MTP would help most. FFN has no "_exps" tensors, so "27B dense" (26.9B params) is still correct about MoE routing, just not about attention uniformity. (2) ZERO MTP TENSORS: exhaustive needle scan found no nextn/mtp/eagle/medusa/draft tensor among all 851; a live llama-server load with --spec-type draft-mtp --spec-draft-n-max 3 refuses to start on BOTH ROCm and Vulkan with an identical error -- "context type MTP requested but model doesn't contain MTP layers" / "failed to create MTP context". The gguf's own general.base_model.0.repo_url is https://huggingface.co/Qwen/Qwen3.6-27B (no -MTP suffix) -- this staged artifact (unsloth/Qwen3.6-27B-GGUF) is the BASE checkpoint, not the MTP-head variant the candidate name assumes. No spec arms exist to measure; none are in this bench. THROUGHPUT MATRIX (d0/d32768/d131072, N=3 at d0/d32768, N=1 at d131072, explicit f16 KV, -fa 1, --load-mode none): ROCm wins prefill decisively at every depth (357.8 vs 302.7 t/s at d0 [+18%]; 210.5 vs 95.7 t/s at d32768 [+120%]) and is the ONLY backend that completed d131072 -- Vulkan device-lost on BOTH allowed attempts at that cell (vk::DeviceLost, GPU wedged, recovered via kernel ring reset each time; ROCm completed cleanly at 95.4 pp / 8.44 tg t/s). Decode is roughly matched, Vulkan slightly ahead (12.83 vs 12.08 t/s at d0, 11.33 vs 10.86 at d32768). ROCm is this bench's serving-config pick on both grounds. KV COST: two-point GTT probe (c=4096, c=131072, real /sys gtt_used readings) gives EXACTLY 64.00 KiB/token -- reconciles to the bit against architecture math (head_count_kv=4 * (256+256) * 2 bytes/f16 * 16 full-attention layers = 65,536 bytes = 64 KiB) and the c=4096 base reading sits within 4 MB of the GGUF's own 15.658 GiB file size. HOUSE GUARD: 4/4 PASS on both the throughput-matrix serving config and (re-run fresh, per protocol) the tau2 serving config -- coherence, non-empty tool call, needle retrieval at depth 8000, isolation skipped at --parallel 1. No cliff detected on either check. FULL 26-TASK TAU2: this candidate could NOT complete the standard run inside the 8h safety ceiling in one session -- rc=124 at 21/26 task ATTEMPTS (20 scored, one task an unrecoverable infrastructure_error with zero messages: "AssistantMessage must have either content or tool_calls", a malformed-output class this bench hit repeatedly across tasks, each occurrence costing several harness-level retries and multiplying wall time). A CONTINUATION session (same build/config/seed, house guard re-run fresh, protocol §9's safe "split across sessions" lever) completed the remaining 5 tasks (21-25) cleanly. COMBINED: 25 of 26 tasks scored (task 15 excluded, infrastructure_error), 16 passed, mean 0.64 -- ABOVE qwen38-27b's 0.577 but well below deepseek-v4-flash (0.846), nemotron3-super (0.769), laguna-s-21 (0.6923) and ornith-35b (0.8846). This is NOT a clean full-26-task completion like every other benched candidate on this board and carries that caveat on every comparison (protocol §9, "bound-limited arms reach different depths/comparing different prefixes compares different work" -- here extended to the one unscored infrastructure_error task). ENERGY IS THE DECISIVE FINDING: total tau2 wall time across both sessions was 40,836s (11.34h) against ornith-35b's 4,193s (70 min) for the SAME 26-task set -- roughly 10x longer, driven by this model's raw ~11-12 t/s decode floor (a third of ornith's ~46 t/s) COMPOUNDED by the repeated malformed-output retries. Combined energy: 2055.19 Wh across both tau2 windows (whole-session, same method as every other board entry) = 128.45 Wh per correct answer -- more than 3x qwen38-27b's previous-worst 38.29 Wh/correct, and roughly 20x ornith-35b's leading 6.48 Wh/correct. GUARD PASSING (no cliff at the cheap-probe tier) did not predict this -- the guard clears binary/hazard levers in seconds and was never designed to catch a real-workload throughput collapse, which is exactly why the full performance sweep and capability suite exist as separate, expensive tiers (protocol §1b). Verdict: SCREENED->BENCHED per house policy (benched records the tier was completed, not that the candidate is recommended) -- this candidate is the current WORST full-bench candidate on this board by Wh-per-correct-answer, despite a passing house guard and a directionally-correct SMOKE screen. See the model page for the full writeup.
- 2026-08-17 · benched → benched
Depth-ladder backfill (protocol extension, same date): ROCm d204800 added -- 68.46 pp / 7.22 tg t/s, guard reused from run-0359 (identical fingerprint: build 3653e6d, ROCm, f16 KV, -fa 1, --load-mode none -- only depth differs, not a fingerprint field). Monotonic with the existing d0/d32768/d131072 series (357.8/210.5/95.4 pp t/s falling to 68.46 at d204800; 12.08/10.86/8.44 tg falling to 7.22) -- no cliff, no scatter anomaly. d262144 (model-max) was ATTEMPTED but killed unfinished after ~43 minutes to preserve session wrap-up time inside a hard overnight deadline; not recorded as a result, and not a cliff signal -- d204800 alone took 1875s (31 min), so d262144 running past 43 min without finishing is consistent with the depth scaling already visible, not evidence of a new failure. Vulkan not attempted at either new depth: it already device-lost at BOTH d131072 attempts, below the "previously survived its prior max tested depth" bar this backfill round set for adding Vulkan cells. Energy: not joined before session release (recorded explicitly per run, not left silent) -- HA history retention ~10 days, join before 2026-08-27.
- note
Vision audit (2026-08-10): parent IS vision-capable - Qwen's own materials describe Qwen3.6-27B as natively multimodal in a single unified checkpoint (same as Qwen3.6-35B-A3B), handling images and video alongside text. Confirmed on our staged artifact: Qwen3.6-27B-Q4_K_M.gguf reads general.architecture=qwen35 (llama.cpp reuses the 3.5-era arch code for 3.6), carries the "image-text-to-text" tag, and has qwen35.rope.dimension_sections=[11,11,10,0] - MRoPE sectioning implemented natively in src/models/qwen35.cpp on our build, not requiring a separate VL arch. unsloth/Qwen3.6-27B-GGUF (source repo for our download) publishes mmproj-F16.gguf (927,607,360 bytes), not staged. Runtime: LLM_ARCH_QWEN35 is registered in our build (commit 3653e6d), and clip.cpp implements PROJECTOR_TYPE_QWEN3VL generically on the LLM side. Verdict: mmproj-download-away, not text-only - and notably this is a DENSE vision-capable candidate, worth weighing against the MoE Qwen3.5/3.6 vision candidates on that basis alone.
gpt-oss-120b
The migration's planned daily-driver upgrade and the speed bar for the whole field at ~55 t/s Vulkan.
gate history (4)
- 2026-06-20 · listed
Listed as the planned daily-driver upgrade and as the speed control the rest of the field is measured against.
- 2026-06-27 · listed → acquired
Downloaded/staged at UD-Q4_K_XL for the Phase 3C field. Later record trace confirmed cfg-0016/cfg-0017/cfg-0018 used the staged UD-Q4_K_XL artifact; older Q4_K_M copies remain archive-only and are not the published bench artifact.
- 2026-06-27 · acquired → screened
Phase A: fastest in the field at ~5s time-to-correct, 100% tool-call gate.
- 2026-08-08 · screened → benched
Throughput rows recorded. NOT cleared for use: returns verbatim answers from previous requests at --parallel 1, reproduced 6/6, and scored 0.00 on tau2 thinking-off. Kept at bench rather than rejected because the staleness mechanism is undiagnosed and may be a harness fault rather than the model.
evidence: clm-0025
- note
Vision audit (2026-08-10): parent is text-only by OpenAI's own design and documentation - gpt-oss ships with no vision encoder in the line, and OpenAI explicitly points multimodal use cases at their hosted API instead. Staged artifact (arch gpt-oss) carries no vision tensors. Verdict: text-only, not a vision-lane candidate - and separately still not cleared for use per the staleness finding above.
Qwen3.5-122B-A10B (MTP)
Production incumbent and the quality ceiling that still fits. 262K context; MTP measured adopt-worthy at +50% decode with a ~10% cold-path prefill tax.
gate history (4)
- 2026-06-20 · listed
Selected as the quality-ceiling candidate that fits 96GB at Q4.
- 2026-06-27 · listed → acquired
Downloaded UD-Q4_K_M for the Phase 3C field.
- 2026-06-27 · acquired → screened
Phase A agentic eval: ~11s time-to-correct, 100% tool-call gate. Matches the 35B despite 3.5x the params, so latency is reasoning-bound not decode-bound.
- 2026-08-08 · screened → benched
Fully characterised on performance: depth to 204.8k, KV quant both ways, FA both ways, speculation curve, cache trace, grammar ceiling.
- note
Vision audit (2026-08-10): parent IS vision-capable, and not as an afterthought - Qwen bills Qwen3.5-122B-A10B as natively multimodal (early-fusion trained on text+image+video tokens), and this is confirmed directly on our staged artifact: Qwen3.5-122B-A10B-UD-Q4_K_M reads general.architecture=qwen35moe, carries the "image-text-to-text" tag, and has qwen35moe.rope.dimension_sections=[11,11,10,0] - the same MRoPE sectioning llama.cpp's dedicated qwen3vl/qwen3vlmoe arches use for spatial image positions, implemented natively inside src/models/qwen35moe.cpp on our build rather than requiring a separate VL arch tag. unsloth's Qwen3.5-122B-A10B-GGUF (the repo we pulled the main weights from) publishes a matching mmproj-F16.gguf (908,724,960 bytes) that we have not staged. Runtime: LLM_ARCH_QWEN35MOE is registered in our build (commit 3653e6d) and clip.cpp implements PROJECTOR_TYPE_QWEN3VL, which is arch-agnostic on the LLM side. Verdict: mmproj-download-away, not text-only - this is the strongest vision-lane candidate already on the box, and it is our current production incumbent. Caveat: an open upstream issue (#21268, CUDA-only in its report) describes CLIP-graph operator gaps with this exact model+mmproj pairing causing OOM rather than a clean load - unverified on ROCm/Vulkan, so budget a debug pass before treating this as a guaranteed win.
Out of the pipeline — every exit carries its evidence
blocked (0)
none
rejected (1)
Gemma 4 26B-A4B
rejected on: clm-0102 con-0005
The efficient fast-daily candidate: ~85 t/s at near-31B quality. If it holds up it is a far cheaper reflex tier than anything currently deployed.
gate history (5)
- 2026-06-20 · listed
Speed/quality ratio is the best claimed in the field for a reflex tier.
- 2026-08-10 · listed → acquired
UD-Q4_K_M, 16 GB, downloaded. Tool-call format needs checking - flagged at listing time as the known risk.
- 2026-08-10 · acquired → screened
Screened: FIT ok (13s), coherence/toolcall/isolation pass, needle retrieval FAILED at depth 8000. Smoke 0.75 over 4 scored tasks (one task did not score — unexamined). Reflex-tier speed claim not yet measured; the needle failure caps usable depth pending a retrieval-depth sweep.
- 2026-08-18 · screened → rejected
HG-001 bounded promotion retried the shallower retrieval-envelope check and stopped at the first declared depth: house guard 3/4, with coherence and native tool-call arguments passing but needle retrieval lost at 2048 tokens. Per the card stop gate, no tau2 smoke, throughput, Vulkan arm or energy measurement was admitted. Existing evidence rejects the reflex-tier promotion path; it does not isolate a broad model-quality failure from the retrieval/probe-specific failure mode.
- 2026-08-21 · rejected → rejected
Upstream #26750 draft-mtp defect addendum added as community report evidence, not a new HaloBench run and not a gate change (remains rejected per the HG-001 2026-08-18 verdict). This candidate's MTP variant (mtp-gemma-4-26B-A4B-it) is the draft model in the #26750 chain; the 2026-08-21 zanphear follow-up isolates the acceptance-rate collapse as spec-path-specific (prose 66.67%->52.38% across b10261 to b10532 on GB10 CUDA, while the same builds are byte-identical in raw no-spec prefill/decode). Recorded for MTP-correctness provenance only: it runs on CUDA/GB10, not our ROCm/Vulkan surface, so it does not change this candidate's or any model's measured serving result.
- note
Vision audit (2026-08-10): parent is multimodal (text+image, all Gemma 4 sizes) - confirmed on our exact staged artifact too: the downloaded gemma-4-26B-A4B-it-UD-Q4_K_M.gguf reads general.architecture=gemma4 and carries the "image-text-to-text" tag. Our llama.cpp build (commit 3653e6d) registers LLM_ARCH_GEMMA4 and clip.cpp implements PROJECTOR_TYPE_GEMMA4V for this size class - runtime support is real, not aspirational. What's missing is the mmproj file: unsloth/gemma-4-26B-A4B-it-GGUF (the repo we pulled the main weights from) publishes mmproj-F16.gguf (1,193,058,784 bytes) alongside it, and we have not staged it. One download away from vision-ready. Caution: a live upstream issue (#21402) reports a SIGABRT in clip_model_loader::load_tensors when loading the Gemma4 mmproj on CUDA - unverified whether that reproduces on our ROCm/Vulkan build, so budget a debug pass rather than assume a clean load on first try.
generated from the candidates collection · a gate change without a reason cannot be recorded (schema-enforced) · a rejection without evidence fails the build · cards sitting over 30 days in acquired/screened carry the amber dot — work-queue honesty without publishing a work queue