Qwen3.6-27B (staged as "qwen36-27b-mtp" — the artifact has no MTP path)benchedguard 4/4
②Verdict
A hybrid architecture (qwen35, 16 of 64 blocks full-attention via full_attention_interval=4, the other 48 SSM/linear-attention with no growing KV cache — same shape class as ornith-35b and the production 122B) served at Q4_K_M, the house-standard quant — with NO speculation path at all: this job's own from-scratch GGUF header parser found zero nextn/mtp/eagle/medusa/draft tensors among all 851 tensors, and a live --spec-type draft-mtp load attempt refuses to start on BOTH backends with an identical error clm-0073. This is a real correction to the candidate's own record, which assumed uniform dense attention specifically because it would benefit most from an MTP path that does not exist in this staged artifact — the gguf's own base_model metadata points at plain Qwen/Qwen3.6-27B, not an MTP-head variant. Backend is a genuine finding: ROCm wins prefill decisively at every depth (357.8/210.5/95.4 t/s at d0/d32768/d131072 vs Vulkan's 302.7/95.7/—) and is the only backend that survives this candidate's deepest matrix cell — Vulkan device-lost on both allowed attempts at d131072, matching the laguna-s-21/ deepseek-v4-flash precedent this silicon has shown before clm-0074. The KV cache costs a measured 64.00 KiB/token — the LARGEST figure among this board's hybrid full-bench candidates, reconciled two independent ways (file-size match to within 3.8 MB, and an exact match to architecture math from the measured layer count and head count) clm-0075. Capability clears the house guard cleanly (4/4, twice, no cliff on either check) and the standard 26-task tau2 airline set — but only across TWO sessions: the first hit the 8h safety ceiling at 21 of 26 task attempts and a same-build continuation was needed to finish the remaining five. Combined: 25 of 26 scored (one task an unrecoverable infrastructure_error, not a 0), 16 passed, mean 0.6400 — clearing qwen38-27b's 0.577 floor but trailing every other full-bench candidate on this board clm-0076. The energy figure is what actually decides this candidate's standing: combined tau2 wall time was 11.34 hours for the SAME 26-task set ornith-35b completed in 70 minutes, driven by a raw ~11-12 t/s decode floor (a third of ornith-35b's ~46 t/s) compounded by repeated harness-level retries from a malformed empty-turn output class this bench measured directly — "AssistantMessage must have either content or tool_calls." Combined energy: 2055.19 Wh across both sessions, 128.45 Wh per correct answer — more than 3x qwen38-27b's 38.29 Wh/correct and roughly 20x ornith-35b's cited 6.48 Wh/correct on the same protocol clm-0076. The headline finding of this bench is not the numbers alone but what produced them: a passing house guard (the cheap cliff-detector, 4/4 twice) and a directionally-consistent SMOKE screen (0.80 mean, 2026-08-13) gave no warning of either the architecture-record error or the real-workload throughput collapse — exactly the gap the guard/performance-sweep/capability-suite tier split exists to catch, and exactly why "screened" and "benched" are different gates on this board rather than one being redundant with the other.
③Best configuration
| model | Qwen3.6-27B-Q4_K_M.gguf · Q4_K_M |
| engine | ggml-org/llama.cpp 3653e6d · rocm · host aihydra (igpu) |
| flags | -ngl 999 -fa on -c 32768 -ctk f16 -ctv f16 --jinja --reasoning-format deepseek -n 4096 --parallel 1 --load-mode none |
| template | not recorded at test time |
| tree | upstream — stock |
config record cfg-0099
every cell generated from the record at build time · throughput cells from cfg-0097 (same build 3653e6d, fa on, f16 KV) · provenance: measured-here throughout
③aDecode against context depth
④Other configurations tested — each as a delta against best
| variant | Δ decode | Δ turns (paired tasks) | Δ energy | note | records |
|---|---|---|---|---|---|
| ROCm vs Vulkan backend · same quant | +6.3% @0 | — | -8.5% | Vulkan decodes marginally faster at every depth it completed (+6% at d0, +4% at d32768), but ROCm wins prefill by a wide margin (+18% at d0, +120% at d32768) and is the only backend that survives this candidate's deepest matrix cell — Vulkan device-lost on both attempts at d131072. Net pick: ROCm, on completeness and prefill, not decode. | clm-0074 |
| Vulkan device-loss — the precedent holds again | — | — | — | Vulkan device-lost on BOTH allowed attempts at d131072 (vk::DeviceLostError, GPU wedged, kernel ring reset recovered it each time) — matching the laguna-s-21 and deepseek-v4-flash precedent at comparable depth on this silicon. ornith-35b broke that streak once; this candidate does not, so the streak-break reads as ornith-specific rather than a general fix. | clm-0074 |
| KV cache cost (fit) | — | — | — | 64.00 KiB/token (f16 KV, ROCm) — the LARGEST full-bench figure on this board (ornith-35b 20.00, nemotron3-super 8.00, deepseek-v4-flash 7.13), consistent with carrying more full-attention layers (16 of 64) and more KV heads (head_count_kv=4) than any hybrid peer measured here. Reconciled two ways: the c=4096 base reading agrees with the GGUF's own file size to within 3.8 MB, and the per-token slope matches architecture math (head_count_kv × (k_len+v_len) × 2 bytes × 16 full-attention layers = 64 KiB) exactly. Projected GTT at c=131,072 is ~23.4 GiB against the 120 GiB boot window — fit was never this candidate's constraint. | clm-0075 |
| Architecture correction — not dense-uniform, no MTP | — | — | — | This candidate's own listing rationale assumed uniform dense attention plus an MTP draft path. Neither is true of the staged artifact: it is a HYBRID architecture (16/64 full-attention blocks, 48/64 SSM/linear-attention), and a live --spec-type draft-mtp load attempt refuses to start on both backends with "model doesn't contain MTP layers" — the gguf is the base Qwen3.6-27B checkpoint, not an MTP-head variant. No spec/MTP arms exist in this bench because there is no speculation path to measure. | clm-0073 |
| τ² capability + energy — the worst efficiency on this board | — | — | — | Combined across two sessions (the standard run hit the 8h safety ceiling at 21/26 attempts and needed a same-config continuation to finish): 25 of 26 tasks scored, 16 passed, mean 0.6400 — above qwen38-27b's 0.577 but well below every other cited full-bench candidate here. Energy is the decisive number: 2055.19 Wh combined = 128.45 Wh per correct answer, more than 3x qwen38-27b's 38.29 Wh/correct and ~20x ornith-35b's cited 6.48 Wh/correct — driven by a raw decode floor a third of ornith-35b's and repeated harness-level retries from a malformed empty-turn output class this bench observed directly. | clm-0076 |
deltas computed at build time from the named runs, paired-task traces and metered windows — a hazard row shows words where a delta would mislead
⑤Open questions
- Whether the malformed-empty-turn output class ("AssistantMessage must have either content or tool_calls") is a template/sampling issue specific to this build and chat template, or an intrinsic property of this checkpoint — it was not isolated or reproduced outside the tau2 harness in this session, and no other full-bench candidate on this board has shown the same signature at this rate.
- Whether an MTP-head variant of Qwen3.6-27B exists as a separate published checkpoint/GGUF this lab could acquire — the vendor's own "MTP" naming and this candidate's original why_listed rationale assumed one exists; this bench only establishes that the STAGED unsloth artifact (base Qwen3.6-27B) does not carry it.
- Whether this candidate's real-workload decode floor (~11-12 t/s) is itself anomalous for a 26.9B dense model on this silicon at this quant, or whether it reflects some other confound in this session's config not yet isolated — no controlled comparison against a same-size dense peer at the same quant exists on this board yet.
- Root cause of the 10 failed/errored tasks (3, 7, 8, 12, 14, 15, 18, 21, 23, 24 — counting task 15's infrastructure_error) — no shared pattern in tool-call count or duration is visible from this bench's own summary; a transcript read would be needed to say more.
- Whether the vision claim from the candidate's own note (image-text-to-text tag, MRoPE sectioning, mmproj-download-away) holds for this specific GGUF — out of scope for this throughput/capability bench and not tested here.