Models › qwen36-27b-mtp
⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.

Qwen3.6-27B (staged as "qwen36-27b-mtp" — the artifact has no MTP path)benchedguard 4/4

27B densehybrid (qwen35) · 16/64 full-attention + 48/64 SSM/linear-attention layers · 26.9B dense (no MoE) — NO MTP/speculation tensors in this GGUFquant held: Q4_K_M · house-standard quant (protocol.json quants.expected)first measured 2026-06-27latest run 2026-08-17

②Verdict

A hybrid architecture (qwen35, 16 of 64 blocks full-attention via full_attention_interval=4, the other 48 SSM/linear-attention with no growing KV cache — same shape class as ornith-35b and the production 122B) served at Q4_K_M, the house-standard quant — with NO speculation path at all: this job's own from-scratch GGUF header parser found zero nextn/mtp/eagle/medusa/draft tensors among all 851 tensors, and a live --spec-type draft-mtp load attempt refuses to start on BOTH backends with an identical error clm-0073. This is a real correction to the candidate's own record, which assumed uniform dense attention specifically because it would benefit most from an MTP path that does not exist in this staged artifact — the gguf's own base_model metadata points at plain Qwen/Qwen3.6-27B, not an MTP-head variant. Backend is a genuine finding: ROCm wins prefill decisively at every depth (357.8/210.5/95.4 t/s at d0/d32768/d131072 vs Vulkan's 302.7/95.7/—) and is the only backend that survives this candidate's deepest matrix cell — Vulkan device-lost on both allowed attempts at d131072, matching the laguna-s-21/ deepseek-v4-flash precedent this silicon has shown before clm-0074. The KV cache costs a measured 64.00 KiB/token — the LARGEST figure among this board's hybrid full-bench candidates, reconciled two independent ways (file-size match to within 3.8 MB, and an exact match to architecture math from the measured layer count and head count) clm-0075. Capability clears the house guard cleanly (4/4, twice, no cliff on either check) and the standard 26-task tau2 airline set — but only across TWO sessions: the first hit the 8h safety ceiling at 21 of 26 task attempts and a same-build continuation was needed to finish the remaining five. Combined: 25 of 26 scored (one task an unrecoverable infrastructure_error, not a 0), 16 passed, mean 0.6400 — clearing qwen38-27b's 0.577 floor but trailing every other full-bench candidate on this board clm-0076. The energy figure is what actually decides this candidate's standing: combined tau2 wall time was 11.34 hours for the SAME 26-task set ornith-35b completed in 70 minutes, driven by a raw ~11-12 t/s decode floor (a third of ornith-35b's ~46 t/s) compounded by repeated harness-level retries from a malformed empty-turn output class this bench measured directly — "AssistantMessage must have either content or tool_calls." Combined energy: 2055.19 Wh across both sessions, 128.45 Wh per correct answer — more than 3x qwen38-27b's 38.29 Wh/correct and roughly 20x ornith-35b's cited 6.48 Wh/correct on the same protocol clm-0076. The headline finding of this bench is not the numbers alone but what produced them: a passing house guard (the cheap cliff-detector, 4/4 twice) and a directionally-consistent SMOKE screen (0.80 mean, 2026-08-13) gave no warning of either the architecture-record error or the real-workload throughput collapse — exactly the gap the guard/performance-sweep/capability-suite tier split exists to catch, and exactly why "screened" and "benched" are different gates on this board rather than one being redundant with the other.

verdict written 2026-08-17 · every number above stands next to the claim chip that carries it

③Best configuration

modelQwen3.6-27B-Q4_K_M.gguf · Q4_K_M
engineggml-org/llama.cpp 3653e6d · rocm · host aihydra (igpu)
flags-ngl 999 -fa on -c 32768 -ctk f16 -ctv f16 --jinja --reasoning-format deepseek -n 4096 --parallel 1 --load-mode none
templatenot recorded at test time
treeupstream — stock

config record cfg-0099

decode @ 0
12.08 t/s
run-0347 · CV 0%
decode @ 32k
10.86 t/s
run-0349 · CV 0%
prefill @ 0
357.84 t/s
run-0346 · CV 0.7%
prefill @ 32k
210.51 t/s
run-0348 · CV 1.6%
τ² airline · thinking off
0.700 ±0.201
run-0358 · passed 14/20
Wh per correct answer (τ², thinking off)
128.45 Wh
eng-0144 run-0358 run-0360 · 3.89 p per answer

every cell generated from the record at build time · throughput cells from cfg-0097 (same build 3653e6d, fa on, f16 KV) · provenance: measured-here throughout

③aDecode against context depth

ROCmVulkan
051015032k65k131k204.8kdecode t/scontext depth (tokens)serving context 32k12.08 t/s @ depth 0 · run-0347 · CV 0% · N=310.86 t/s @ depth 32k · run-0349 · CV 0% · N=38.44 t/s @ depth 131k · run-0351 · CV 0% · N=112.83 t/s @ depth 0 · run-0353 · CV 0% · N=311.33 t/s @ depth 32k · run-0355 · CV 0% · N=3Vulkan · 11.33ROCm · 8.44
1/3 reps per cell · max CV 0.0% · build 3653e6d / 3653e6d · ROCm completed all three tested depths (0, 32768, 131072); Vulkan completed only the first two — it device-lost on BOTH allowed attempts at d131072 (GPU wedged, recovered via kernel ring reset each time). ROCm wins prefill decisively at every depth Vulkan also reached (+18% at d0, +120% at d32768); decode is close, with Vulkan marginally ahead where it has a number (+6% at d0, +4% at d32768). · records: run-0347 run-0349 run-0351 run-0353 run-0355

④Other configurations tested — each as a delta against best

variantΔ decodeΔ turns (paired tasks)Δ energynoterecords
ROCm vs Vulkan backend · same quant+6.3% @0—-8.5%Vulkan decodes marginally faster at every depth it completed (+6% at d0, +4% at d32768), but ROCm wins prefill by a wide margin (+18% at d0, +120% at d32768) and is the only backend that survives this candidate's deepest matrix cell — Vulkan device-lost on both attempts at d131072. Net pick: ROCm, on completeness and prefill, not decode. clm-0074
Vulkan device-loss — the precedent holds again———Vulkan device-lost on BOTH allowed attempts at d131072 (vk::DeviceLostError, GPU wedged, kernel ring reset recovered it each time) — matching the laguna-s-21 and deepseek-v4-flash precedent at comparable depth on this silicon. ornith-35b broke that streak once; this candidate does not, so the streak-break reads as ornith-specific rather than a general fix. clm-0074
KV cache cost (fit)———64.00 KiB/token (f16 KV, ROCm) — the LARGEST full-bench figure on this board (ornith-35b 20.00, nemotron3-super 8.00, deepseek-v4-flash 7.13), consistent with carrying more full-attention layers (16 of 64) and more KV heads (head_count_kv=4) than any hybrid peer measured here. Reconciled two ways: the c=4096 base reading agrees with the GGUF's own file size to within 3.8 MB, and the per-token slope matches architecture math (head_count_kv × (k_len+v_len) × 2 bytes × 16 full-attention layers = 64 KiB) exactly. Projected GTT at c=131,072 is ~23.4 GiB against the 120 GiB boot window — fit was never this candidate's constraint. clm-0075
Architecture correction — not dense-uniform, no MTP———This candidate's own listing rationale assumed uniform dense attention plus an MTP draft path. Neither is true of the staged artifact: it is a HYBRID architecture (16/64 full-attention blocks, 48/64 SSM/linear-attention), and a live --spec-type draft-mtp load attempt refuses to start on both backends with "model doesn't contain MTP layers" — the gguf is the base Qwen3.6-27B checkpoint, not an MTP-head variant. No spec/MTP arms exist in this bench because there is no speculation path to measure. clm-0073
τ² capability + energy — the worst efficiency on this board———Combined across two sessions (the standard run hit the 8h safety ceiling at 21/26 attempts and needed a same-config continuation to finish): 25 of 26 tasks scored, 16 passed, mean 0.6400 — above qwen38-27b's 0.577 but well below every other cited full-bench candidate here. Energy is the decisive number: 2055.19 Wh combined = 128.45 Wh per correct answer, more than 3x qwen38-27b's 38.29 Wh/correct and ~20x ornith-35b's cited 6.48 Wh/correct — driven by a raw decode floor a third of ornith-35b's and repeated harness-level retries from a malformed empty-turn output class this bench observed directly. clm-0076

deltas computed at build time from the named runs, paired-task traces and metered windows — a hazard row shows words where a delta would mislead

⑤Open questions

⑥Provenance

bench host aihydra · rocm · ggml-org/llama.cpp 3653e6d
discipline 3/1 reps per throughput cell · scatter published per cell (max CV 1.6%) · guard or written waiver on every performance series
window 2026-06-27 → 2026-08-17