Home › Evidence › Records › clm-0073
⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.

clm-0073

measured-herehigh ●●●
citable URL: https://halobench.com/records/clm-0073/ — this address never moves; the anchor /records/#clm-0073 keeps resolving

The staged qwen36-27b-mtp GGUF (unsloth/Qwen3.6-27B-GGUF, Q4_K_M) is NOT the uniform-dense-attention, MTP-capable model this candidate's own record assumed. A from-scratch GGUF header parse (851 tensors, every metadata KV) found general.architecture=qwen35, qwen35.full_attention_interval=4: 16 of 64 blocks carry standard full-attention tensors (head_count_kv=4, key/value length 256), the other 48 carry ssm_a/ssm_alpha/ssm_beta/ssm_conv1d/ssm_dt.bias/ssm_norm/ssm_out — a HYBRID architecture (same shape class as ornith-35b and the production 122B family), not the uniform dense-attention case the candidate's why_listed rationale argued MTP's larger speed-up would apply to. FFN carries no "_exps" tensors, so "27B dense" (26.9B total params, confirmed by llama-bench's own model_n_params field) remains correct about MoE routing — the correction is about attention uniformity only. Separately: zero MTP/speculation tensors exist in this artifact. An exhaustive needle scan of all 851 tensor names found no nextn/mtp/eagle/medusa/draft tensor anywhere, and this was confirmed empirically, not just by absence: a live llama-server load with --spec-type draft-mtp --spec-draft-n-max 3 refuses to start on BOTH the ROCm and Vulkan binaries with the identical error — "context type MTP requested but model doesn't contain MTP layers" / "failed to create MTP context". The gguf's own general.base_model.0.repo_url metadata field is https://huggingface.co/Qwen/Qwen3.6-27B (no "-MTP" suffix) — this staged artifact is the BASE Qwen3.6-27B checkpoint, not the MTP-head variant the candidate's display name and listing rationale assumed. No spec/MTP arms exist in this bench because there is no speculation path to measure.

verified 2026-08-16 · volatility low

Note — the record's own working

METHOD: gguf_probe2.py, a from-scratch GGUF metadata + tensor-name parser with no numpy/gguf-py dependency, run against ~/models/qwen36-27b/Qwen3.6-27B-Q4_K_M.gguf on aihydra. Every metadata KV dumped unfiltered; all 851 tensor names scanned for ssm/mamba/conv1d/hybrid/linear-attention AND nextn/mtp/draft/eagle/medusa/a_log/dt_bias/x_proj/in_proj/out_proj markers. Per-block tensor schema grouped into exactly 2 classes across all 64 blocks, matching qwen35.full_attention_interval=4 exactly (blocks 3,7,11,15,19,23,27,31,35,39,43,47, 51,55,59,63 = 16 of 64 are full-attention). LIVE CONFIRMATION: a direct llama-server invocation on both binaries (~/src/llama.cpp/build/bin, ROCm, and ~/src/llama.cpp-vk3653e6d/build/bin, Vulkan), identical flags except -dev, with --spec-type draft-mtp --spec-draft-n-max 3 added. Both processes fail to load with the same log line: "llama_init_from_model: context type MTP requested but model doesn't contain MTP layers" / "common_speculative_init_result: failed to create MTP context" / "load_model: failed to create MTP context" — server exits, never becomes healthy. Raw logs preserved: aihydra ~/bench-results/qwen36-27b-mtp-fullbench/ mtp-attempt-rocm.log, mtp-attempt-vulkan.log. This closes a real gap in the candidate's own record: its 2026-06-20 "listed" and 2026-06-27/2026-08-13 "screened" entries all assumed dense-uniform attention plus an MTP path, neither of which the staged artifact actually has. The 2026-08-13 screen's own SMOKE result (0.80 mean, joint-strongest on this box at the time) is unaffected by this correction — it was measured against the real artifact regardless of what the candidate record believed about its architecture — but the "dense architecture — the case where MTP's larger dense speed-up applies" why_listed rationale is now known to not describe what was actually screened.

Cited by — computed at build time, never stored

model pages qwen36-27b-mtp