Qwen3-Coder-Next (Qwen)benchedguard 4/4
②Verdict
A hybrid mixture-of-experts (qwen3next architecture, 12 of 48 blocks carrying full-attention layers via full_attention_interval=4, the other 36 gated- DeltaNet-class SSM/linear-attention with no growing KV cache — the same architecture family as this lab's production 122B and the already-benched ornith-35b, just deeper and roughly 2.3x larger) served at Q8_0 — this candidate's own designated primary quant, chosen because it fit-projected comfortably against the 120 GiB GTT window even at full model-max context, a PROTOCOL_OVERRIDE against protocol.json's quants.expected list on the same rationale ornith-35b and qwen38-27b already established — with no speculation path: this job's own from-scratch GGUF header parser found zero nextn/mtp/ eagle/medusa/draft tensors among all 807 tensors across all 4 shards.
Backend is a genuine, nuanced finding here, not a formality: Vulkan wins decode at every measured depth in its five-depth ladder and prefill at the two shallowest cells, but ROCm overtakes prefill at the two deepest cells (d204800, d262144) — this page's cited case showing Vulkan surviving cleanly through model-max (zero device-loss across all 10 matrix cells, extending the ornith-35b clean-streak comparison to this 80B candidate) while ALSO losing ground on prefill at depth. Decode remains the dominant driver of interactive throughput, so Vulkan still serves the capability arm. KV cost is a measured 24 KiB/token (f16 KV, derived from the GGUF's own attention-layer metadata) — the higher than the cited hybrid comparator figures on this page, but still a rounding error against the 120 GiB window: even this candidate's own 262,144-token model-max projects to roughly 6 GiB of KV, weights alone account for the other 79 GiB.
Capability is where this candidate's story turns: the FULL 26-task tau2 airline run completed end to end (rc=0, not wall-bound) at 0.5385 mean reward (14/26), 192 tool-call messages, 0 empty turns — VALID under the SMOKE gate, but below qwen38-27b's 0.577 on the cited same-protocol airline comparison. This is a coder model — built and marketed for agentic coding, not customer-service dialogue — being scored on tau2's airline domain purely for cross-model comparability with the cited same- protocol candidates; the number says how a coding specialist performs OFF its home ground, not what it can do on the work it is actually built for. Energy inverts the usual pattern: 4.7896 Wh per correct answer, 0.1451 pence at the standing tariff — lower than ornith-35b's cited 6.4834 Wh record on the same standard-26-task tau2 protocol. That pairs a low energy-per-correct-answer measurement with weak airline-domain reward, the inverse of Ornith's high-reward / low-energy record rather than a claim of better capability.
This efficiency figure was screened against the then-published benched models' Wh-per-correct-answer records before publishing (qwen35-122b's 5.713 Wh/correct comes from a 5-task smoke window predating the standard-26-task convention — a different scale, not a like-for-like comparator). The confound worth naming plainly: a lower-reward run that reaches user_stop in fewer effective turns is not the same phenomenon as efficient problem-solving, and this verdict does not claim this candidate is intrinsically more energy-efficient PER UNIT OF CAPABILITY than ornith-35b or any other comparator — only that the measured Wh-per-correct-answer figure, on the identical protocol, is low for the cited comparator set, and that it arrives paired with weak airline-domain reward rather than a higher score.
③Best configuration
| model | Qwen3-Coder-Next-Q8_0-00001-of-00004.gguf · Q8_0 |
| engine | ggml-org/llama.cpp 3653e6d · vulkan · host aihydra (igpu) |
| flags | -ngl 999 -fa on -c 32768 -dev Vulkan0 --jinja --reasoning-format deepseek -n 4096 --parallel 1 --load-mode none --slots |
| template | not recorded at test time |
| tree | upstream — stock |
config record cfg-0111
every cell generated from the record at build time · throughput cells from cfg-0113 (same build 3653e6d, fa on, f16 KV) · provenance: measured-here throughout
③aDecode against context depth
④Other configurations tested — each as a delta against best
| variant | Δ decode | Δ turns (paired tasks) | Δ energy | note | records |
|---|---|---|---|---|---|
| Vulkan backend decode lead · same quant | +17% @0 | — | +29% | Vulkan wins decode at every depth measured (+17% at d0, +19% at d32768, +24% at d131072/d204800/d262144) and prefill at the two shallowest cells (+29% at d0, +30% at d32768) — but ROCm overtakes prefill at the two deepest cells (d204800 +26.2%, d262144 +52.0% in ROCm's favour). Vulkan is the backend this bench serves the capability arm on, matching the house convention of following decode (the dominant driver of interactive throughput) even where prefill splits by depth. | clm-0079 |
| KV cache cost (fit) | — | — | — | 24 KiB/token (f16 KV) — higher than the cited hybrid full-bench comparator figures on this page, consistent with this candidate carrying more full-attention layers than those cited comparators (12 of 48, vs ornith-35b's 10 of 40) even though the fraction of full-attention blocks is similar. Still cheap in absolute terms against this box: even at this candidate's own model-max context (262,144 tokens), KV projects to roughly 6 GiB against the 120 GiB GTT window — weights (79 GiB) dominate the footprint by an order of magnitude, and fit was never a risk at any depth this bench measured. | clm-0080 |
| No speculation / MTP | — | — | — | This job's own from-scratch GGUF header parser scanned all 807 tensors across all 4 shards and found zero nextn/mtp/eagle/medusa/draft tensors — genuinely no speculative-decode path exists in this artifact, corroborated independently by the base model's own config.json before download. Every throughput and capability number in this bench is plain decode. | clm-0081 |
| Wh-per-correct-answer with low airline-domain reward | — | — | — | 4.7896 Wh/correct is a low same-protocol energy figure beside this page's cited full-bench comparators (ornith-35b 6.4834, laguna-s-21 7.2588, deepseek-v4-flash 24.29, nemotron3-super 32.43, qwen38-27b 38.29) — but it arrives with a weak airline-domain capability score (0.5385 mean reward, below qwen38-27b's 0.577). The efficiency figure is measured on the identical protocol as those cited comparator records, but it is not accompanied by a capability lead and should not be read as one. | clm-0082 |
deltas computed at build time from the named runs, paired-task traces and metered windows — a hazard row shows words where a delta would mislead
⑤Open questions
- Root cause of the 12 failed tasks (0, 7, 8, 10, 11, 14, 15, 18, 20, 21, 23, 24) — no shared pattern in tool-call count or duration is visible from this bench's own summary; a transcript read would be needed to say more, and a coding-domain tau2 task set (if one existed) would be a more informative capability read for this specific candidate than the airline domain.
- Whether the efficiency lead survives a controlled energy-per-unit-of- capability comparison against ornith-35b — the raw Wh-per-correct-answer figures are on the same protocol, but this candidate's lower reward plausibly reaches its terminal user_stop in fewer costly turns on the tasks it fails, which is a different mechanism from doing more useful work per watt-hour on tasks it solves.
- This candidate's actual home-domain performance (agentic coding) is unmeasured by this bench entirely — tau2's airline domain was chosen for cross-model comparability, not as a coding-capability instrument; a coding-specific agentic benchmark would be the natural follow-up before any purpose-fit recommendation for this model.
- The ROCm prefill lead at depth (d204800, d262144) appears here on a candidate that ALSO survived Vulkan cleanly through model-max — worth a closer look at whether this is architecture-specific (48 layers / larger total size) or a general trend as this programme's candidates get larger.