Models › qwen3-coder-next-80b
⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.

Qwen3-Coder-Next (Qwen)benchedguard 4/4

80B total / ~3B active (A3B-class MoE, 512 experts / 10 active + 1 shared)hybrid MoE (qwen3next) · 12/48 full-attention + 36/48 gated-DeltaNet SSM layers · 80B total / ~3B active (512 experts / 10 active + 1 shared) — no MTP/speculation path in this GGUFquant held: Q8_0 · designated primary quant, fit-projected and confirmed comfortable against the 120 GiB GTT windowfirst measured 2026-08-17latest run 2026-08-17

②Verdict

A hybrid mixture-of-experts (qwen3next architecture, 12 of 48 blocks carrying full-attention layers via full_attention_interval=4, the other 36 gated- DeltaNet-class SSM/linear-attention with no growing KV cache — the same architecture family as this lab's production 122B and the already-benched ornith-35b, just deeper and roughly 2.3x larger) served at Q8_0 — this candidate's own designated primary quant, chosen because it fit-projected comfortably against the 120 GiB GTT window even at full model-max context, a PROTOCOL_OVERRIDE against protocol.json's quants.expected list on the same rationale ornith-35b and qwen38-27b already established — with no speculation path: this job's own from-scratch GGUF header parser found zero nextn/mtp/ eagle/medusa/draft tensors among all 807 tensors across all 4 shards.

Backend is a genuine, nuanced finding here, not a formality: Vulkan wins decode at every measured depth in its five-depth ladder and prefill at the two shallowest cells, but ROCm overtakes prefill at the two deepest cells (d204800, d262144) — this page's cited case showing Vulkan surviving cleanly through model-max (zero device-loss across all 10 matrix cells, extending the ornith-35b clean-streak comparison to this 80B candidate) while ALSO losing ground on prefill at depth. Decode remains the dominant driver of interactive throughput, so Vulkan still serves the capability arm. KV cost is a measured 24 KiB/token (f16 KV, derived from the GGUF's own attention-layer metadata) — the higher than the cited hybrid comparator figures on this page, but still a rounding error against the 120 GiB window: even this candidate's own 262,144-token model-max projects to roughly 6 GiB of KV, weights alone account for the other 79 GiB.

Capability is where this candidate's story turns: the FULL 26-task tau2 airline run completed end to end (rc=0, not wall-bound) at 0.5385 mean reward (14/26), 192 tool-call messages, 0 empty turns — VALID under the SMOKE gate, but below qwen38-27b's 0.577 on the cited same-protocol airline comparison. This is a coder model — built and marketed for agentic coding, not customer-service dialogue — being scored on tau2's airline domain purely for cross-model comparability with the cited same- protocol candidates; the number says how a coding specialist performs OFF its home ground, not what it can do on the work it is actually built for. Energy inverts the usual pattern: 4.7896 Wh per correct answer, 0.1451 pence at the standing tariff — lower than ornith-35b's cited 6.4834 Wh record on the same standard-26-task tau2 protocol. That pairs a low energy-per-correct-answer measurement with weak airline-domain reward, the inverse of Ornith's high-reward / low-energy record rather than a claim of better capability.

This efficiency figure was screened against the then-published benched models' Wh-per-correct-answer records before publishing (qwen35-122b's 5.713 Wh/correct comes from a 5-task smoke window predating the standard-26-task convention — a different scale, not a like-for-like comparator). The confound worth naming plainly: a lower-reward run that reaches user_stop in fewer effective turns is not the same phenomenon as efficient problem-solving, and this verdict does not claim this candidate is intrinsically more energy-efficient PER UNIT OF CAPABILITY than ornith-35b or any other comparator — only that the measured Wh-per-correct-answer figure, on the identical protocol, is low for the cited comparator set, and that it arrives paired with weak airline-domain reward rather than a higher score.

verdict written 2026-08-17 · every number above stands next to the claim chip that carries it

③Best configuration

modelQwen3-Coder-Next-Q8_0-00001-of-00004.gguf · Q8_0
engineggml-org/llama.cpp 3653e6d · vulkan · host aihydra (igpu)
flags-ngl 999 -fa on -c 32768 -dev Vulkan0 --jinja --reasoning-format deepseek -n 4096 --parallel 1 --load-mode none --slots
templatenot recorded at test time
treeupstream — stock

config record cfg-0111

decode @ 0
44.13 t/s
run-0415 · CV 0%
decode @ 32k
36.97 t/s
run-0427 · CV 0%
prefill @ 0
614.00 t/s
run-0414 · CV 0%
prefill @ 32k
450.49 t/s
run-0426 · CV 0%
τ² airline · thinking off
0.538 ±0.192
run-0433 · passed 14/26
turns to done · median (all tasks)
30
run-0433 · successes only: 27 · max 46
wall-clock to done · median (successes only)
1.0 min
run-0433 · failures excluded — they have no done
Wh per correct answer (τ², thinking off)
4.79 Wh
eng-0168 run-0433 · 0.15 p per answer

every cell generated from the record at build time · throughput cells from cfg-0113 (same build 3653e6d, fa on, f16 KV) · provenance: measured-here throughout

③aDecode against context depth

VulkanROCm
01020304050065k131k204.8k262.1kdecode t/scontext depth (tokens)serving context 32k44.09 t/s @ depth 0 · run-0417 · CV 0% · N=144.13 t/s @ depth 0 · run-0415 · CV 0% · N=144.19 t/s @ depth 0 · run-0419 · CV 0% · N=136.97 t/s @ depth 32k · run-0427 · CV 0% · N=136.97 t/s @ depth 32k · run-0431 · CV 0% · N=137.06 t/s @ depth 32k · run-0429 · CV 0% · N=126.40 t/s @ depth 131k · run-0421 · CV 0% · N=121.73 t/s @ depth 204.8k · run-0423 · CV 0% · N=119.15 t/s @ depth 262.1k · run-0425 · CV 0% · N=137.59 t/s @ depth 0 · run-0397 · CV 0% · N=137.65 t/s @ depth 0 · run-0399 · CV 0% · N=137.73 t/s @ depth 0 · run-0401 · CV 0% · N=130.43 t/s @ depth 32k · run-0409 · CV 0% · N=130.91 t/s @ depth 32k · run-0413 · CV 0% · N=131.84 t/s @ depth 32k · run-0411 · CV 0% · N=121.29 t/s @ depth 131k · run-0403 · CV 0% · N=117.51 t/s @ depth 204.8k · run-0405 · CV 0% · N=115.46 t/s @ depth 262.1k · run-0407 · CV 0% · N=1Vulkan · 19.15ROCm · 15.46
1 rep per cell · max CV 0.0% · build 3653e6d / 3653e6d · Both backends completed all five ladder depths cleanly (0, 32768, 131072, 204800, 262144/model-max) — Vulkan lost the GPU device zero times across all 10 cells, extending the ornith-35b clean-streak comparison to this 80B full-bench candidate. Vulkan leads decode at every depth including the new ceiling, but the prefill lead INVERTS past d204800 — ROCm overtakes prefill at the two deepest cells, a pattern this programme has not seen on a candidate that also survived Vulkan cleanly to model-max. · records: run-0417 run-0415 run-0419 run-0427 run-0431 run-0429 run-0421 run-0423 run-0425 run-0397 run-0399 run-0401 run-0409 run-0413 run-0411 run-0403 run-0405 run-0407

④Other configurations tested — each as a delta against best

variantΔ decodeΔ turns (paired tasks)Δ energynoterecords
Vulkan backend decode lead · same quant+17% @0—+29%Vulkan wins decode at every depth measured (+17% at d0, +19% at d32768, +24% at d131072/d204800/d262144) and prefill at the two shallowest cells (+29% at d0, +30% at d32768) — but ROCm overtakes prefill at the two deepest cells (d204800 +26.2%, d262144 +52.0% in ROCm's favour). Vulkan is the backend this bench serves the capability arm on, matching the house convention of following decode (the dominant driver of interactive throughput) even where prefill splits by depth. clm-0079
KV cache cost (fit)———24 KiB/token (f16 KV) — higher than the cited hybrid full-bench comparator figures on this page, consistent with this candidate carrying more full-attention layers than those cited comparators (12 of 48, vs ornith-35b's 10 of 40) even though the fraction of full-attention blocks is similar. Still cheap in absolute terms against this box: even at this candidate's own model-max context (262,144 tokens), KV projects to roughly 6 GiB against the 120 GiB GTT window — weights (79 GiB) dominate the footprint by an order of magnitude, and fit was never a risk at any depth this bench measured. clm-0080
No speculation / MTP———This job's own from-scratch GGUF header parser scanned all 807 tensors across all 4 shards and found zero nextn/mtp/eagle/medusa/draft tensors — genuinely no speculative-decode path exists in this artifact, corroborated independently by the base model's own config.json before download. Every throughput and capability number in this bench is plain decode. clm-0081
Wh-per-correct-answer with low airline-domain reward———4.7896 Wh/correct is a low same-protocol energy figure beside this page's cited full-bench comparators (ornith-35b 6.4834, laguna-s-21 7.2588, deepseek-v4-flash 24.29, nemotron3-super 32.43, qwen38-27b 38.29) — but it arrives with a weak airline-domain capability score (0.5385 mean reward, below qwen38-27b's 0.577). The efficiency figure is measured on the identical protocol as those cited comparator records, but it is not accompanied by a capability lead and should not be read as one. clm-0082

deltas computed at build time from the named runs, paired-task traces and metered windows — a hazard row shows words where a delta would mislead

⑤Open questions

⑥Provenance

bench host aihydra · vulkan · ggml-org/llama.cpp 3653e6d
discipline 1 reps per throughput cell · scatter published per cell (max CV 0.0%) · guard or written waiver on every performance series
window 2026-08-17 → 2026-08-17