Models › qwen36-35b

Qwen3.6-35B-A3Bbenchedguard 4/4

35B / 3B activeMoE · no NextN layers in this GGUFquant held: UD-Q4_K_XLfirst measured 2026-06-27latest run 2026-08-08

Verdict

A 35-billion-parameter mixture-of-experts model activating 3B per token, held at UD-Q4_K_XL — the reflex-tier case. It decodes at 2.3x the 122B on identical hardware with prefill above 1000 tok/s, and passes the capability guard 4/4 clm-0023. Flash attention is not optional at depth: without it the model loses 88% of its decode from empty context to 131k, against 44% with it clm-0028. Speculation buys almost nothing on varied prompts, and MTP is unavailable because this build carries no NextN layers clm-0026. The community's 121 tok/s headline is four streams pooled on a different quant format — the honest lever on this model is ROCmFP4, worth roughly 27% on the no-speculation floor, at the price of a third fork divergence clm-0026. Agentic capability is unmeasured: the τ² protocol has not been run on this model, and its reflex-tier standing currently rests on throughput plus a guard. A blind eight-task conceptual screen placed it 7/8 on that rubric clm-0004.

verdict written 2026-08-13 · every number above stands next to the claim chip that carries it

Best configuration

modelQwen3.6-35B-A3B-UD-Q4_K_XL.gguf · UD-Q4_K_XL
engineggml-org/llama.cpp 3653e6d · rocm · host aihydra (igpu)
flags-ngl 999 -fa 1 -b 2048 -ub 512 -ctk f16 -ctv f16 -t 16
templatenot recorded at test time
treeupstream — stock

backfilled aged evidence — reconstructed from the archive · config record cfg-0012

decode @ 0
51.01 t/s
run-0031 · CV 0.2%
decode @ 32k
42.45 t/s
run-0035 · CV 0.1%
prefill @ 0
1079.38 t/s
run-0030 · CV 0.5%
prefill @ 32k
598.68 t/s
run-0034 · CV 0.6%
Wh per correct answer
unmeasured
open question ↓

every cell generated from the record at build time · throughput cells from cfg-0012 (same build 3653e6d, fa on, f16 KV) · provenance: measured-here throughout

③aDecode against context depth

f16 KV · fa onq8_0 KV · stockflash attention off
0102030405060032k65k131k204.8kdecode t/scontext depth (tokens)51.01 t/s @ depth 0 · run-0031 · CV 0.2% · N=351.12 t/s @ depth 0 · run-0087 · CV 0% · N=349.80 t/s @ depth 4k · run-0033 · CV 0.1% · N=342.45 t/s @ depth 32k · run-0035 · CV 0.1% · N=342.45 t/s @ depth 32k · run-0089 · CV 0.1% · N=328.46 t/s @ depth 131k · run-0091 · CV 0.1% · N=350.07 t/s @ depth 0 · run-0093 · CV 1.1% · N=335.28 t/s @ depth 32k · run-0095 · CV 0.1% · N=318.48 t/s @ depth 131k · run-0097 · CV 0% · N=350.88 t/s @ depth 0 · run-0081 · CV 1.1% · N=322.84 t/s @ depth 32k · run-0083 · CV 0.1% · N=35.91 t/s @ depth 131k · run-0085 · CV 0% · N=3f16 KV · fa on · 28.46q8_0 KV · stock · 18.48flash attention off · 5.91
3 reps per cell · max CV 1.1% · build 3653e6d · records: run-0031 run-0087 run-0033 run-0035 run-0089 run-0091 run-0093 run-0095 run-0097 run-0081 run-0083 run-0085

Other configurations tested — each as a delta against best

variantΔ decodeΔ turns (paired tasks)Δ energynoterecords
flash attention off-79% @131kInvisible at empty context, catastrophic at depth — and quantised KV cannot create a context at all without flash attention. clm-0028
q8_0 KV · stock build-35% @131kNo dequant-patch arm has been run on this model; what the patch recovers here is an open question, not an assumption. clm-0022
speculation (ngram-mod / ngram-cache)Close to worthless on varied prompts — the scatter is roughly three times the mean gain, ngram-cache buys nothing, and draft-mtp fails cleanly because this GGUF carries no NextN layers. Measured in the spec A/B harness; the arms are recorded in the claim, not ingested as runs. clm-0026

deltas computed at build time from the named runs, paired-task traces and metered windows — a hazard row shows words where a delta would mislead

Open questions

Provenance

bench host aihydra · rocm · ggml-org/llama.cpp 3653e6d
discipline 3 reps per throughput cell · scatter published per cell (max CV 0.6%) · guard chain on every performance series
window 2026-06-27 → 2026-08-08