Models › qwen38-flash-next
⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.

Qwen3.8-Flash-Next (Unsloth first look + KingJones full-STRIX)benchedno guard

~180B stored / 125B MoE (512 experts, ~6B active) + 51B N-gram/PLE tables + 4B MTP head + visionQwen4-arch MoE (qwen4exp) · 125B MoE / ~6B active + 51B N-gram/PLE tables + hybrid QSA attention + vision; PLE quantisation, tensor placement and head availability are artifact-specificquant held: Two non-transferable series · Unsloth UD-Q4_K_XL/Vulkan first look · KingJones Q4_0_ROCmFP4_STRIX/HIP bounded campaignfirst measured 2026-08-26latest run 2026-08-30

②Verdict

This page preserves two distinct Qwen3.8-Flash-Next histories instead of blending their artifacts, runtimes or backends. The Unsloth UD-Q4_K_XL/Vulkan first look remains visible under the held interpretation in clm-0125. The KingJones full-STRIX ROCmFP4/HIP campaign supersedes only its own earlier receipt-build stop and is current as BOUNDED/PARTIAL in clm-0127: standard capability coverage, corrected served QSA rows and separate cache traces are retained alongside the infrastructure-error denominator, a disclosed post-hoc energy join, incomplete PLE adapter, engine scatter and full-window refusal. The strict Tau2 result remains the 8k-budget 19/24 evaluated result and 19/26 all-attempt coverage: exactly two original calls reached the inherited ceiling, and a later two-task 16k recovery is held rather than folded because its transport capture was not lossless. Neither series is a production or role-fit recommendation, and no publisher number transfers.

A third, separate serving series is the one now actually in fleet production: the agention ROCmFP4-FAST-v2-ple16 imatrix quant on the agention/Laurent Vulkan fork (cfg-0182, run-0653 / eng-0277 / clm-0128). It is distinct from both the Unsloth and KingJones arms in artifact, runtime and method, and no number transfers between them. Unlike the first-look Unsloth path it keeps the entire n-gram/PLE table GPU-resident (per-head ple16) and ships a working MTP head, so it serves single-slot with adaptive speculative decoding: ~34 tok/s decode on the live tau2 workload (a served source-code depth sweep holds 16-26 t/s across 32-200K, rising with MTP acceptance; 44-47 on shallow generated code), prefill ~285 t/s, warm agent-turn TTFT ~0.6 s, loading in ~60 s and fitting fully on GPU (~95-106 GiB GTT) with no host-RAM thrash. That speed is why it was adopted for production. On tau2 airline it scored 21/25 (0.84, run-0653) at 12.74 Wh/correct (eng-0277) — but that run used the model's DEFAULT reasoning effort in the deployed config and is N=1, NOT reasoning-matched to the Unsloth reasoning=low baseline, so it is not a cross-series capability or energy ranking. The honest read: fastest to serve and GPU-resident, capability and energy pending a matched reasoning=low re-run.

verdict written — · every number above stands next to the claim chip that carries it

③Best configuration

modelQwen3.8-Flash-Next-Q4_0-ROCmFP4-STRIX.gguf · Q4_0_ROCmFP4_STRIX · rev c6770d7442a06bf1d78edf28cec83e1ec93afdd34664c23ff898807b6b9349fa
enginekingjones30/ROCmFPX 36e9acd · HIP/gfx1151 · host aihydra (igpu)
flagsCorrected served cohort: llama-server -c 262144 -ngl 999 -fa on -fit off -ctk f16 -ctv f16 -b 2048 -ub 512 -t 16 --jinja -rea off --ctx-checkpoints 512 --checkpoint-every-n-tokens 2048 -np 1. One slot, temperature 0, no MTP/draft/speculation. Runtime-default mmap is an explicit model-intrinsic exception: this artifact's PLE path depends on lazy file-backed mapping and is not comparable to the house --load-mode none floor.
templatenot recorded at test time
treefork — carrying con-0017

config record cfg-0180
carrying The current KingJones full-STRIX record depends on a carried ROCmFPX fork and a publisher-specific quant. Its model-page comparison is source context, not evidence for any measured-here value. con-0017 clm-0127

τ² airline · KingJones full-STRIX · evaluated-only denominator; two infrastructure errors retained separately
0.792 ±0.162
run-0639 · passed 19/24
Wh per correct answer (τ², KingJones full-STRIX · evaluated-only denominator; two infrastructure errors retained separately)
20.54 Wh

every cell generated from the record at build time · throughput cells from cfg-0180 (same build 36e9acd, fa on, f16 KV) · provenance: measured-here throughout

③aDecode against context depth

Vulkan (RADV) engine floor — first-look, non-servedagention ROCmFP4-FAST-v2 · revised depth-sweep (served + MTP)
0102030032k65k131k204.8kdecode t/scontext depth (tokens)first-look serving context 32k22.82 t/s @ depth 0 · run-063314.09 t/s @ depth 32k · run-06347.09 t/s @ depth 131k · run-06355.73 t/s @ depth 204.8k · run-063616.10 t/s @ depth 32k · run-065416.00 t/s @ depth 65k · run-065522.80 t/s @ depth 131k · run-065626.10 t/s @ depth 204.8k · run-0657agention ROCmFP4-FAST-v2 · revised depth-sweep (served + MTP) · 26.10Vulkan (RADV) engine floor — first-look, non-served · 5.73
build 250b614 / 5e085d1 · Historical non-served engine-floor ladder, not intended-config or production throughput. Vulkan/RADV r3 fresh-process decode was 22.82 t/s at d0, 14.09 at d32768, 7.09 at d131072, 5.73 at the >200k production depth d204800; prefill 401.8 / 201.3 / 86.9 / 63.9 t/s respectively. d245760 failed as a matrix cell (opaque — rc=0, truncated output, no stderr error) but d204800 ran clean standalone. 262144 hits a Vulkan workgroup-count GGML_ASSERT (cap -c <= 262140). ROCm is not charted: its mmap-load path stalls on gfx1151, so Vulkan/RADV is the only backend that runs this model here. The comparison series is a DIFFERENT method and a DIFFERENT (agention ROCmFP4-FAST-v2) quant, not comparable cell-for-cell to the hero: single (N=1) served-path requests through the proxy on source-code content with adaptive MTP, tokenizer-trimmed to land exactly on 32768/65536/131072/204800 — decode 16.1 / 16.0 / 22.8 / 26.1 t/s (rising with MTP acceptance 0.49 -> 0.86, i.e. holding to 200K on code), prefill 298 -> 127 t/s (run-0654..run-0657). It shows the deployed serving config's real code-generation speed at depth, not an engine floor. · records: run-0633 run-0634 run-0635 run-0636 run-0654 run-0655 run-0656 run-0657

⑤Open questions

⑥Provenance

bench host aihydra · HIP/gfx1151 · kingjones30/ROCmFPX 36e9acd
discipline no chart-cell CV aggregate · guard or written waiver on every performance series
window 2026-08-26 → 2026-08-30