Qwen3.8-Flash-Next (Unsloth first look + KingJones full-STRIX)benchedno guard
②Verdict
This page preserves two distinct Qwen3.8-Flash-Next histories instead of blending their artifacts, runtimes or backends. The Unsloth UD-Q4_K_XL/Vulkan first look remains visible under the held interpretation in clm-0125. The KingJones full-STRIX ROCmFP4/HIP campaign supersedes only its own earlier receipt-build stop and is current as BOUNDED/PARTIAL in clm-0127: standard capability coverage, corrected served QSA rows and separate cache traces are retained alongside the infrastructure-error denominator, a disclosed post-hoc energy join, incomplete PLE adapter, engine scatter and full-window refusal. The strict Tau2 result remains the 8k-budget 19/24 evaluated result and 19/26 all-attempt coverage: exactly two original calls reached the inherited ceiling, and a later two-task 16k recovery is held rather than folded because its transport capture was not lossless. Neither series is a production or role-fit recommendation, and no publisher number transfers.
A third, separate serving series is the one now actually in fleet production: the agention ROCmFP4-FAST-v2-ple16 imatrix quant on the agention/Laurent Vulkan fork (cfg-0182, run-0653 / eng-0277 / clm-0128). It is distinct from both the Unsloth and KingJones arms in artifact, runtime and method, and no number transfers between them. Unlike the first-look Unsloth path it keeps the entire n-gram/PLE table GPU-resident (per-head ple16) and ships a working MTP head, so it serves single-slot with adaptive speculative decoding: ~34 tok/s decode on the live tau2 workload (a served source-code depth sweep holds 16-26 t/s across 32-200K, rising with MTP acceptance; 44-47 on shallow generated code), prefill ~285 t/s, warm agent-turn TTFT ~0.6 s, loading in ~60 s and fitting fully on GPU (~95-106 GiB GTT) with no host-RAM thrash. That speed is why it was adopted for production. On tau2 airline it scored 21/25 (0.84, run-0653) at 12.74 Wh/correct (eng-0277) — but that run used the model's DEFAULT reasoning effort in the deployed config and is N=1, NOT reasoning-matched to the Unsloth reasoning=low baseline, so it is not a cross-series capability or energy ranking. The honest read: fastest to serve and GPU-resident, capability and energy pending a matched reasoning=low re-run.
③Best configuration
| model | Qwen3.8-Flash-Next-Q4_0-ROCmFP4-STRIX.gguf · Q4_0_ROCmFP4_STRIX · rev c6770d7442a06bf1d78edf28cec83e1ec93afdd34664c23ff898807b6b9349fa |
| engine | kingjones30/ROCmFPX 36e9acd · HIP/gfx1151 · host aihydra (igpu) |
| flags | Corrected served cohort: llama-server -c 262144 -ngl 999 -fa on -fit off -ctk f16 -ctv f16 -b 2048 -ub 512 -t 16 --jinja -rea off --ctx-checkpoints 512 --checkpoint-every-n-tokens 2048 -np 1. One slot, temperature 0, no MTP/draft/speculation. Runtime-default mmap is an explicit model-intrinsic exception: this artifact's PLE path depends on lazy file-backed mapping and is not comparable to the house --load-mode none floor. |
| template | not recorded at test time |
| tree | fork — carrying con-0017 |
config record cfg-0180
carrying The current KingJones full-STRIX record depends on a carried ROCmFPX fork and a publisher-specific quant. Its model-page comparison is source context, not evidence for any measured-here value. con-0017 clm-0127
every cell generated from the record at build time · throughput cells from cfg-0180 (same build 36e9acd, fa on, f16 KV) · provenance: measured-here throughout
③aDecode against context depth
⑤Open questions
- Can a future separately authorized tasks-23/24 collection capture the lossless adapter boundary needed to distinguish model output exhaustion from parser or serialization loss, without changing the strict 8k record?
- What causes the source-config mmap engine scatter at the shallow repeated rows, and can it be removed without changing the artifact/runtime boundary?
- Can a corrected all-tensor mincore adapter recover complete PLE working-set evidence on a future independent trace without retroactively filling this run?
- What does a matched UD-Q2 control and a cross-host pair (second Bosgame M5) show — is the Q4 result quant-robust and reproducible on identical silicon?
- Does ROCm 7.2.4 materially change the bounded served path under matched prompt bytes, counts and checkpoint flags?
- How does the production agention ROCmFP4-FAST-v2 series compare once matched? Its 21/25 (0.84) and 12.74 Wh/correct were default-reasoning + N=1 vs the Unsloth reasoning=low baseline; a reasoning=low, n=3 re-run on the same tau2 suite is the decisive within-series head-to-head for capability, energy-per-correct and decode-at-depth.