Home › Evidence › Records › qwen38-flash-next
⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.

qwen38-flash-next

live
citable URL: https://halobench.com/records/qwen38-flash-next/ — this address never moves; the anchor /records/#qwen38-flash-next keeps resolving
Qwen3.8-Flash-Next @ agention ROCmFP4-FAST-v2, single slot + adaptive MTP · kind model · engine igpu · tier daily-driver
runs on aibeast · config cfg-0182
⌁ current state live · ⌁ days in production 28

Lifecycle — append-only

2026-09-02
live · run-0653 clm-0128
The production serving model on aibeast since the 2026-09-02 Vulkan cutover: the agention ROCmFP4-FAST-v2-ple16 imatrix quant on the agention/Laurent Vulkan fork, served single-slot with adaptive speculative decoding (MTP) via fast-model.service. It keeps the full n-gram/PLE tables GPU-resident and ships a working MTP head, so it fits fully on GPU (~95-106 GiB GTT) and serves at ~34 tok/s decode on the live tau2 workload (a served source-code depth sweep holds 16-26 t/s across 32-200K, rising with MTP acceptance; 44-47 on shallow generated code), prefill ~285 tok/s, warm agent-turn TTFT ~0.6 s, loading in ~60 s. Adopted for production on serve speed and GPU-residency; Warden and Hermes are wired to it. Capability and energy remain N=1 and reasoning-unmatched to the Unsloth reasoning=low baseline (tau2 airline 21/25 = 0.84 at 12.74 Wh/correct, run-0653 / eng-0277), so those are not cross-series rankings — see clm-0128 and the model page for the bounded reading. The tau2 bench itself (cfg-0182) ran on the aihydra bench host; this node records where the same served quant runs in production, which is aibeast.