Home › Evidence › Records › clm-0128
⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.

clm-0128

measured-herehigh ●●●
citable URL: https://halobench.com/records/clm-0128/ — this address never moves; the anchor /records/#clm-0128 keeps resolving

Qwen3.8-Flash-Next ROCmFP4-FAST-v2-ple16 (agention imatrix quant, 87.06 GiB @ 4.23 bpw, everything GPU-resident incl. the per-head n-gram/PLE table) on the agention/Laurent Vulkan fork (cfg-0182, LaurentZuijdwijk/llama.cpp branch vulkan/qwen4exp-rocmfpx commit 5e085d12, build b10809, gfx1151/RADV) — first benchmark of the ROCmFP4 quant path on this fleet. ADMITTED CLAIM 1 (speed holds at depth, served path): single-slot with adaptive MTP, temp 0, source-code content, tokenizer-exact depths — decode 16.1 / 16.0 / 22.8 / 26.1 tok/s at 32K / 64K / 128K / 200K, prefill 298 -> 127 tok/s (run-0654..run-0657). Decode rises with MTP acceptance (0.49 -> 0.86 across the sweep) rather than falling with depth, i.e. it HOLDS to 200K on code; the ordering is N=1-noisy. Real agentic decode averaged ~34 tok/s over the tau2 run (run-0653); shallow generated code (red-black tree, JSON) reaches 44-47 tok/s at 0.86-0.91 acceptance; prose is lower (~18-22 tok/s, ~0.5 acceptance). Warm agent-turn TTFT ~0.6 s (cached prefix). ADMITTED CLAIM 2 (fits fully on GPU, stable): 87.06 GiB weights stay GPU- resident (GTT ~95-106 GiB at 1 slot, ~16-27 GiB free for KV), loads in 57-64 s, no host-RAM thrash and no OOM. The per-head ple16 layout keeps the ~51B n-gram/PLE table on GPU (no -ot offload); a joined-table variant with --ngram-on-disk exists for off-GPU placement but was not needed. 3-slot + adaptive-MTP loads and serves stably (no crash), but decode is bandwidth- bound so concurrency does not beat single-slot aggregate at depth. ADMITTED CLAIM 3 (capability, with a matched-comparison caveat): tau2 airline full-26 (run-0653, seed 42, claude-haiku-4.5 simulator, deterministic reward) scored 21/25 = 0.84 (26 attempted, task 10 excluded as an infrastructure error with 0 messages — a user-simulator/cloud fault). All 25 scored terminated clean; tool-calling and 196k single-slot context served the full agentic workload with no grammar or overflow failures. This sits BELOW the same-suite UD-Q4_K_XL result (run-0632, 24/26 = 0.9231) and the Ciru-IU4 result (23/26) — but it is NOT reasoning-matched: this run used the model's DEFAULT reasoning effort in the deployed serving config, whereas run-0632 used reasoning_effort=low. ADMITTED CLAIM 4 (energy, same caveat): wall-metered join (eng-0277, HA counter-difference) = 286.3 Wh over 6701 s, 267.5 Wh active above the 10.1 W idle floor, = 12.74 Wh per correct answer (0.386 p @ 30.3 p/kWh). This is HIGHER (worse) than UD-Q4_K_XL's 9.90 Wh/correct (eng-0275) — the raw-decode speed advantage does NOT translate to energy-per-correct, because the default reasoning effort generates more tokens and the run solved fewer tasks (21 vs 24). A reasoning=low re-run is the matched comparison. BLOCKED CLAIMS (explicitly NOT made): single trial (N=1), no variance bound; the decode-depth cells are N=1 served-path measurements, not the n=3 llama-bench ladder used for cfg-0178, and MTP acceptance (hence decode) is noisy cell-to-cell. NOT reasoning-matched to run-0632 (default vs low), so the capability and Wh-per-correct comparisons conflate quant, reasoning depth, and serving-vs-bench config — no "faster AND better/worse" conclusion is drawn. Quant quality "within 2.5% PPL" is a vendor (agention) claim, not independently perplexity-measured here. Lab result under agention's qwen-community license; the fork is a third-party WIP branch (verified branch tip, but not a released/immutable upstream build). This quant IS in production use on the fleet's serving box as of this date, but that is an operational choice for speed, not a capability endorsement over UD-Q4_K_XL.

verified 2026-09-03 · volatility medium
evidence cfg-0182 run-0653 eng-0277

Note — the record's own working

First benchmark of the agention ROCmFP4-FAST-v2 quant, under lab custody of AI Hydra, driven by the same operator flow as the UD-Q4_K_XL entry. The headline is speed-at-depth + clean GPU-resident fit; the capability and energy numbers are admitted but carry load-bearing caveats (N=1, default reasoning not matched to the UD-Q4_K_XL reasoning=low baseline). The obvious next step is a reasoning=low, n=3 matched re-run before any "which quant is better" claim.

Cited by — computed at build time, never stored

model pages qwen38-flash-next
candidate gate history qwen38-flash-next
docs models/qwen35-122b