Home › Evidence › Records › clm-0079
⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.

clm-0079

measured-herehigh ●●●
citable URL: https://halobench.com/records/clm-0079/ — this address never moves; the anchor /records/#clm-0079 keeps resolving

Qwen3-Coder-Next 80B (qwen3next hybrid, 512 experts/10 active + 1 shared, ~3B active) Q8_0's own throughput matrix has Vulkan leading decode at EVERY depth measured (44.13/37.0/26.4/21.73/19.15 t/s at d0/d32768/d131072/d204800/d262144 vs ROCm's 37.66/31.06/21.29/17.51/15.46) and prefill at the two shallowest cells (614.0/450.4 vs 476.6/346.8 t/s at d0/d32768) -- but the pattern INVERTS at the two deepest cells: ROCm overtakes prefill at d204800 (124.6 vs 98.75 t/s, Vulkan -20.7%) and by a wider margin at d262144, this candidate's model-max (104.56 vs 68.77 t/s, Vulkan -34.2%). Vulkan lost the GPU device zero times across all 10 matrix cells (both backends, all 5 depths, 2-attempt cap never triggered) -- a clean streak matching ornith-35b's precedent-breaking result and extending it to this programme's largest full-bench candidate so far (80B total, vs ornith's 35B). Decode remains the dominant real-world driver of interactive throughput, so Vulkan is the backend this bench serves the capability arm on, same choice as ornith-35b/nemotron3-super/deepseek-v4-flash despite the deep-prefill inversion this candidate newly shows.

verified 2026-08-17 · volatility low
evidence run-0396 run-0397 run-0406 run-0407 run-0414 run-0415 run-0424 run-0425

Note — the record's own working

Full matrix: cfg-0112 (ROCm, stock 3653e6d) and cfg-0113 (Vulkan, stock 3653e6d6d), both Q8_0, default f16 KV, -fa 1, --load-mode none, N=3 fresh-process reps at d0/d32768, N=1 at d131072/d204800/d262144 per protocol.json's throughput_ladder. Guard chain wired from birth: every performance run in both series carries `guard:` pointing at its backend's matching 4/4 house guard (run-0394 ROCm, run-0395 Vulkan -- the tau2-serving-config guard, since a backend swap is the `binary` lever the guard exists to catch). Scatter checked at every cell via sweep.sh's own CV gate; no cell exceeded ~1% run-to-run variance where N=3 applied. The deep-prefill inversion (ROCm ahead past d204800) is a genuine new finding for this programme -- every prior full-bench candidate that survived Vulkan to comparable depth (ornith-35b) kept Vulkan ahead on prefill too, or lost the device entirely (laguna-s-21, deepseek-v4-flash). This candidate is the first to show Vulkan surviving cleanly AND losing the prefill lead at depth, a distinct outcome from either precedent.

Cited by — computed at build time, never stored

model pages qwen3-coder-next-80b
candidate gate history qwen3-coder-next-80b