⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.
clm-0059
measured-heremed ●●○superseded
citable URL: https://halobench.com/records/clm-0059/ — this address never moves; the anchor /records/#clm-0059 keeps resolving
On gfx1151 at f16 KV, DeepSeek-V4-Flash-0731 UD-IQ3_XXS is a fifth model measured under the per-model, per-phase backend rule (clm-0050), and it lands on the same side as Qwen3.8-27B (clm-0054): ROCm swept the full throughput matrix clean (d0 through d262144) while Vulkan lost the GPU device on both allowed attempts at every depth >=32768 (rc=134/SIGABRT, vk::DeviceLostError; kernel evidence: amdgpu ring timeout, ring reset, "device wedged, but recovered through reset" — 12 reset cycles total across the four failed cells). Even at the one depth Vulkan completed, d0, ROCm decode was already ahead (15.30 vs 12.44 t/s, +23%) while prefill was close (140.80 vs 134.73, ROCm +4.5%); ROCm's prefill lead widens with depth on every model measured so far on this chip. Serving this model at any useful context on this box requires ROCm — Vulkan cannot be trusted to survive a conversation that grows past 32k.
superseded by clm-0084 — the corrected statement lives there; this record keeps its URL and full text
corrections 2026-08-18: The stock 3653e6d Vulkan measurements remain valid for that build, but the general conclusion that Vulkan cannot survive past d0 is superseded by the clean baf6360be carried-fork series through d262144 (clm-0084).
Note — the record's own working
METHOD — deepseek-v4-flash-fullbench matrix: stock llama.cpp 3653e6d (ROCm) / 3653e6d6d (Vulkan, prefix-matches per house convention), f16 KV mandatory (candidate's COMMUNITY HAZARD note — quantised K trips an incoherence-rotation bug on this model), fa on, -ngl 999, --load-mode none, pp1024/tg256, llama-bench defaults otherwise (-b 2048/-ub 512). N=3 fresh-process reps at d0/d32768 (stddev <=0.6 t/s on every cell, well inside the 3% scatter-flag threshold — a healthy code path per protocol §0), N=1 above per protocol Tier1 ("the number does not move" past 32K). Vulkan capped at 2 attempts per cell per house policy; both attempts failed identically at every depth tried.
pp1024 / tg256 by depth (median of N):
| depth | ROCm | Vulkan | vk/rocm pp | vk/rocm tg | |---|---|---|---|---| | 0 | 140.80 / 15.30 | 134.73 / 12.44 | 0.957 | 0.813 | | 32,768 | 81.28 / 12.35 | DEVICE LOST | — | — | | 65,536 | 58.62 / 10.50 | DEVICE LOST | — | — | | 131,072 | 36.86 / 8.86 | DEVICE LOST | — | — | | 262,144 | 21.52 / 6.86 | DEVICE LOST | — | — |
DEVICE-LOST CELLS carry the process and kernel evidence verbatim (run-0300 through run-0303; full journalctl -k excerpt at aihydra ~/bench-results/deepseek-v4-flash-fullbench/matrix/vulkan-devicelost-journalctl-k.txt), the same discipline as qwen38-27b's d131072 cells (clm-0054) and the Nemotron OOM cells — a failed cell is a result, not a gap to paper over. The failure recovers without a reboot (kernel resets the compute ring each time) but is deterministic enough at 2/2 across four depths that the fast-fail policy applied.
THIS MODEL RAN CLEAN TO d262144 ON ROCm — the deepest cell measured in this programme so far (qwen38-27b's matrix stopped at d131072). See clm-0060 (fit) for why this model tolerates depth so much more cheaply than a conventional hybrid-attention model: its KV cache is ~9x smaller per token.