Home › Evidence › Records › clm-0065
⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.

clm-0065

measured-heremed ●●○
citable URL: https://halobench.com/records/clm-0065/ — this address never moves; the anchor /records/#clm-0065 keeps resolving

On gfx1151, Laguna-S-2.1 Q4_K_M is a sixth model measured under the per-model, per-phase backend rule (clm-0050), and unlike nemotron3-super and deepseek-v4-flash (both Vulkan-favoured on this box), ROCm wins outright here: ROCm swept the full throughput matrix clean at every depth tried (d0, d32768, d131072 -- 300.7/23.68 t/s down to 116.43/7.19 t/s pp/tg), while Vulkan completed only d0 and d32768 (256.9/13.29 and 51.6/12.27 t/s pp/tg, ROCm ahead on every figure at every depth both backends completed) and lost the GPU device on both allowed attempts at d131072 (rc=134, vk::DeviceLostError; kernel evidence: amdgpu ring timeout, ring reset FAILED, a MODE2 GPU reset, "device wedged, but recovered through reset" -- the same device-loss class as clm-0054/clm-0059's precedent, now a fourth confirmed occurrence on this silicon). ROCm's decode lead is largest at d0 (23.68 vs 13.29 t/s, +78%) and narrows but does not close at d32768 (15.53 vs 12.27 t/s, +26%); ROCm's prefill lead widens sharply with depth (0.957x at d0 -> 4.17x at d32768, ROCm ahead both times). Serving this model at any useful context on this box requires ROCm -- Vulkan cannot be trusted to survive a conversation that grows past 32k, and even where it survives it is already behind on both phases.

verified 2026-08-16 · volatility medium
evidence run-0317 run-0318 run-0319 run-0320 run-0321 run-0322 run-0323 run-0324 run-0325 run-0326 run-0327

Note — the record's own working

METHOD -- laguna-s21-fullbench matrix: stock llama.cpp 3653e6d (ROCm) / 3653e6d6d (Vulkan, prefix-matches per house convention), f16 KV (matches the 2026-08-10 screen's serving config), fa on, -ngl 999, --load-mode none, pp1024/tg256, llama-bench defaults otherwise (-b 2048/-ub 512, not passed explicitly). N=3 fresh-process reps at d0/d32768 (stddev <=1.01 t/s on every cell, CoV <=0.88% on every cell -- well inside the 3% scatter-flag threshold, a healthy code path per protocol §0), N=1 at d131072 per protocol Tier1 / deepseek-v4-flash convention. Vulkan capped at 2 attempts per cell per house policy; both attempts at d131072 failed identically. pp1024 / tg256 by depth (mean of N): | depth | ROCm | Vulkan | vk/rocm pp | vk/rocm tg | |---|---|---|---|---| | 0 | 300.67 / 23.68 | 256.92 / 13.29 | 0.855 | 0.561 | | 32,768 | 215.11 / 15.53 | 51.64 / 12.27 | 0.240 | 0.790 | | 131,072 | 116.43 / 7.19 | DEVICE LOST | -- | -- | DEVICE-LOSS CELL (run-0327): attempt 1 (08:21:07Z-08:38:22Z, 1035s) and attempt 2 (08:38:22Z-08:53:03Z, 881s) both terminated identically -- "radv/amdgpu: The CS has been cancelled because the context is lost. This context is innocent." -> vk::Queue::submit: ErrorDeviceLost, rc=134. The job's own per-cell dmesg capture (matrix/vulkan-d131072-rep1.dmesg.txt) is 0 BYTES -- a capture bug in queue-laguna-s21-fullbench-matrix.sh: its on-failure `dmesg` call ran without `sudo` and silently produced nothing under this host's permissions (the aihydra user has passwordless sudo per /etc/sudoers.d/90-aihydra; the script just never invoked it for this call site). This is a gap in the SCRIPT, not evidence the failure did not happen. The kernel evidence was independently re-pulled this session via `sudo dmesg -T` scoped to the failure window and confirms the device-loss class exactly: 08:37:54Z ring comp_1.2.0 timeout, signaled seq=7255308 -> ring reset -> recovered 08:38:25Z ring comp_1.2.0 timeout, signaled seq=7255338 -> ring reset -> recovered 08:44:46Z Fence fallback timer expired on ring comp_1.2.0 08:53:05Z ring comp_1.2.0 timeout, signaled seq=7302878 -> ring reset FAILED -> GPU reset begin! (MODE2 reset) -> GPU reset succeeded, trying to resume -> SMU resumed successfully -> GPU reset(21) succeeded -> "device wedged, but recovered through reset" -> "*ERROR* Failed to initialize parser -125!" The box self-recovered via the MODE2 reset on the second failed attempt -- no reboot needed, no lingering wedge, matching the clm-0054/clm-0059 precedent exactly. Fix needed in queue-laguna-s21-fullbench-matrix.sh for future runs: prefix its on-failure `dmesg` call with `sudo`, or its evidence file silently stays empty every time. Laguna's SWA-heavy layer mix (36/48 sliding-window + 12/48 global attention, per the candidate record) does not spare it from this class -- unlike deepseek-v4-flash's MLA-compressed cache, which ran clean on ROCm through d262144 and was never tested to the point of Vulkan failure at comparable depth in this programme's matrix design.

Cited by — computed at build time, never stored

model pages laguna-s-21
candidate gate history laguna-s-21