con-0018
open
citable URL: https://halobench.com/records/con-0018/ — this address never moves; the anchor /records/#con-0018 keeps resolving
kind pr · upstream ggml-org/llama.cpp
opened 2026-08-26 · https://github.com/ggml-org/llama.cpp/pull/27311
problem — On gfx1151 (AMD Strix Halo) HIP/ROCm with --parallel>1, the host can overwrite graph inputs while the GPU is still reading them asynchronously on the UMA pool, so responses leak / corrupt across concurrent requests (upstream #25992). Measured here: on fresh master c5a5535e the multi-slot Qwen3.8-27B worker degraded to EMPTY output within ~4 minutes of 4-slot tau2 load (run-0660). The same race also collapses speculative-decode acceptance under concurrent MTP (#27572) and corrupts tool arrays (#27579).
PR #27311 "Scheduler UMA ring buffer (+ sanitizer and fixes)" adds an input ring buffer on UMA devices so the host cannot clobber in-flight graph inputs. OPEN / unmerged as of the 2026-08-26 head commit. Carried on aihydra as branch pr27311 (commit c530ea79c) and is the load-bearing dependency of the Qwen3.8-27B worker serving config (cfg-0183): 0 EMPTY across a 106-minute 4-slot run vs the master build's ~4-minute failure (run-0660), and it is why the worker can serve --parallel 4 at all where the earlier Q8_0 config had to pin --parallel 1. A reproducibility hazard until merged: any rebuild of this worker must fetch the PR branch, not stock master. See cfg-0183, clm-0130.
Cited by — computed at build time, never stored
claims clm-0130
candidate gate history qwen38-27b