Home › Evidence › Records › con-0019
⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.

con-0019

carrying
citable URL: https://halobench.com/records/con-0019/ — this address never moves; the anchor /records/#con-0019 keeps resolving
kind fork · upstream ROCm/rocm-systems
opened 2026-08-26 · https://github.com/pwilkin/rocm-systems/tree/ilintar-experiments
carrying — reproducibility hazard: in use but not upstream. Every config listing this id in its divergence runs a fork.
problem — Strix Halo (gfx1151) decode is dispatch/overhead-bound, not compute-bound: each generated token issues hundreds of tiny GPU kernels, and the per-launch CPU cost of building and submitting HIP command packets becomes the bottleneck, leaving the GPU waiting. Standard HIP graphs help, but rebuild/re-validate the command buffer each replay.
"Retained PM4 dispatch" — a custom AMD HIP/ROCm runtime (rocm-systems ilintar-experiments @ 78d1160, the CLR/hipamd graph path) that keeps a graph's low-level PM4 command buffer resident and replays it, cutting per-token launch overhead. Activated by ENABLE_RETAINED_PM4=1 -> DEBUG_HIP_GRAPH_PM4=1 with HIP graphs on. This is a RUNTIME change (libamdhip64), not a llama.cpp change. Carried on aihydra as the custom rocr+clr runtime behind the pwilkin 27B build (cfg-0185/cfg-0186). It is the ONE piece of the pwilkin stack with NO upstream PR: the author states "all of this is in upstream PRs except for the PM4 stuff which I don't expect to be accepted (but I might retire for proper HRX support when it's there)." A novel-fork reproducibility hazard: any rebuild must build this custom runtime from source; there is no merge path. Measured effect on our worker model is small (~+4%, clm-0135) — it is not the source of the pwilkin speed advantage. The pwilkin llama.cpp changes, by contrast, ARE upstream PRs (e.g. #27311, con-0018).

Cited by — computed at build time, never stored

claims clm-0135
candidate gate history qwen38-27b