con-0016
open
citable URL: https://halobench.com/records/con-0016/ — this address never moves; the anchor /records/#con-0016 keeps resolving
kind pr · upstream ggml-org/llama.cpp
opened 2026-08-24 · https://github.com/ggml-org/llama.cpp/pull/25863#issuecomment-5397699526
Raw upstream comment by alexpooley on open PR #25863, retrieved directly from GitHub on 2026-08-25. The report is HIP/ROCm on the integrated gfx1151 surface, using Qwen3-4B-Instruct-2507 Q8_0, llama-perplexity, c4096, b4096, 12 chunks, and variable ubatch. Its master-versus-proposed-fix PPL table is stable at ub4096 (9.1542 versus 9.1542), but degrades on master as prompt chunking increases: ub2048 11.8957 versus 9.1527, ub1024 2259.7396 versus 9.1557, and ub512 49142.0371 versus 9.1594. The proposed explanation is that the scheduler does not protect caller-writable input during chunking; the author disclosed AI assistance for identifying and implementing the patch, while stating that the final change was personally tested and reviewed.
This is a source-bound upstream report, not a HaloBench result or a validated fix. It shares HIP/gfx1151 and low-ubatch shape with the local Qwen3.5-122B production-optimisation matrix (cfg-0151, b2048/ub512), but differs in model, weight quant, tool, corpus, batch, context and runtime. The local four-item guard (run-0588) passed but is not a chunked-input forward correctness test; it cannot prove this source does or does not apply. Therefore no recorded Qwen3.5 number is reinterpreted or invalidated, but future or modified HIP/gfx1151 low-ubatch performance work needs the new exact-fingerprint chunked-input control before a claim is admitted.
Cited by — computed at build time, never stored
docs benchmark-protocol