Home › Evidence › Records › con-0010
⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.

con-0010

open
citable URL: https://halobench.com/records/con-0010/ — this address never moves; the anchor /records/#con-0010 keeps resolving
kind report · upstream ggml-org/llama.cpp
opened 2026-08-22 · https://github.com/ggml-org/llama.cpp/issues/27122#issuecomment-5383228244
llama.cpp #27122 comment by mazinist (2026-08-22T23:51:13Z) independently confirms the MTP/CUDA multi-GPU lockup originally reported by tripletto, and validates a workaround on a completely different platform from the earlier zyxyunxin note (#issuecomment-5310571689, 2026-08-17). Hardware: Threadripper Pro 3945WX (WRX80, Gigabyte MC62-G40), 4x RTX A4000 16GB (PCIe 4.0 x16, NO NVLink / no P2P — nvidia-smi topo shows NODE), Ubuntu 24.04, kernel 6.14, closed driver 595.84, llama.cpp build 749f688fc (Aug 21, CUDA 13.2). Repro: Qwen3.8-27B-UD-Q6_K, --split-mode tensor -ts 25,25,25,25, --spec-type draft-mtp --spec-draft-n-max 3, 131072-token deep-context prefill via llama-benchy. 6/6 hard crashes, dying 2-4 min into the prefill; symptom on this platform is harsher than a lockup — the entire machine hard-powers-off (BMC logs "Power off/down", no Xid/AER/panic). KEY DISCRIMINATOR: the same workload on 2 GPUs (which uses the internal 2-device AllReduce path, not the meta-backend) is completely stable with MTP enabled, so the defect is specific to the multi-GPU tensor-split MTP path. WORKAROUND CONFIRMED: LLAMA_GRAPH_REUSE_DISABLE=1 lets the exact crash config survive the full 131072 prefill (610.82 t/s prefill, 27.53 t/s decode @131K, 38.43 t/s @4K), consistent with PR #24549's mechanism (graph reuse leaves dangling per-device tensor references when MTP/target contexts share memory under SPLIT_MODE_TENSOR). Also matches the reporter's notes: no-MTP tensor is stable, layer split never affected, and lockup frequency scales with --spec-draft-n-max. Upstream diagnostic record only: surface is CUDA multi-GPU, not our single-GPU ROCm/Vulkan Strix Halo gfx1151 aihydra box, so no HaloBench benchmark claim is made from these figures.

Cited by — computed at build time, never stored

claims clm-0114
candidate gate history qwen38-27b
docs benchmark-protocol