Home › Evidence › Records › clm-0114
⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.

clm-0114

communitymed ●●○
citable URL: https://halobench.com/records/clm-0114/ — this address never moves; the anchor /records/#clm-0114 keeps resolving

The llama.cpp MTP/spec-path multi-GPU fragility corroborates across a second, platform, with a validated workaround. On llama.cpp #27122 a CUDA/RTX-A4000-4x tensor-split build (Threadripper Pro 3945WX, no NVLink/P2P, closed driver 595.84, build 749f688fc) hard-crashes 6/6 on a Qwen3.8-27B Q6_K + draft-mtp n_max=3 deep 131072-token prefill under --split-mode tensor (symptom is a full machine power-off with BMC "Power off/down", no Xid/AER/panic), while the SAME workload on 2 GPUs (internal 2-device AllReduce path, not the meta-backend) is completely stable with MTP enabled. The defect is specific to the multi-GPU meta-backend MTP path. LLAMA_GRAPH_REUSE_DISABLE=1 lets the exact crash config survive the full 131072 prefill (610.82 t/s prefill, 27.53 t/s decode @131K, 38.43 t/s @4K), consistent with PR #24549's mechanism (graph reuse leaves dangling per-device tensor references when MTP/target contexts share memory under SPLIT_MODE_TENSOR). The corollary for the house: a HIGH --spec-draft- n-max interacts with the multi-GPU MTP path where the crash frequency scales with n-max.

verified 2026-08-23 · volatility medium
evidence con-0010

Note — the record's own working

Corroborating arm of the MTP-correctness dossier, independent platform confirmation of the MTP CUDA multi-GPU lockup axis (originally tripletto, now independently reproduced by mazinist on a different GPU/OS/platform, with the earlier zyxyunxin 2-GPU note). Reported surface is CUDA multi-GPU tensor split; HaloBench is single-GPU (aihydra, Strix Halo APU, ROCm/Vulkan), so this does NOT change any HaloBench boundary, config, run or production claim — it confirms the spec-path fragility and the value of a shallow/adaptive draft depth. No HaloBench benchmark number inherits from these figures.

Cited by — computed at build time, never stored

candidate gate history qwen38-27b
docs benchmark-protocol