Home › Evidence › Records › cfg-0002

cfg-0002

oom riskfork
citable URL: https://halobench.com/records/cfg-0002/ — this address never moves; the anchor /records/#cfg-0002 keeps resolving
modelunsloth/Qwen3.5-122B-A10B-MTP-GGUF · UD-Q4_K_M · rev a91c2f7
engineggml-org/llama.cpp b6142 · rocm · host aibeast (igpu)
flags-ngl 999 --no-mmap --parallel 1 -fa 1 --jinja --spec-type draft-mtp --spec-draft-n-max 6 --ctx-checkpoints 4
samplingtemp 1 · top_p 0.95
template7b21e9c4a5f60d18
treefork — carrying con-0001

Memory — static estimate vs observed peak

context 200,000 × 1 slot(s) = 200,000 tokens
⌁ total 80.4 GiB of 96 GiB pool · ⌁ headroom 15.6 GiB
⌁ risk utilisation 83.8% — basis estimate (the estimate has never predicted an OOM — clm-0001)
oom risk acknowledged (ack-0001, 2026-08-03, review 2026-09-15) — Runs at 83.8% of pool and has been the daily driver since 2026-06-28. Accepted deliberately: the 122B does not fit under 80% at any usable context, and the alternative is a materially weaker model. Mitigations after the 2026-07-21 OOM panic: --ctx-checkpoints capped at 4, desktop stack removed, swap raised to 16G. Revisit when aihydra arrives and the load can move.

Capability basis

unmeasured — performance numbers on this config stand on a guard alone, not an established capability

production node(s): qwen35-122b-a10b

Runs on this config (1)

rundatekindsuitemetrics
run-00022026-07-20performancemtp-decode-sweep@v1decode_tps=36.5 · prefill_tps=144 · ttft_ms=0 · empty_arg_calls=0

Cited by — computed at build time, never stored