cfg-0002
oom riskfork
citable URL: https://halobench.com/records/cfg-0002/ — this address never moves; the anchor /records/#cfg-0002 keeps resolving
| model | unsloth/Qwen3.5-122B-A10B-MTP-GGUF · UD-Q4_K_M · rev a91c2f7 |
| engine | ggml-org/llama.cpp b6142 · rocm · host aibeast (igpu) |
| flags | -ngl 999 --no-mmap --parallel 1 -fa 1 --jinja --spec-type draft-mtp --spec-draft-n-max 6 --ctx-checkpoints 4 |
| sampling | temp 1 · top_p 0.95 |
| template | 7b21e9c4a5f60d18 |
| tree | fork — carrying con-0001 |
Memory — static estimate vs observed peak
context 200,000 × 1 slot(s) = 200,000 tokens
⌁ total 80.4 GiB of 96 GiB pool · ⌁ headroom 15.6 GiB
⌁ risk utilisation 83.8% — basis estimate (the estimate has never predicted an OOM — clm-0001)
oom risk acknowledged (ack-0001, 2026-08-03, review 2026-09-15) — Runs at 83.8% of pool and has been the daily driver since 2026-06-28. Accepted deliberately: the 122B does not fit under 80% at any usable context, and the alternative is a materially weaker model. Mitigations after the 2026-07-21 OOM panic: --ctx-checkpoints capped at 4, desktop stack removed, swap raised to 16G. Revisit when aihydra arrives and the load can move.
Capability basis
unmeasured — performance numbers on this config stand on a guard alone, not an established capability
production node(s): qwen35-122b-a10b
Runs on this config (1)
| run | date | kind | suite | metrics |
|---|---|---|---|---|
| run-0002 | 2026-07-20 | performance | mtp-decode-sweep@v1 | decode_tps=36.5 · prefill_tps=144 · ttft_ms=0 · empty_arg_calls=0 |