Home › Evidence › Records › clm-0009

clm-0009

communitymed ●●○
citable URL: https://halobench.com/records/clm-0009/ — this address never moves; the anchor /records/#clm-0009 keeps resolving

A ROCmFP4 iMatrix quant of our exact model (Qwen3.5-122B-A10B) reports 60.70 GiB and 28.505 tok/s decode with MTP OFF — against our measured 20.96 tok/s MTP-off at UD-Q4_K_M. That is roughly +36% decode for ~5-6 GB less memory, if it reproduces.

verified 2026-08-05 · volatility high

Note — the record's own working

Source: vmlinux/Qwen3.5-122B-A10B-ROCmFP4-iMatrix-GGUF on HuggingFace, updated 2026-07-26. Built with ROCmFPX's Q4_0_ROCMFP4_STRIX_LEAN recipe and tested ONLY on gfx1151 — our exact silicon. Published figures: 60.70 GiB, greedy decode 28.505 tok/s, 4,277-token prefill 356.9 tok/s, and BF16 mean KLD 0.041366 +/- 0.002531. NOTABLE ON METHOD: they publish KL divergence WITH an uncertainty interval — the same instrument we selected for KV-quant quality, applied to weight quantisation. That gives us a comparable reference point for what a small quality delta looks like on this model, which we did not previously have. WHY IT MATTERS FOR PLACEMENT: 60.70 GiB against our ~66 GB frees ~5-6 GB. Config 2 in model-phases.md has the 122B and 35B co-resident at ~108 of ~120 GiB, which we flagged as uncomfortably tight given our OOM history. This quant would meaningfully relieve that — possibly the difference between viable and reckless. ⚠ THE COST: it uses custom ROCmFP4 tensor types and **stock llama.cpp cannot load it**. It requires the ROCmFPX runtime. That is a THIRD carried fork alongside con-0001 (our checkpoint sidecar) and the community quantized-KV fix — and our reproducibility debt is already a stated concern. Worse for measurement: a different runtime is a different config fingerprint, so ROCmFP4 numbers CANNOT sit in the same series as our existing ones (protocol section 8). It is a new experiment, not a faster version of the old one. Our Tier 1 pipeline also assumes stock llama-bench, which would need the ROCmFPX build. ALSO: the MTP heads ship as a SEPARATE 2.14 GiB companion file, where our current GGUF embeds them. Flags: "experimental", 1,929 downloads, 6 likes — low adoption, so we would be early rather than following. TEST IT AS A LAB-BOX CANDIDATE, not a production swap: same suite, same context, against our own UD-Q4_K_M baseline, with the runtime difference recorded as the config change it is. If +36% MTP-off holds and MTP scales on top, it plausibly lands near 42 tok/s against our current 36.5.

Cited by — computed at build time, never stored

model pages qwen35-122b
docs model-phases