clm-0002
communitymed ●●○
citable URL: https://halobench.com/records/clm-0002/ — this address never moves; the anchor /records/#clm-0002 keeps resolving
Multi-token prediction gains MORE at heavier quantisation, not less. On this silicon Q8_0 saw 2.44x (7.7 -> 18.1 t/s) against Q4_K_M's 1.81x (12.1 -> 21.2).
verified 2026-08-03 · volatility medium
Note — the record's own working
Counterintuitive, and it matters for quant selection: baseline decode here is bandwidth-saturated, so MTP's only lever — fewer memory passes per token — pays more the heavier the weights. Partially offsets the cost of a higher quant. Not yet reproduced on our own hardware; the quant x MTP matrix is queued for the lab box. Our own measured figure is the 122B at UD-Q4_K_M: 20.96 t/s single-slot baseline to 31.17 with MTP n=3 (+49%), and 36.5 t/s after tuning to n=6.