clm-0003
measured-heremed ●●○
citable URL: https://halobench.com/records/clm-0003/ — this address never moves; the anchor /records/#clm-0003 keeps resolving
MTP's speedup tracks how predictable the text is: +81% on code, +21% on freeform, with the real agent workload mix landing around +29-40%.
verified 2026-07-20 · volatility low
evidence run-0002
Note — the record's own working
The headline "+50%" is the favourable end of a range, not a constant. Draft acceptance measured at 75%, averaging ~5.0 tokens per speculative call. Prefill paid a ~10% tax, discounted in practice by a ~90% cache-hit rate — so the tax lands only on cold turns. Single-slot cost zero decode, which is what made the architecture viable.
Cited by — computed at build time, never stored
model pages qwen35-122b