Home › Evidence › Records › clm-0007

clm-0007

inferredlow ●○○
citable URL: https://halobench.com/records/clm-0007/ — this address never moves; the anchor /records/#clm-0007 keeps resolving

A GPU cache wall around 32-40 MB, past which read speed drops roughly 4x, is a candidate explanation for our own prefill degradation from ~354 tok/s at small context to ~144 tok/s at 128K.

verified 2026-08-03 · volatility medium
evidence run-0003

Note — the record's own working

The wall was measured on another 128GB Strix Halo box with a purpose-built benchmark; the originating theory about where the collapse should begin did NOT match the measured curves, so the mechanism is real but not yet modelled. Our ~2.5x prefill drop is the right order for a ~4x bandwidth cliff partially amortised. TWO REASONS TO BE CAUTIOUS: this is a GPU-cache effect, entirely distinct from the pool-utilisation metric in clm-0001 — conflating them would be a mistake; and the reporter's dense model collapsed while his MoE shrugged the same conditions off, which is unexplained. We run MoE, so the q4_0 mitigation that helped his dense model may not transfer. Falsifiable: run the same cache benchmark on aibeast and compare the knee against our measured prefill curve.

Cited by — computed at build time, never stored