Home › Evidence › Records › clm-0135
⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.

clm-0135

measured-herehigh ●●●
citable URL: https://halobench.com/records/clm-0135/ — this address never moves; the anchor /records/#clm-0135 keeps resolving

The retained-PM4 runtime is a real but small, lossless dispatch optimisation — NOT the source of the pwilkin speed advantage. Running OUR worker model (UD-Q4_K_XL + self-spec MTP, cfg-0187) on the pwilkin build with retained-PM4 toggled: 20.17 t/s at 32k with it ON (run-0672) vs 19.45 OFF (run-0673) — +3.7%, identical acceptance (0.766), same binary and model so the delta is purely the runtime. That ~20 t/s ceiling for our high-capability model sits far below the pwilkin candidate's 27.98 at the same depth (clm-0133). Therefore the ~45% single-stream lead comes from the IQ4_XS quant (far fewer bytes read per token than UD-Q4_K_XL) plus the DFlash2 draft — i.e. exactly the levers that cost the 8 capability points (clm-0132) — and cannot be recovered by adopting the runtime. The retained-PM4 change is also the one piece of the pwilkin stack with no upstream PR: the author states he does not expect it to be accepted and may retire it for HRX support (con-0019); everything else is upstream PRs (e.g. #27311, con-0018).

verified 2026-09-06 · volatility low
evidence run-0672 run-0673 con-0019

Note — the record's own working

Clean isolation: cfg-0187 is our exact worker model + self-spec draft-mtp n_max 2 on the pwilkin binary + custom retained-PM4 HIP runtime; run-0672 (ENABLE_RETAINED_PM4=1 -> DEBUG_HIP_GRAPH_PM4=1, HIP graphs on) vs run-0673 (=0, graphs disabled) differ only by that env toggle. PM4-off (19.45) matches our worker's own baseline (19.23, run-0661), confirming the runtime is the only variable. Consequence for adoption: there is no free lunch — the worker's speed cannot be lifted to pwilkin's by the runtime; buying that speed means accepting the lower-capability IQ4_XS + DFlash2 stack. The remaining unexhausted lever is a higher-capability fast quant (a calibrated IQ4 that holds competence closer to Q4_K_XL), which is a quant-quality question, not a runtime/spec one.

Cited by — computed at build time, never stored

model pages qwen38-27b
candidate gate history qwen38-27b