⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.
clm-0132
measured-herehigh ●●●
citable URL: https://halobench.com/records/clm-0132/ — this address never moves; the anchor /records/#clm-0132 keeps resolving
The reproduced pwilkin/ilintar Strix Halo stack (IQ4_XS-imatrix + DFlash2 draft on the retained-PM4 runtime, cfg-0185) scored 83.3% on tau2 airline 0-25 — 20 of 24 scored passed, 4-slot, reasoning low, 2 cloud-user-sim infra errors excluded — versus the worker's 91.7% (run-0658) on the identical task set, harness (668d3bc), seed, reasoning and concurrency. So the worker holds an 8-point capability lead. The DFlash2 correctness gate was clean: 0 empty outputs across 16 probes over the 66-minute run (run-0666), so the prior DFlash2 sentinel-empty concern did not recur at --parallel 4. This is NOT a clean single-lever A/B: quant (UD-Q4_K_XL -> IQ4_XS-imatrix), speculation (self-spec MTP -> DFlash2 draft model) and build/runtime all differ at once. Separately, the build reproduced the author's own headline (31.5k prompt, 256 gen, DFlash2 width 3): decode 26.9 t/s vs the claimed 26.256, prefill 325.7 vs 256.84, acceptance 0.573 vs 0.607.