Home › Evidence › Records › clm-0132
⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.

clm-0132

measured-herehigh ●●●
citable URL: https://halobench.com/records/clm-0132/ — this address never moves; the anchor /records/#clm-0132 keeps resolving

The reproduced pwilkin/ilintar Strix Halo stack (IQ4_XS-imatrix + DFlash2 draft on the retained-PM4 runtime, cfg-0185) scored 83.3% on tau2 airline 0-25 — 20 of 24 scored passed, 4-slot, reasoning low, 2 cloud-user-sim infra errors excluded — versus the worker's 91.7% (run-0658) on the identical task set, harness (668d3bc), seed, reasoning and concurrency. So the worker holds an 8-point capability lead. The DFlash2 correctness gate was clean: 0 empty outputs across 16 probes over the 66-minute run (run-0666), so the prior DFlash2 sentinel-empty concern did not recur at --parallel 4. This is NOT a clean single-lever A/B: quant (UD-Q4_K_XL -> IQ4_XS-imatrix), speculation (self-spec MTP -> DFlash2 draft model) and build/runtime all differ at once. Separately, the build reproduced the author's own headline (31.5k prompt, 256 gen, DFlash2 width 3): decode 26.9 t/s vs the claimed 26.256, prefill 325.7 vs 256.84, acceptance 0.573 vs 0.607.

verified 2026-09-06 · volatility medium
evidence run-0665 run-0666

Note — the record's own working

METHOD — tau2-bench airline, harness 668d3bc, --seed 42 --num-trials 1 --task-ids 0..25 --max-concurrency 4 --max-steps 100, user-sim openrouter/anthropic/claude-haiku-4.5 (temp 0). Agent: pwilkin llama.cpp d3b5cc4 + custom retained-PM4 runtime, IQ4_XS-imatrix target + DFlash2 draft, temperature 0 / max_tokens 8192, reasoning_effort low. Direct :5803 (capability is proxy-invariant, so comparable to the worker's proxy run-0658). Window 09:27:14-10:33:13Z. Reproduction of the author headline used repro_bench at a 31,497-token prompt with DFlash2 width 3. The 8-point gap is genuine capability, not degradation (0 empties, no max-steps cuts): the IQ4_XS-imatrix quant and/or DFlash2 spec path cost competence that the worker's UD-Q4_K_XL + self-spec MTP retains. Energy: 10.26 Wh per correct answer (eng-0280), clm-0134. Throughput: clm-0133.

Cited by — computed at build time, never stored

model pages qwen38-27b
candidate gate history qwen38-27b