Home › Evidence › Records › clm-0133
⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.

clm-0133

measured-herehigh ●●●
citable URL: https://halobench.com/records/clm-0133/ — this address never moves; the anchor /records/#clm-0133 keeps resolving

The pwilkin candidate wins single-stream throughput decisively but only ties the worker under concurrency. Served single-stream decode: 27.98 / 23.18 / 22.12 / 17.07 t/s at 32k / 64k / 131k / 205k (cfg-0186) — a consistent +37% to +50% over the worker's 19.23 / 16.99 / 14.71 / 12.49 (cfg-0184) across every depth, monotonic with no cliff (and no Vulkan-style device loss, since this is the HIP/retained-PM4 build). But multi-slot aggregate reaches 47.3 t/s at 4 concurrent slots (run-0667) — essentially tied with the worker's ~46 (run-0659) — and it gets there by scaling only 1.3x from a high single-stream base (36.5 -> 47.3 @1->4) with very uneven per-slot rates (DFlash2 draft contention), whereas the worker scales ~2x evenly from ~22. So pwilkin front-loads latency-per-task; the worker scales with concurrency; they converge at 4-slot aggregate.

verified 2026-09-06 · volatility medium
evidence run-0667 run-0668 run-0669 run-0670 run-0671

Note — the record's own working

Single-stream depth via slotpin (drive_depth_served, code content, 256-token gens), cfg-0186, --parallel 1, DFlash2 width 3, retained-PM4 build. DFlash2 acceptance rose with depth (0.62 -> 0.79). Multi-slot: 4 varied ~4k-prompt requests fired concurrently via slotpin; aggregate decode summed across slots. The single-stream lead is real and depth-robust; the multi-slot tie is the operative fact for a worker that receives concurrent tasks. Why the lead exists is isolated in clm-0135 (it is the quant + spec, not the retained-PM4 runtime).

Cited by — computed at build time, never stored

model pages qwen38-27b
candidate gate history qwen38-27b