Home › Evidence › Records › clm-0130
⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.

clm-0130

measured-herehigh ●●●
citable URL: https://halobench.com/records/clm-0130/ — this address never moves; the anchor /records/#clm-0130 keeps resolving

The Qwen3.8-27B dense worker earns its slot on THROUGHPUT UNDER CONCURRENCY, not single-stream speed, and only on the leak-fixed build. Served single-stream decode falls off gently with depth and never cliffs: 19.2 / 17.0 / 14.7 / 12.5 t/s at 32k / 65k / 131k / 205k, prefill 239 / 141 / 74 / 50 t/s, self-spec MTP acceptance 0.76-0.84 (run-0661..run-0664, cfg-0184, single slot). Under a real 4-slot agentic workload the same config sustains ~12.8 t/s per slot at mean concurrency 3.6 — a derived aggregate of ~46 t/s of useful work (run-0659). BINARY BUILD FINDING: this multi-slot behaviour requires PR #27311. The identical flags on fresh master c5a5535e degraded to EMPTY output within ~4 minutes of 4-slot load (the #25992 cross-request leak; monitor 13:34:25Z OK -> 13:38:29Z EMPTY), whereas the pr27311 build (c530ea79c) held 0 EMPTY across the full 106-minute run (run-0660). The earlier Q8_0 serve config pinned --parallel 1 precisely to avoid #25992; pr27311 is what makes a multi-slot dense worker viable at all.

verified 2026-09-06 · volatility medium
evidence run-0659 run-0660 run-0661 run-0662 run-0663 run-0664 con-0018

Note — the record's own working

DEPTH SWEEP — served path via the slotpin proxy (drive_depth_served, code content, tokenizer-trimmed to exact depths, 256-token generations), cfg-0184, --parallel 1, pr27311 build. Monotonic, no device loss, no cliff to d204800; d262144 (model-max) was not taken on the served path (prompt+gen exceeded -c and returned HTTP 400), so the single-stream served ceiling recorded here is d204800. MULTI-SLOT — read from proxy metrics.jsonl for the τ² window (run-0658): 469 decode turns, per-slot mean 12.8 t/s (min 1.9, peak 23.7), concurrency mean 3.6 with 350 of 471 rows at 4.0. Aggregate ~46 t/s is per-slot-mean x mean-concurrency on a VARIED workload — not the shared-n-gram-pool artifact identical-prompt multi-slot would give. BUILD HAZARD — PR #27311 "Scheduler UMA ring buffer (+ sanitizer and fixes)" is an OPEN, unmerged upstream fix for the gfx1151 UMA async-output race behind #25992; it also restores speculative-decode acceptance under concurrent MTP (#27572). Carried on aihydra as branch pr27311 (con-0018). server_flags_degraded is false ON THIS BUILD; it would be true on stock master under --parallel>1, which is the whole point of the guard (run-0660). Serving this worker multi-slot on a non-pr27311 build is a data-integrity hazard, not a performance preference.

Cited by — computed at build time, never stored

model pages qwen38-27b
candidate gate history qwen38-27b