Qwen3.8-27Bbenchedno guard
②Verdict
A 27-billion-parameter dense model with a hybrid attention layout that makes long context nearly free: the KV cache costs 64 KiB per token, 2.00 GiB at 32k, and both quants fit the full 262k declared context with more than 75 GiB of headroom clm-0056. The capability case pre-registered for it was met: 1.000 on the five-task τ² smoke against the 0.80 bar recorded before measurement, and 0.577 over the 26 tasks its capped full run completed clm-0057.
Speed is the limitation, not competence: the dense floor decodes at 7.8 t/s, speculation via the MTP head is therefore mandatory, and its 2.33x lift to roughly 18 t/s is hard-capped at n_max 3 because higher settings corrupt generations outright in live sessions — a failure this lab characterised and staged for upstream clm-0055. Backend choice is not free either: decode is backend-indifferent, but Vulkan prefill halves at working depth and the device is lost entirely at 131k, so serving this model at depth requires ROCm clm-0054. The shallow probe picked the other backend — depth reversed it, which is the standing per-model, per-phase rule doing its work clm-0050.
Energy, retroactively joined the day after this bench first published: the tau2 arm cost 38.29 Wh per correct answer, whole-session, 1.16 pence at the standing tariff clm-0057; and MTP's throughput lift does not convert to energy savings at the same ratio — mean power rises 9-23% under speculation depending on backend and quant, so the ~2.3x decode speedup buys roughly 1.8-2.1x less energy per token, not 2.3x clm-0055.
Autonomous-tuning screening evidence received later does not change this production selection. Its alternative server path lacks the paired floor and served-path cells, configuration guard, artifact identities, and energy joins required for a production recommendation. A bounded cache-reuse trace is retained as an investigation result, not a replacement capability or throughput result clm-0120 clm-0121.
WORKER PROFILE (2026-09). A second serving profile was measured for this model in the role it actually suits — a dense worker sitting beside a faster interactive model, taking each task of a project in turn. It is a genuinely different configuration, not a re-run of the Q8_0 screen: UD-Q4_K_XL, self-speculative MTP at n_max 2, four slots, reasoning_effort low, on the PR #27311 leak-fix build (cfg-0183). On this config the model scored 91.7% on tau2 airline 0-25 clm-0129 — a striking distance above the screen's capped 0.577, but across four changed fingerprint fields at once, so it is a new measurement rather than a before/after; the leading (untested) suspect is the earlier run's 4096-token cap truncating tool calls clm-0129. Throughput is where the worker earns its slot: served single-stream decode holds a no-cliff curve to 205k clm-0130, and under a real four-slot workload the ~13 t/s per-slot decode becomes ~46 t/s of aggregate useful work clm-0130 at 15.18 Wh per correct answer, well under the single-slot Q8_0 run's 38.29 clm-0131. The load-bearing dependency is PR #27311: the same flags on stock master leak empty output within minutes under concurrency (#25992), which is exactly why the 2026-08 recommendation pinned --parallel 1 clm-0130. This is now the model's recommended serving config (best.serve_config = cfg-0183, promoted 2026-09-06): for a dense worker that receives concurrent tasks it is the stronger choice on every axis measured — capability, throughput per watt, and leak-safety under load. The earlier Q8_0 single-slot profile (cfg-0066) remains on record as the highest-quality single-stream point and the origin of this model's capability, backend and n_max findings; it is superseded as the recommendation, not deleted.
THIRD-PARTY SCREEN (2026-09). The pwilkin/ilintar Strix Halo stack — IQ4_XS-imatrix + a DFlash2 draft on a custom "retained-PM4" HIP runtime — was reproduced and screened against the worker. Its author's own build reproduced faithfully (26.9 t/s vs a claimed 26.256 on a 31.5k prompt), and it is genuinely faster single-stream: +37% to +50% decode at every depth to 205k, and cheaper per correct answer (10.26 vs 15.18 Wh) clm-0133 clm-0134. But it costs 8 capability points on the identical tau2 0-25 (83.3% vs 91.7%) clm-0132, and under real concurrency its aggregate only ties the worker (~47 vs ~46 t/s at 4 slots) because it scales poorly from a high single-stream base clm-0133. The one lever that might have given the worker that speed without the capability cost — the retained-PM4 runtime — was isolated on the worker's own model and found to be ~+4% and lossless, not the source of the gain; the gain is intrinsic to the IQ4_XS quant + DFlash2, i.e. the capability trade itself clm-0135. So the worker (cfg-0183) stands as the capability-first recommendation, and the speed/quality trade is confirmed real and not circumventable by the runtime. (A separate halo-box fork was also screened: its merged build measured ~= this pwilkin stack on our models, and its headline "65.6 t/s" figure is an unbacked README claim on an unmerged branch, with no artifact — nothing to reproduce.)
③Best configuration
| model | Qwen3.8-27B-UD-Q4_K_XL.gguf · UD-Q4_K_XL · rev unsloth Dynamic V3.0 (downloaded 2026-09-03; file 17,559,178,144 bytes) |
| engine | ggml-org/llama.cpp c530ea7 · rocm · host aihydra (igpu) |
| flags | -ngl 999 -fa on -c 49152 -b 2048 -ub 2048 -ctk f16 -ctv f16 --parallel 4 --cont-batching --slots --slot-save-path /var/lib/slotpin/kv-cache --cache-ram 8192 --reasoning-effort low --spec-type draft-mtp --spec-draft-n-max 2 --jinja --metrics --host 127.0.0.1 --port 5802 (fronted by the slotpin KV-warmth proxy on :8090, slots=4 max_inflight=4) |
| sampling | temperature 0 · max_tokens 8192 |
| template | not recorded at test time |
| tree | fork — carrying con-0018 |
config record cfg-0183
every cell generated from the record at build time · throughput cells from cfg-0184 (same build c530ea7, fa on, f16 KV) · provenance: measured-here throughout
③aDecode against context depth
④Other configurations tested — each as a delta against best
| variant | Δ decode | Δ turns (paired tasks) | Δ energy | note | records |
|---|---|---|---|---|---|
| UD-Q4_K_XL quant · same backend | +42% @32k | — | -22% | The comparability arm decodes faster, as a dense model reading fewer bytes per token must — the capability screen ran on Q8_0, so the quant's own agentic score is unmeasured, not inherited. The window energy follows the same direction: fewer bytes per token costs less power for the same 3-rep window, not just less wall time. | clm-0054 |
| Vulkan backend · same quant | -0.2% @32k | — | +2.0% | Decode is a wash — the delta that matters is not in this column. Vulkan prefill collapses to roughly half of ROCm's at working depth, and at the deepest cell Vulkan lost the device on every attempt while ROCm completed the matrix. Window energy is a wash too, in step with decode. | clm-0054 run-0265 run-0266 |
| draft-mtp · n_max 3 | +133% | — | — | Not optional on this model — the floor is unusably slow — and not safe above this setting either; the ceiling is a correctness bound, not a tuning preference. | run-0268 clm-0055 |
| draft-mtp · n_max above 3 hazard | — | generations collapse to 1 token after accumulated session volume, and the broken cells report 1,000,000 t/s | The EOS cliff: the speculative path drives the target's own next-token distribution onto end-of-turn in live varied-prompt sessions. Sweeps that average unfiltered throughput flatter exactly the broken settings. | clm-0055 | |
| worker profile (UD-Q4_K_XL · self-spec MTP n_max 2 · 4-slot · pr27311) vs the Q8_0 single-slot screen | — | — | -42% | A whole-PROFILE change for the dense-worker role, not a single-lever delta, so read it as such: quant, build, reasoning_effort and slot count all move together [clm-0129]. The worker config scored 91.7% on tau2 0-25 [clm-0129], held a no-cliff served depth curve to 205k and sustained ~46 t/s aggregate under real 4-slot load [clm-0130], and cost 15.18 Wh per correct answer against the earlier single-slot Q8_0 run's 38.29 — the energy column here [clm-0131]. The lift is unlocked by PR #27311: multi-slot on stock master goes empty within minutes (#25992), which is why the earlier config pinned --parallel 1 [clm-0130]. | clm-0129 clm-0130 clm-0131 |
| pwilkin candidate (IQ4_XS-imatrix · DFlash2 draft · retained-PM4 runtime) vs the worker | +46% @32k | — | -39% | A reproduced third-party stack, screened against the worker. It wins single-stream decisively — +37% to +50% decode at every depth, the decode column here shows d32k [clm-0133] — and is cheaper per correct answer (10.26 vs 15.18 Wh, the energy column) [clm-0134]. But it costs 8 capability points (83.3% vs 91.7% on the identical tau2 0-25) [clm-0132], and its multi-slot aggregate only ties the worker (~47 vs ~46 t/s @4) because it scales poorly from a high single-stream base [clm-0133]. A genuine speed/quality trade, not a free win. Not a clean single-lever A/B (quant, spec and build all differ). DFlash2 correctness gate was clean (0 empty / 66 min). | clm-0132 clm-0133 clm-0134 |
| retained-PM4 runtime (isolated on the worker's own model) | +3.7% @32k | — | — | Running the worker's own UD-Q4_K_XL + self-spec MTP on the pwilkin build with retained-PM4 toggled shows the runtime is only ~+4% and lossless (same binary/model, identical acceptance) — NOT the source of the candidate's lead. That lead is the IQ4_XS quant + DFlash2, i.e. the capability trade; it cannot be recovered via the runtime. Retained-PM4 is also the one piece with no upstream PR (con-0019) [clm-0135]. | clm-0135 |
deltas computed at build time from the named runs, paired-task traces and metered windows — a hazard row shows words where a delta would mislead
⑤Open questions
- The EOS cliff's root cause — the target's forward pass is demonstrably corrupted at n_max ≥ 4, the MTP head's internal state is the suspect, and the upstream issue is not yet filed clm-0055.
- ANSWERED 2026-08-15, partially: a same-day cause-isolation matrix flipped backend alone (ROCm, everything else held at the failing Vulkan/Q8_0/f16 baseline) at n_max=5 and found zero anomalies across 15 requests and an extended 45-request confirmation, vs 11/15 degenerate on Vulkan in the identical session shape. Backend is the dominant single-variable gate measured so far — but this is a 15-45-request-depth result, not a production-scale one, and the underlying mechanism is still not instrumented; "ROCm doesn't show it here" is not "ROCm is immune." The n_max≤3 hard cap stands unchanged on any backend clm-0055.
- The full 50-task τ² mean — the wall-bound run completed 26 tasks; the remaining tasks and an uncapped-output arm would make the score comparable to the historical series clm-0057.
- ANSWERED 2026-08-17: d204800 and d262144 (model-max) both measured on ROCm (cfg-0103) — 61.00/51.23 pp t/s, 5.449/5.025 tg t/s, monotonic continuation of the d0/d32768/d131072 series with no cliff. Vulkan was NOT re-attempted at these depths (it was already device-lost at d131072, below this backfill round's bar of "previously survived its prior max tested depth"), so the ROCm-vs-Vulkan-at-depth comparison itself remains as it was at d131072, just extended further on the surviving backend alone.
- UD-Q5_K_XL — the community's preferred quant for this model was not screened here; the bracket around it (Q8_0 above, UD-Q4_K_XL below) is clm-0054.
- DFlash2 is now a live upstream/community lead, not a recommendation: PR #27342's comments show possible workload-specific gains over draft-MTP on Qwen3.8-27B, but also fresh hazards around multimodal/M-RoPE injection, high draft depth, and multi-agent scaling. Treat it as a future matched lab check only after the implementation and artifact path are stable; no current HaloBench number or serving verdict inherits from it clm-0090.
- ANSWERED 2026-08-20, negatively: the matched HO-004 lab check proved stock MTP n_max=3 and DFlash2 Q8_0 width 4 activation on the v0.6.5 portable Vulkan path, but every arm failed Stage A sentinel-empty-content checks on code/toolish prompts, so no DFlash2 guard, throughput, energy-efficiency, production recommendation, quality equivalence, or stock MTP n_max>=4 safety claim is admitted clm-0098.
- The worker profile's 0.577 -> 0.917 movement (run-0271 -> run-0658, same tasks 0-25, same harness/seed) is unattributed: quant, build, reasoning_effort, max_tokens and slot count all changed at once. A matched single-lever isolation — hold everything and vary only agent max_tokens (4096 vs 8192), then only reasoning_effort (medium vs low) — would say which lever carries the gain. The leading hypothesis is the output cap truncating tool calls clm-0129.
- aihydra's energy delta_w still rests on the 10.1 W canonical box-idle floor (idle-aihydra-2026-08), but a spot idle reading in the 2026-09 campaign was ~20 W, most likely the forced-high GPU power state plus the now-resident strix-halo fan-control daemon. A fresh post-fan-control idle-baseline would refine delta_w across the energy series (it does not touch wh_per_task, which uses wh_total) clm-0131.
- The one unexhausted speed lever for this worker: a higher-capability FAST quant. The pwilkin screen showed the ~45% single-stream lead comes from IQ4_XS reading far fewer bytes per token than UD-Q4_K_XL, at an 8-point capability cost, and that the runtime (retained-PM4) recovers only ~4% clm-0135. So the open question is whether a calibrated IQ4-class quant can hold capability closer to Q4_K_XL while keeping the byte-count speed win — a quant-quality experiment, not a runtime or spec one clm-0132 clm-0133.