Models › qwen38-27b
⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.

Qwen3.8-27Bbenchedno guard

27B densedense · hybrid attention (16 of 64 layers carry KV) · MTP headquant held: Q8_0 (unsloth) · UD-Q4_K_XL comparability armfirst measured 2026-08-15latest run 2026-09-06

②Verdict

A 27-billion-parameter dense model with a hybrid attention layout that makes long context nearly free: the KV cache costs 64 KiB per token, 2.00 GiB at 32k, and both quants fit the full 262k declared context with more than 75 GiB of headroom clm-0056. The capability case pre-registered for it was met: 1.000 on the five-task τ² smoke against the 0.80 bar recorded before measurement, and 0.577 over the 26 tasks its capped full run completed clm-0057.

Speed is the limitation, not competence: the dense floor decodes at 7.8 t/s, speculation via the MTP head is therefore mandatory, and its 2.33x lift to roughly 18 t/s is hard-capped at n_max 3 because higher settings corrupt generations outright in live sessions — a failure this lab characterised and staged for upstream clm-0055. Backend choice is not free either: decode is backend-indifferent, but Vulkan prefill halves at working depth and the device is lost entirely at 131k, so serving this model at depth requires ROCm clm-0054. The shallow probe picked the other backend — depth reversed it, which is the standing per-model, per-phase rule doing its work clm-0050.

Energy, retroactively joined the day after this bench first published: the tau2 arm cost 38.29 Wh per correct answer, whole-session, 1.16 pence at the standing tariff clm-0057; and MTP's throughput lift does not convert to energy savings at the same ratio — mean power rises 9-23% under speculation depending on backend and quant, so the ~2.3x decode speedup buys roughly 1.8-2.1x less energy per token, not 2.3x clm-0055.

Autonomous-tuning screening evidence received later does not change this production selection. Its alternative server path lacks the paired floor and served-path cells, configuration guard, artifact identities, and energy joins required for a production recommendation. A bounded cache-reuse trace is retained as an investigation result, not a replacement capability or throughput result clm-0120 clm-0121.

WORKER PROFILE (2026-09). A second serving profile was measured for this model in the role it actually suits — a dense worker sitting beside a faster interactive model, taking each task of a project in turn. It is a genuinely different configuration, not a re-run of the Q8_0 screen: UD-Q4_K_XL, self-speculative MTP at n_max 2, four slots, reasoning_effort low, on the PR #27311 leak-fix build (cfg-0183). On this config the model scored 91.7% on tau2 airline 0-25 clm-0129 — a striking distance above the screen's capped 0.577, but across four changed fingerprint fields at once, so it is a new measurement rather than a before/after; the leading (untested) suspect is the earlier run's 4096-token cap truncating tool calls clm-0129. Throughput is where the worker earns its slot: served single-stream decode holds a no-cliff curve to 205k clm-0130, and under a real four-slot workload the ~13 t/s per-slot decode becomes ~46 t/s of aggregate useful work clm-0130 at 15.18 Wh per correct answer, well under the single-slot Q8_0 run's 38.29 clm-0131. The load-bearing dependency is PR #27311: the same flags on stock master leak empty output within minutes under concurrency (#25992), which is exactly why the 2026-08 recommendation pinned --parallel 1 clm-0130. This is now the model's recommended serving config (best.serve_config = cfg-0183, promoted 2026-09-06): for a dense worker that receives concurrent tasks it is the stronger choice on every axis measured — capability, throughput per watt, and leak-safety under load. The earlier Q8_0 single-slot profile (cfg-0066) remains on record as the highest-quality single-stream point and the origin of this model's capability, backend and n_max findings; it is superseded as the recommendation, not deleted.

THIRD-PARTY SCREEN (2026-09). The pwilkin/ilintar Strix Halo stack — IQ4_XS-imatrix + a DFlash2 draft on a custom "retained-PM4" HIP runtime — was reproduced and screened against the worker. Its author's own build reproduced faithfully (26.9 t/s vs a claimed 26.256 on a 31.5k prompt), and it is genuinely faster single-stream: +37% to +50% decode at every depth to 205k, and cheaper per correct answer (10.26 vs 15.18 Wh) clm-0133 clm-0134. But it costs 8 capability points on the identical tau2 0-25 (83.3% vs 91.7%) clm-0132, and under real concurrency its aggregate only ties the worker (~47 vs ~46 t/s at 4 slots) because it scales poorly from a high single-stream base clm-0133. The one lever that might have given the worker that speed without the capability cost — the retained-PM4 runtime — was isolated on the worker's own model and found to be ~+4% and lossless, not the source of the gain; the gain is intrinsic to the IQ4_XS quant + DFlash2, i.e. the capability trade itself clm-0135. So the worker (cfg-0183) stands as the capability-first recommendation, and the speed/quality trade is confirmed real and not circumventable by the runtime. (A separate halo-box fork was also screened: its merged build measured ~= this pwilkin stack on our models, and its headline "65.6 t/s" figure is an unbacked README claim on an unmerged branch, with no artifact — nothing to reproduce.)

verdict written 2026-09-06 · every number above stands next to the claim chip that carries it

③Best configuration

modelQwen3.8-27B-UD-Q4_K_XL.gguf · UD-Q4_K_XL · rev unsloth Dynamic V3.0 (downloaded 2026-09-03; file 17,559,178,144 bytes)
engineggml-org/llama.cpp c530ea7 · rocm · host aihydra (igpu)
flags-ngl 999 -fa on -c 49152 -b 2048 -ub 2048 -ctk f16 -ctv f16 --parallel 4 --cont-batching --slots --slot-save-path /var/lib/slotpin/kv-cache --cache-ram 8192 --reasoning-effort low --spec-type draft-mtp --spec-draft-n-max 2 --jinja --metrics --host 127.0.0.1 --port 5802 (fronted by the slotpin KV-warmth proxy on :8090, slots=4 max_inflight=4)
samplingtemperature 0 · max_tokens 8192
templatenot recorded at test time
treefork — carrying con-0018

config record cfg-0183

decode @ 32k
19.23 t/s
prefill @ 32k
238.90 t/s
draft-mtp decode speedup
2.33×
run-0268 · 18.25 vs 7.82 t/s floor
τ² airline · reasoning low
0.917 ±0.111
run-0658 · passed 22/24
Wh per correct answer (τ², reasoning low)
15.18 Wh
eng-0278 run-0658 run-0659 · 0.46 p per answer

every cell generated from the record at build time · throughput cells from cfg-0184 (same build c530ea7, fa on, f16 KV) · provenance: measured-here throughout

③aDecode against context depth

ROCm · Q8_0ROCm · UD-Q4_K_XLVulkan · Q8_0
051015065k131k204.8k262.1kdecode t/scontext depth (tokens)serving context 32k7.81 t/s @ depth 0 · run-0246 · CV 0% · N=37.29 t/s @ depth 32k · run-0248 · CV 0% · N=36.11 t/s @ depth 131k · run-0250 · CV 0% · N=35.45 t/s @ depth 204.8k · run-0374 · CV 0% · N=15.03 t/s @ depth 262.1k · run-0376 · CV 0% · N=111.44 t/s @ depth 0 · run-0252 · CV 0.1% · N=310.36 t/s @ depth 32k · run-0254 · CV 0% · N=38.14 t/s @ depth 131k · run-0256 · CV 0% · N=37.86 t/s @ depth 0 · run-0258 · CV 0% · N=37.27 t/s @ depth 32k · run-0260 · CV 0% · N=3ROCm · UD-Q4_K_XL · 8.14Vulkan · Q8_0 · 7.27ROCm · Q8_0 · 5.03
1/3 reps per cell · max CV 0.1% · build 3653e6d / 3653e6d · The Vulkan line ends where the record does: its deepest cells lost the GPU device on every attempt and are recorded as failed runs, not points. The ROCm · Q8_0 line was extended 2026-08-17 (depth-ladder backfill) to d204800 and d262144 (this candidate's real declared context ceiling, not a YaRN projection) — 61.00/51.23 pp t/s and 5.449/5.025 tg t/s, monotonic with no cliff. ROCm's lead past d131072 is now measured on this model, not extrapolated from peers. · records: run-0246 run-0248 run-0250 run-0374 run-0376 run-0252 run-0254 run-0256 run-0258 run-0260

④Other configurations tested — each as a delta against best

variantΔ decodeΔ turns (paired tasks)Δ energynoterecords
UD-Q4_K_XL quant · same backend+42% @32k—-22%The comparability arm decodes faster, as a dense model reading fewer bytes per token must — the capability screen ran on Q8_0, so the quant's own agentic score is unmeasured, not inherited. The window energy follows the same direction: fewer bytes per token costs less power for the same 3-rep window, not just less wall time. clm-0054
Vulkan backend · same quant-0.2% @32k—+2.0%Decode is a wash — the delta that matters is not in this column. Vulkan prefill collapses to roughly half of ROCm's at working depth, and at the deepest cell Vulkan lost the device on every attempt while ROCm completed the matrix. Window energy is a wash too, in step with decode. clm-0054 run-0265 run-0266
draft-mtp · n_max 3+133%——Not optional on this model — the floor is unusably slow — and not safe above this setting either; the ceiling is a correctness bound, not a tuning preference. run-0268 clm-0055
draft-mtp · n_max above 3 hazard—generations collapse to 1 token after accumulated session volume, and the broken cells report 1,000,000 t/s The EOS cliff: the speculative path drives the target's own next-token distribution onto end-of-turn in live varied-prompt sessions. Sweeps that average unfiltered throughput flatter exactly the broken settings. clm-0055
worker profile (UD-Q4_K_XL · self-spec MTP n_max 2 · 4-slot · pr27311) vs the Q8_0 single-slot screen——-42%A whole-PROFILE change for the dense-worker role, not a single-lever delta, so read it as such: quant, build, reasoning_effort and slot count all move together [clm-0129]. The worker config scored 91.7% on tau2 0-25 [clm-0129], held a no-cliff served depth curve to 205k and sustained ~46 t/s aggregate under real 4-slot load [clm-0130], and cost 15.18 Wh per correct answer against the earlier single-slot Q8_0 run's 38.29 — the energy column here [clm-0131]. The lift is unlocked by PR #27311: multi-slot on stock master goes empty within minutes (#25992), which is why the earlier config pinned --parallel 1 [clm-0130]. clm-0129 clm-0130 clm-0131
pwilkin candidate (IQ4_XS-imatrix · DFlash2 draft · retained-PM4 runtime) vs the worker+46% @32k—-39%A reproduced third-party stack, screened against the worker. It wins single-stream decisively — +37% to +50% decode at every depth, the decode column here shows d32k [clm-0133] — and is cheaper per correct answer (10.26 vs 15.18 Wh, the energy column) [clm-0134]. But it costs 8 capability points (83.3% vs 91.7% on the identical tau2 0-25) [clm-0132], and its multi-slot aggregate only ties the worker (~47 vs ~46 t/s @4) because it scales poorly from a high single-stream base [clm-0133]. A genuine speed/quality trade, not a free win. Not a clean single-lever A/B (quant, spec and build all differ). DFlash2 correctness gate was clean (0 empty / 66 min). clm-0132 clm-0133 clm-0134
retained-PM4 runtime (isolated on the worker's own model)+3.7% @32k——Running the worker's own UD-Q4_K_XL + self-spec MTP on the pwilkin build with retained-PM4 toggled shows the runtime is only ~+4% and lossless (same binary/model, identical acceptance) — NOT the source of the candidate's lead. That lead is the IQ4_XS quant + DFlash2, i.e. the capability trade; it cannot be recovered via the runtime. Retained-PM4 is also the one piece with no upstream PR (con-0019) [clm-0135]. clm-0135

deltas computed at build time from the named runs, paired-task traces and metered windows — a hazard row shows words where a delta would mislead

⑤Open questions

⑥Provenance

bench host aihydra · rocm · ggml-org/llama.cpp c530ea7
discipline scatter published per cell (max CV 0.0%) · guard or written waiver on every performance series
window 2026-08-15 → 2026-09-06