Laguna S 2.1 (poolside)benchedguard 4/4
②Verdict
A hybrid mixture-of-experts (36/48 sliding-window attention layers, 12/48 global, 118B total parameters with ~8B active) served at Q4_K_M — the only quant staged for this candidate — with no speculation path: llama.cpp's tree carries no nextn/draft- tensor wiring for LLM_ARCH_LAGUNA.
Backend is not a close call: this job's own throughput matrix has ROCm sweeping every cell clean (300.7/23.68 t/s at d0 falling to 116.43/7.19 t/s at d131072) while Vulkan survives d0 and d32768 already behind on both phases and then loses the GPU device on both allowed attempts at d131072, kernel-evidenced amdgpu ring resets independently re-pulled after the job's own dmesg capture for that cell turned up empty from a script bug clm-0065 — unlike nemotron3-super and deepseek-v4-flash (both Vulkan-favoured on this box), ROCm wins outright here and is the backend this bench serves the capability arm on. The KV cache costs an estimated 48.0 KiB/token — larger than nemotron3-super's 8.00 and deepseek-v4-flash's 7.13 despite the SWA-heavy layer mix, because the 12 full-attention layers still dominate the aggregate bill — but this figure carries lowered confidence: the underlying two-point sysfs readings were not recoverable from any artifact this job produced, only from prose, so it is published flagged rather than suppressed clm-0066.
Capability clears the standard bar cleanly: the FULL 26-task tau2 airline run completed end to end (rc=0, not wall-bound) at 0.6923 mean reward (18/26), 187 tool-call messages, 0 empty turns — VALID under the SMOKE gate clm-0067. Energy is the standout finding of this bench: 7.2588 Wh per correct answer, 0.22 pence at the standing tariff — lower than the cited same-protocol records for deepseek-v4-flash, nemotron3-super and qwen38-27b, achieved with no speculation at all against a raw ~15.5 t/s decode floor at the serving depth — a small active-parameter footprint (~8B of 117.6B total) again outrunning faster dense or speculated decode on this metric, the same pattern deepseek-v4-flash's MoE showed against qwen38-27b's dense-plus-MTP arm clm-0067.
This same-protocol energy claim was screened against the then-published benched models' Wh-per-correct-answer records before being published; the one lower figure on the site at publication time (qwen35-122b, a 5-task smoke window that pre-dates this programme's standard-26-task convention) is a different scale and protocol, not a like-for-like comparator, and this verdict does not claim to beat it.
③Best configuration
| model | laguna-s-2.1-Q4_K_M.gguf · Q4_K_M |
| engine | ggml-org/llama.cpp 3653e6d · rocm · host aihydra (igpu) |
| flags | -ngl 999 -fa on -c 32768 -dev ROCm0 --load-mode none --parallel 1 --jinja --reasoning-format deepseek -n 4096 --slots |
| template | not recorded at test time |
| tree | upstream — stock |
config record cfg-0092
every cell generated from the record at build time · throughput cells from cfg-0090 (same build 3653e6d, fa on, f16 KV) · provenance: measured-here throughout
③aDecode against context depth
④Other configurations tested — each as a delta against best
| variant | Δ decode | Δ turns (paired tasks) | Δ energy | note | records |
|---|---|---|---|---|---|
| Vulkan backend · same quant | -21% @32k | — | +38% | Unlike nemotron3-super (the one candidate in this programme where Vulkan won decode outright), ROCm is ahead on both phases at every depth both backends completed — decode +26% at d32768, prefill +317% — and Vulkan is the backend that loses the GPU device at this candidate's next depth step. This is not a close call the way it was for nemotron3-super or deepseek-v4-flash. | clm-0065 |
| Backend requirement past the serving depth hazard | — | Vulkan lost the GPU device on both allowed attempts at d131072 (rc=134, vk::DeviceLostError), a fourth confirmed occurrence of this device-loss class on this silicon. The job's own per-cell dmesg capture for this cell is empty — a `sudo`-omission bug in the matrix script, not evidence the failure did not happen — so the kernel ring-reset evidence was independently re-pulled this session and confirms the class exactly. | Serving this model at any context deeper than this bench's tested 32,768-token band should assume ROCm, not Vulkan, until Vulkan is re-tested past d32768 on a fixed build. | clm-0065 | |
| KV cache cost (fit) hazard | — | The 48.0 KiB/token figure this page's fit arithmetic rests on could not be independently re-derived from any surviving artifact this session — the raw two-point GTT sysfs readings behind it exist only in prose (this session's and an earlier delegate's), never in a captured log. Published at lowered confidence, not suppressed. | Larger than nemotron3-super's 8.00 and deepseek-v4-flash's 7.13 KiB/token despite this candidate's SWA-heavy layer mix (36/48 sliding-window) — the 12 full-attention layers still dominate the aggregate KV bill. Even so, projected GTT use at this bench's deepest tested depth (c=131,072) is ~98.1 GiB against the 120 GiB boot window, comfortably inside with ~24.2 GiB headroom. | clm-0066 | |
| Wh-per-correct-answer versus cited same-protocol records | — | — | — | 7.26 Wh/correct is lower than the cited same-protocol full standard-26-task tau2 records for deepseek-v4-flash (24.29), nemotron3-super (32.43) and qwen38-27b (38.29). One lower number exists on the site (qwen35-122b, 5.71 Wh/correct) but from a 5-task smoke window that pre-dates the standard-26-task convention — not a like-for-like comparison, and not what this row claims to beat. | clm-0067 |
deltas computed at build time from the named runs, paired-task traces and metered windows — a hazard row shows words where a delta would mislead
⑤Open questions
- Whether Vulkan survives past d32768 on a fixed build — the matrix only tested one depth step past the serving context before the device-loss cap engaged; "clean at d32768" is not "immune at depth" for this candidate any more than it was for qwen38-27b or deepseek-v4-flash clm-0065.
- A real two-point GTT capture for the KV-cost-per-token figure — this bench's own kvprobe logs contain no gtt_used reading at all, only server load/shutdown lines. The 48.0 KiB/token figure rests entirely on prose from this session and an earlier delegate; a fresh probe with its raw sysfs output actually captured to a file would let this claim move from confidence:low to confidence:medium clm-0066.
- Root cause of the 8 failed tasks (1, 5, 7, 10, 14, 16, 21, 23) — no shared pattern in tool-call count or duration is visible from this bench's own summary; a transcript read would be needed to say more clm-0067.
- Whether this candidate's efficiency lead holds at a deeper serving context — this bench's capability run and throughput matrix both stayed at the 32,768-token band the 2026-08-10 screen already characterised; the matrix's own d131072 cell only measured throughput, not a capability or energy result at that depth.
- Cross-quant comparison — Q4_K_M is the only quant staged for this candidate; no higher- or lower-precision arm exists to compare against.