Models › laguna-s-21
⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.

Laguna S 2.1 (poolside)benchedguard 4/4

118B / ~8B activehybrid MoE (LLM_ARCH_LAGUNA) · 36/48 sliding-window + 12/48 global attention · 118B total / ~8B active — no MTP/speculation path in llama.cppquant held: Q4_K_M · sole staged quantfirst measured 2026-08-10latest run 2026-08-17

②Verdict

A hybrid mixture-of-experts (36/48 sliding-window attention layers, 12/48 global, 118B total parameters with ~8B active) served at Q4_K_M — the only quant staged for this candidate — with no speculation path: llama.cpp's tree carries no nextn/draft- tensor wiring for LLM_ARCH_LAGUNA.

Backend is not a close call: this job's own throughput matrix has ROCm sweeping every cell clean (300.7/23.68 t/s at d0 falling to 116.43/7.19 t/s at d131072) while Vulkan survives d0 and d32768 already behind on both phases and then loses the GPU device on both allowed attempts at d131072, kernel-evidenced amdgpu ring resets independently re-pulled after the job's own dmesg capture for that cell turned up empty from a script bug clm-0065 — unlike nemotron3-super and deepseek-v4-flash (both Vulkan-favoured on this box), ROCm wins outright here and is the backend this bench serves the capability arm on. The KV cache costs an estimated 48.0 KiB/token — larger than nemotron3-super's 8.00 and deepseek-v4-flash's 7.13 despite the SWA-heavy layer mix, because the 12 full-attention layers still dominate the aggregate bill — but this figure carries lowered confidence: the underlying two-point sysfs readings were not recoverable from any artifact this job produced, only from prose, so it is published flagged rather than suppressed clm-0066.

Capability clears the standard bar cleanly: the FULL 26-task tau2 airline run completed end to end (rc=0, not wall-bound) at 0.6923 mean reward (18/26), 187 tool-call messages, 0 empty turns — VALID under the SMOKE gate clm-0067. Energy is the standout finding of this bench: 7.2588 Wh per correct answer, 0.22 pence at the standing tariff — lower than the cited same-protocol records for deepseek-v4-flash, nemotron3-super and qwen38-27b, achieved with no speculation at all against a raw ~15.5 t/s decode floor at the serving depth — a small active-parameter footprint (~8B of 117.6B total) again outrunning faster dense or speculated decode on this metric, the same pattern deepseek-v4-flash's MoE showed against qwen38-27b's dense-plus-MTP arm clm-0067.

This same-protocol energy claim was screened against the then-published benched models' Wh-per-correct-answer records before being published; the one lower figure on the site at publication time (qwen35-122b, a 5-task smoke window that pre-dates this programme's standard-26-task convention) is a different scale and protocol, not a like-for-like comparator, and this verdict does not claim to beat it.

verdict written 2026-08-16 · every number above stands next to the claim chip that carries it

③Best configuration

modellaguna-s-2.1-Q4_K_M.gguf · Q4_K_M
engineggml-org/llama.cpp 3653e6d · rocm · host aihydra (igpu)
flags-ngl 999 -fa on -c 32768 -dev ROCm0 --load-mode none --parallel 1 --jinja --reasoning-format deepseek -n 4096 --slots
templatenot recorded at test time
treeupstream — stock

config record cfg-0092

decode @ 0
23.68 t/s
decode @ 32k
15.53 t/s
prefill @ 0
300.67 t/s
prefill @ 32k
215.11 t/s
τ² airline · thinking off
0.692 ±0.177
run-0329 · passed 18/26
Wh per correct answer (τ², thinking off)
7.26 Wh
eng-0124 run-0329 · 0.22 p per answer

every cell generated from the record at build time · throughput cells from cfg-0090 (same build 3653e6d, fa on, f16 KV) · provenance: measured-here throughout

③aDecode against context depth

ROCmVulkan
0510152025032k65k131k204.8kdecode t/scontext depth (tokens)serving context 32k23.68 t/s @ depth 0 · run-031815.53 t/s @ depth 32k · run-03207.19 t/s @ depth 131k · run-03224.91 t/s @ depth 204.8k · run-0392 · CV 0% · N=113.29 t/s @ depth 0 · run-032412.27 t/s @ depth 32k · run-0326Vulkan · 12.27ROCm · 4.91
1 rep per cell · max CV 0.0% · build 3653e6d / 3653e6d · The Vulkan line ends at d32768 — unlike deepseek-v4-flash, Vulkan survives this candidate's serving depth, but every depth from 131,072 tokens on lost the GPU device on both allowed attempts and is recorded as a failed run, not a point. ROCm is the only backend that completed the deepest matrix cell. A 2026-08-17 depth-ladder extension pushed ROCm further to d204800 (81.95 pp / 4.91 tg t/s, clean); this candidate's true model-max (1,048,576) was NOT attempted — the fit projection (89.41 GiB weights + 48.0 GiB KV = 137.41 GiB) exceeds this box's 120.0 GiB GTT budget by ~17.4 GiB, a recorded refusal, not a silent skip. · records: run-0318 run-0320 run-0322 run-0392 run-0324 run-0326

④Other configurations tested — each as a delta against best

variantΔ decodeΔ turns (paired tasks)Δ energynoterecords
Vulkan backend · same quant-21% @32k—+38%Unlike nemotron3-super (the one candidate in this programme where Vulkan won decode outright), ROCm is ahead on both phases at every depth both backends completed — decode +26% at d32768, prefill +317% — and Vulkan is the backend that loses the GPU device at this candidate's next depth step. This is not a close call the way it was for nemotron3-super or deepseek-v4-flash. clm-0065
Backend requirement past the serving depth hazard—Vulkan lost the GPU device on both allowed attempts at d131072 (rc=134, vk::DeviceLostError), a fourth confirmed occurrence of this device-loss class on this silicon. The job's own per-cell dmesg capture for this cell is empty — a `sudo`-omission bug in the matrix script, not evidence the failure did not happen — so the kernel ring-reset evidence was independently re-pulled this session and confirms the class exactly. Serving this model at any context deeper than this bench's tested 32,768-token band should assume ROCm, not Vulkan, until Vulkan is re-tested past d32768 on a fixed build. clm-0065
KV cache cost (fit) hazard—The 48.0 KiB/token figure this page's fit arithmetic rests on could not be independently re-derived from any surviving artifact this session — the raw two-point GTT sysfs readings behind it exist only in prose (this session's and an earlier delegate's), never in a captured log. Published at lowered confidence, not suppressed. Larger than nemotron3-super's 8.00 and deepseek-v4-flash's 7.13 KiB/token despite this candidate's SWA-heavy layer mix (36/48 sliding-window) — the 12 full-attention layers still dominate the aggregate KV bill. Even so, projected GTT use at this bench's deepest tested depth (c=131,072) is ~98.1 GiB against the 120 GiB boot window, comfortably inside with ~24.2 GiB headroom. clm-0066
Wh-per-correct-answer versus cited same-protocol records———7.26 Wh/correct is lower than the cited same-protocol full standard-26-task tau2 records for deepseek-v4-flash (24.29), nemotron3-super (32.43) and qwen38-27b (38.29). One lower number exists on the site (qwen35-122b, 5.71 Wh/correct) but from a 5-task smoke window that pre-dates the standard-26-task convention — not a like-for-like comparison, and not what this row claims to beat. clm-0067

deltas computed at build time from the named runs, paired-task traces and metered windows — a hazard row shows words where a delta would mislead

⑤Open questions

⑥Provenance

bench host aihydra · rocm · ggml-org/llama.cpp 3653e6d
discipline scatter published per cell (max CV 0.0%) · guard or written waiver on every performance series
window 2026-08-10 → 2026-08-17