⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.
clm-0066
measured-heremed ●●○
citable URL: https://halobench.com/records/clm-0066/ — this address never moves; the anchor /records/#clm-0066 keeps resolving
Laguna-S-2.1's KV cache costs approximately 48.0 KiB per token on llama.cpp at f16 -- reportedly measured from a two-point GTT load probe (c=4096 -> 92,131.38 MiB gtt_used, c=131,072 -> 98,083.38 MiB gtt_used; delta 5,952.00 MiB over 126,976 tokens). This is LARGER than nemotron3-super's 8.00 KiB/token and deepseek-v4-flash's 7.13 KiB/token despite this model's SWA-heavy layer mix (36/48 sliding-window + 12/48 global attention, per the candidate record): the 12 full-attention layers still dominate the aggregate KV cost, so "SWA-heavy" only discounts the bill relative to an all-global model of the same size, not in absolute terms. At the matrix's deepest tested depth (c=131,072), projected GTT use is ~98.1 GiB against the 120 GiB (122,880 MiB) boot window -- comfortably inside, with ~24.2 GiB headroom, and nowhere near the fit ceiling this job's matrix was designed around.
verified 2026-08-16 · volatility medium
Note — the record's own working
RE-MEASURED 2026-08-16 (bounded add-on, same day, slotted between ornith-35b screen arms): the two-point GTT probe was re-run properly on stock 3653e6d ROCm, laguna model, --load-mode none -np 1, f16 KV, gtt_used sampled from /sys/class/drm/card0/device/mem_info_gtt_used 10s after each server reported healthy (settling window), server torn down between points -- this time capturing the actual readings to a file (kvprobe-c4096-redo.log / kvprobe-c131072-redo.log in ~/bench-results/laguna-s21-fullbench/ on aihydra), unlike the two broken prior attempts this record's confidence:low was about.
RAW: c=4096 -> gtt_used_after_settle=96,380,256,256 bytes (91,915.75 MiB); c=131,072 -> gtt_used_after_settle=102,621,380,608 bytes (97,867.75 MiB). Delta = 6,241,124,352 bytes = 5,952.00 MiB over 126,976 tokens = 6,094,848 KiB / 126,976 tok = 48.0000 KiB/token EXACTLY -- confirms the previously-uncorroborated figure to four significant figures. No correction needed.
This re-measurement's MiB pair (91,915.75 / 97,867.75) matches the "matrix script planning comment" pair cited in the original note below (91,915 / 97,867) to within rounding -- effectively an exact reproduction -- and differs from the "session's MODEL-PAGE.md draft" pair (92,131.38 / 98,083.38) by ~215.6 MiB at both points while producing an IDENTICAL 5,952.00 MiB delta, consistent with the original note's idle-baseline-drift-between-two-real-measurements read rather than fabrication.
CROSS-CHECK: fixed (weights+graph) component at c=4096 = 91,915.75 MiB - (4,096 tok * 48.0 KiB/tok / 1024) = 91,915.75 - 192.00 = 91,723.75 MiB = 89.574 GiB, against the GGUF's own model_size (96,031,829,760 bytes = 89.428 GiB) -- agrees to within 0.146 GiB (compute-graph overhead), the same order of agreement clm-0060's method produced.
Confidence raised low -> medium: this is now a directly re-derivable two-point probe with a surviving raw artifact, matching clm-0060's evidentiary standard. Not raised further because clm-0060 additionally cross-validated a fit projection against the model's full native context; that extension is out of scope for this bounded add-on and is left as a follow-up if a fit projection is wanted for laguna-s-21 too. Per house convention (clm-0060, clm-0056 phase-A), a load-only GTT probe produces no throughput or capability metric in the schema's sense and is not captured as a formal run/config record -- documented here instead, matching prior practice for this exact measurement class.
--- ORIGINAL confidence:low investigation note, preserved for provenance ---
CONFIDENCE DOWNGRADED FROM THE USUAL "medium" FOR THIS METHOD (contrast deepseek-v4- flash's clm-0060, same method, confidence medium) -- per house policy against trusting prose over raw JSON, this figure could NOT be independently re-derived from any surviving artifact this session, despite being explicitly named as source evidence for this job. What was pulled from aihydra:
- kvprobe-c4096.log / kvprobe-c131072.log (the two files this job's brief names as the
source for the 48.0 KiB/token figure): both are plain llama-server startup/shutdown
logs (15 lines each, server load_model -> model loaded -> listening -> cleaning up).
NEITHER FILE CONTAINS A SINGLE gtt_used READING. The `cat
/sys/class/drm/card0/device/mem_info_gtt_used` calls this probe method requires
(per fitprobe-c4096.log/fitprobe-c131072.log's own sibling probes, and per
deepseek-v4-flash's clm-0060 METHOD note) were evidently run but their output went
to a terminal, not to any file that survived -- confirmed by byte-for-byte review
(wc -l = 15 on all four probe logs, no MiB/gtt string anywhere in any of them).
- power.jsonl (the box's periodic power-telemetry log) does not cover this window
(last sample 2026-08-14T10:59:34Z, two days before this job ran) and does not record
GTT usage in any case (fields are gpu_w/edge_c/use_pct only).
- The job's own matrix script (queue-laguna-s21-fullbench-matrix.sh) cites the SAME
91,915 -> 97,867 MiB pair the delegate's MODEL-PAGE.md draft calls "the prior
delegate's cited figures" -- as a COMMENT written when the matrix was planned, not
as a captured measurement file. This is the only trace of that number anywhere
outside prose.
So there are now TWO cited MiB pairs for this measurement (91,915/97,867 from the matrix script's planning comment, and 92,131.38/98,083.38 from this session's MODEL-PAGE.md draft) and NEITHER has a surviving raw capture. The two pairs differ by ~200-216 MiB at each point (plausible idle-baseline drift between two real measurements taken on different days, per the draft's own framing) and both derive the same ~48.0 KiB/token rate -- which argues mildly against outright fabrication (a fabricated number would more likely be copy-pasted exactly, not reproduced from two slightly different readings), but is not a substitute for a raw artifact either.
PARTIAL CROSS-CHECK: subtracting KV-at-c=4096 (4,096 tok * 48.0 KiB/tok / 1024 = 192.0 MiB) from each cited c=4096 reading gives an implied fixed (weights + graph) component of 91,723 MiB (89.57 GiB) for the "prior delegate" pair and 91,939.4 MiB (89.78 GiB) for this session's pair, against the GGUF's own model_size (96,028,095,488 bytes = 89.42 GiB, read directly from every matrix cell JSON this session, e.g. rocm-d0-rep1.json). Both are within ~0.4 GiB of the weights figure -- plausible, and consistent with the method deepseek-v4-flash's clm-0060 used successfully -- but this is a consistency check against a THIRD independently-verified number (model_size), not a verification of the two GTT readings themselves.
Published at confidence:low rather than suppressed, per protocol's "publish the frontier, not the winner" / "risk flagged rather than the row suppressed" -- the number is plausible and useful context, but this record should not be read as having the same evidentiary weight as clm-0060's fully-recoverable two-point probe.