clm-0058
measured-heremed ●●○
citable URL: https://halobench.com/records/clm-0058/ — this address never moves; the anchor /records/#clm-0058 keeps resolving
clm-0035's retracted quality-drop story does not reappear at n=14 on the PATCHED (ce7689f) build: task-matched against an f16 control, quantised KV on Qwen3.6-35B-A3B-UD-Q4_K_XL scores q8_0 0.857 against f16 0.786 — q8_0 AHEAD, not behind. 13 of the 14 matched tasks scored identically in both arms; the single exception is the one f16 failed and q8_0 passed. Combined with clm-0022's speed result, quantised KV on the patched build costs nothing measurable in speed or quality on the models measured so far.
Note — the record's own working
METHOD — hard-split, task-matched tau2 airline arms on the same server binary (`~/src/llama.cpp-kvfix/build/bin`, commit ce7689f — the KV-dequant cherry-pick behind clm-0022's speed result), differing only in `-ctk`/`-ctv` (f16 vs q8_0). Qwen3.6-35B-A3B-UD-Q4_K_XL, ROCm 7.1.0, `-fa on -c 32768 --parallel 1 --load-mode none`, pinned Haiku-4.5 simulator (openrouter/anthropic, temperature 0), seed 42. `bench/queue-kv-quality-patched.sh`'s design: the run window is split HARD in half up front, f16 runs first against a 50-task target list capped at its slice, and whatever it completed and scored (14 of 16 sims — tasks 8 and 14 hit `infrastructure_error`) becomes the EXACT task-id list q8_0 then runs. Both arms are task-matched by construction, not post-hoc intersection — cfg-0070/run-0272 (f16), cfg-0071/run-0273 (q8_0).
| arm | scored | mean_reward | paired_mean | tool_call_msgs | infra_err | verdict | |---|---|---|---|---|---|---| | f16 | 14/16 | 0.786 | 0.786 | 110 | 2 (tasks 8, 14) | VALID | | q8_0 | 14/14 | 0.857 | 0.857 | 83 | 0 | VALID |
Both arms pass the SMOKE validity gate (protocol.json): nonzero tool-call messages, zero empty assistant turns, so neither mean is a do-nothing artifact (armcompare.json, aihydra `~/bench-results/kv-quality-patched-20260815T052214Z/`).
**13 of the 14 matched tasks scored identically across arms** (task 15 failed in both; task 7 failed in both; the other 11 passed in both). The sole exception is **task 4**: f16 scored 0.0 (`user_stop`, 9 tool calls, 176.5s), q8_0 scored 1.0 (`user_stop`, 7 tool calls, 45.7s) — q8_0 passed the task f16 failed, not the reverse. That single task is the entire gap between the two paired means (0.857 − 0.786 = 0.071, exactly 1/14).
**THIS IS THE OTHER HALF OF clm-0022, AND IT REVERSES clm-0035's DIRECTION.** clm-0022 established the dequant patch restores q8_0 KV *speed* to parity with (and past) f16, on the STOCK-vs-patched 122B comparison. clm-0035 measured a quality cost for q8_0 on the STOCK build (5-task tau2, f16 1.00 vs q8_0 0.60) — a finding clm-0036 then retracted for being fully inside this harness's ~0.40 run-to-run noise band at n=5. Nobody had run the quality question on the PATCHED build until now. At n=14, patched q8_0 does not merely avoid costing quality — it is nominally ahead of its own f16 control, by a margin the size of one task. **Combined with clm-0022, the reading is: on the patched build, quantised KV costs nothing measurable in either speed or quality — on the models measured so far.**
CAVEATS, stated plainly: - **n=14, one task's worth of the whole result.** A single matched task is
0.071 of the paired mean — the entire f16/q8_0 gap here is one task's outcome,
not a distributed trend. This is not the n=5/noise-band problem clm-0036
diagnosed (14 is closer to a decision than 5 is), but it is still a single
trial per task, not a repeated-arm measurement, and should be read as
suggestive rather than dispositive on its own.
- **Different subject than clm-0035, deliberately.** clm-0035's retracted
quality-drop story was measured on Qwen3.5-122B-A10B; this run substitutes
Qwen3.6-35B-A3B for time (the box was needed for a release window,
`bench/queue-kv-quality-patched.sh`'s header). This is a NEW measurement on a
smaller model, not a replication of clm-0035 on the patched build — the
122B's own patched-build KV quality is still unmeasured.
- **max_tokens-capped, not comparable to the historical uncapped tau2 series.**
Both arms ran with `max_tokens: 4096` (tau2 `--agent-llm-args`, mirrored by
`-n 4096` server-side) after an earlier attempt on this same suite ran an
agent turn unbounded into a stuck `<think>` block and stalled the queue. The
cap is identical across both arms here, so the f16-vs-q8_0 comparison itself
is unaffected, but neither arm's raw mean is directly comparable to
clm-0035/clm-0036's uncapped numbers or to the historical tau2 series
generally.
- **f16's two infrastructure_errors are the known empty-content harness quirk,
not a KV effect.** Tasks 8 and 14 both failed with zero tool calls and zero
duration — the same harness-vs-serving-config discrepancy class clm-0053
documents (trust the harness only as far as it agrees with the serving
configuration it claims to describe). They are excluded from both
`tasks_total` and the paired comparison, not counted against f16's score.
ENERGY (retroactive join, computed 2026-08-15, eng-0082/eng-0083 — neither arm had an energy join at the time this claim was first written). Wh-per-correct- answer over the identical 14 task-matched ids: f16 14.88 Wh/correct (163.696 Wh total, 11/14 passed), q8_0 8.66 Wh/correct (103.874 Wh total, 12/14 passed) — a 42% reduction. This is not merely the quality-neutral result clm-0058's `text` reports: q8_0's smaller KV cache also cut wall time nearly in half over the same tasks (2755s vs 4303s, window-for-window), so it costs less energy AND scores higher on this pair, not a trade between them. Both figures are WHOLE-SESSION energy (simulator wait time included, not agent-only decode), same caveat as every other tau2 arm's Wh-per-correct on this site.
CROSS-REFERENCE — clm-0022 (the speed result this claim's other half rests on, measured on the 122B); clm-0035 (the retracted stock-build quality-drop story this claim answers, on the same subject clm-0022 used); clm-0036 (the noise characterisation that retracted clm-0035, and the reason n=14 rather than n=5 was worth the box time here).
Cited by — computed at build time, never stored
docs coverage