⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.
clm-0129
measured-herehigh ●●●
citable URL: https://halobench.com/records/clm-0129/ — this address never moves; the anchor /records/#clm-0129 keeps resolving
On its proposed dense-WORKER serving config (cfg-0183: UD-Q4_K_XL, the pr27311 leak-fix build, self-speculative draft-mtp n_max 2, --parallel 4, reasoning_effort low, greedy, fronted by the slotpin proxy), Qwen3.8-27B scored 91.7% on tau2-bench airline tasks 0-25 — 22 of 24 scored passed, run end-to-end through the proxy under a real 4-slot concurrent workload, with 0 empty assistant turns over 106 minutes. Two of the 26 attempted tasks were infrastructure errors (empty AssistantMessage from the cloud user-simulator, openrouter/haiku-4.5) and are excluded from the denominator, as run-0271 excluded its too_many_errors task. Harness rollup: Read Actions 24/24, Write Actions 26/27, DB Match 23/24; LLM-judge agent errors 0. This is NOT a clean A/B against the 2026-08 screen's 0.577 (run-0271, same harness 668d3bc, same seed 42, same tasks 0-25): four fingerprint fields changed at once — quant (Q8_0 -> UD-Q4_K_XL), build (3653e6d -> pr27311), reasoning_effort (medium -> low) and agent max_tokens (4096 -> 8192), plus --parallel (1 -> 4). The leading hypothesis for the gap is the output cap — run-0271's 4096-token cap under reasoning=medium plausibly truncated agent turns mid-tool-call, which reasoning=low with an 8192 cap avoids — but that is a hypothesis, not an isolated measurement.
METHOD — tau2-bench airline, harness 668d3bcd135c02aa3438f987ef45735b7c163ee3, --domain airline --seed 42 --num-trials 1 --task-ids 0..25 --max-concurrency 4 --max-steps 100, user simulator pinned to openrouter/anthropic/claude-haiku-4.5 (temperature 0). Agent: local llama-server on cfg-0183 (pr27311 build c530ea79c, UD-Q4_K_XL, self-spec draft-mtp n_max 2, -c 49152, --parallel 4), agent sampling temperature 0.0 / max_tokens 8192, reasoning_effort low set both at the server (--reasoning-effort low) and in agent-llm-args. Window 2026-09-05 13:41:58-15:27:37Z.
reasoning=low VALIDATED FOR CAPABILITY, not just speed — the worker default holds full competence at this task set, so choosing it for throughput does not trade away the agentic score. n_max 2 is safely under this model's n_max>=4 EOS-cliff ceiling (clm-0055).
CONFIDENCE high on the number as a measurement of cfg-0183 (n=24 scored, 0 empty turns, 100-step ceiling never truncating a task at max_steps); MEDIUM would apply to any claim that a single lever CAUSED the lift over run-0271, which this deliberately does not make.
OPEN QUESTION — a matched single-lever isolation (hold quant/build/parallel, vary only max_tokens; then only reasoning_effort) would attribute the 0.577 -> 0.917 movement. Until then the two numbers are two honest measurements of two different configs, not a before/after. ENERGY: 15.18 Wh per correct answer (eng-0278), see clm-0131.