Home › Evidence › Records › clm-0118
⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.

clm-0118

measured-herehigh ●●●
citable URL: https://halobench.com/records/clm-0118/ — this address never moves; the anchor /records/#clm-0118 keeps resolving

HK-RERUN-REASONOFF verdict on the Kairic Edge TheRock build with reasoning DISABLED (-rea off, cfg-0172) — the fix path clm-0116 prescribed for the reasoning-ON empty-assistant defect. Verdict: ADMIT-with-corrections applied verbatim (HK-RERUN-REASONOFF-REVIEW, t_3b03c623). ADMITTED CLAIM 1 — tau2 full-26 capability, reasoning-OFF: reward 0.423 (11/26 passed, total_reward 11.0), 26/26 simulations COMPLETED with ZERO empty-assistant turns and zero empty-argument tool calls — the prior full-26's 60%-abort / "AssistantMessage must have either content or tool_calls" failure mode is eliminated by -rea off. The defect was a server-config interaction, not model capability. Wall 3181 s at ~13 t/s single-stream decode. ADMITTED CLAIM 2 — served-path spec-on/off throughput ratio ~1.0: matched 10/10 rung sweep (rungs 0/32768/131072/204800, ±256 depth-fill probes) of the same server with and without --kairic-edge; tg 13.1-13.2 t/s flat across the entire depth ladder in BOTH arms, per-cell deltas within noise. Includes d32768, stable in both states (the prior served sweep crashed there). --kairic-edge buys NO served-path decode throughput on this workload. Energy counterpoint (eng-0268 vs eng-0269): the spec=on band drew 17.1 Wh vs 22.2 Wh spec=off for identical tokens (~23% less wall energy); one observation, not an established effect. ADMITTED CLAIM 3 — d32768 stability: previously-crashing fill now completes rc=0 in both spec states. Production-config validation: build e1da26bb8 ("version: 80"), binary e121a4d3, gdn_output_fallback=48 (tau2 server) / 0 (all 20 sweep cells), served id qwen3.8-27b-kairic-edge — all re-verified on-box this ingest. BLOCKED CLAIMS (explicitly NOT made): no reasoning-ON reward comparison (prior full-26 incomplete/do-not-score under G6); no reasoning-ON server throughput claim (crashed). No production recommendation flips on energy alone from a single band pair.

verified 2026-08-23 · volatility medium
evidence clm-0116 cfg-0172 run-0618 run-0619 run-0620 eng-0267 eng-0268 eng-0269

Note — the record's own working

Ingest of t_b7766d91 under the reviewer boundary from t_3b03c623. Every number re-verified against raw aihydra artifacts (/home/aihydra/bench-results/hk001-kairic-rerun-reasonoff-r1/) this session: tau2-reasonoff-full26-summary.json (mean_reward 0.4230769…, 11/26, empty_assistant_turns=0, infrastructure_errors=0), specsweep-summary.txt (20/20 cells, tg flat 13.1-13.2), fingerprint-full26.txt (binary/model/build hashes match cfg-0164 lineage). All 21 energy windows joined from HA history same-day (well within ~10d retention; deadline was 2026-09-02) — no unjoined_pending_ingestion records needed. [2026-08-24 COMPARABILITY ANNOTATION — HK-V12-ANNOTATE, per t_b05473c9] Recorded runs used the v1.1 Kairic Edge runtime, whose native IU4 M65 verifier could flip a greedy target token on low-margin reproduced cases; throughput figures are UNAFFECTED (GGUF/sidecars unchanged), but the correctness of recorded tau2 outputs carries this caveat. Re-runs for capability comparability (full-26 tau2) are justified and gated behind the v1.2 runtime tag. All future Kairic Edge runs MUST use kairic-edge-qwen38-27b-v1.2 @ 205a3e5f (strict-compact M65 verifier default).

Cited by — computed at build time, never stored

claims clm-0119
candidate gate history qwen38-27b