Home › Evidence › Records › clm-0119
⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.

clm-0119

measured-herehigh ●●●
citable URL: https://halobench.com/records/clm-0119/ — this address never moves; the anchor /records/#clm-0119 keeps resolving

HK-001-REASONING-MATRIX verdict on the Kairic Edge TheRock build of Qwen3.8-27B-IU4 (cfg-0173) — 5-cell tau2 smoke matrix isolating the reasoning lever on identical tasks/seeds/judge. Verdict: ADMIT-with-corrections applied verbatim (HK-MATRIX-REVIEW, t_8869ab61). ADMITTED CLAIM 1 — reasoning-ON improves smoke capability: all four reasoning-ON cells (R1-B6k unlimited; R2-B8k unlimited; R3-B8k-L budget=1024; R4-B8k-M budget=2048) scored 0.8 mean reward (4/5) vs reasoning-OFF R0-B4k at 0.6 (3/5). Same tasks 0-4, seed 42, claude-haiku-4.5 simulator — comparison fair. ADMITTED CLAIM 2 — budget capping preserves quality at this depth: capped cells R3 (1024) and R4 (2048) match unlimited R1/R2 at 0.8. ADMITTED CLAIM 3 — operational hygiene held across the matrix: 5/5 cells rc=0 with ZERO infra_errors and ZERO empty_assistant_turns; guard 4/4 PASS on R0-B4k, waived for R1-R4 under a documented binary-invariant rationale (model binary/sidecars/quant unchanged). ADMITTED COST SIGNAL — reasoning-ON roughly doubles wall clock (1808-1942 s vs 901 s) and wall energy (~84-92 Wh vs 43.2 Wh per cell; matrix total 394.4 Wh, eng-0270..eng-0274). RECOMMENDED CONFIG (reviewer): R2-B8k (max_tokens=8192, reasoning=on, budget=-1) for the full-26 confirmation run. REVIEWER CORRECTION applied verbatim: R0 smoke 0.6 differs from the full-26 reasoning-OFF baseline 0.423 (t_2330c08d lineage) — expected 5-task vs 26-task variance, not a protocol issue. BLOCKED CLAIMS (explicitly NOT made): these are 5-task smoke results — no full-26 capability claim, no production recommendation flip beyond the reviewer's config pick for the NEXT run, no energy effect claims beyond the single-matrix observation above.

Note — the record's own working

Ingest of HK-MATRIX-INGEST under the ADMIT_WITH_CORRECTIONS boundary from t_8869ab61. Every number re-verified against raw aihydra artifacts (/home/aihydra/bench-results/hk001-kairic-reasoning-matrix-r1/) this session: tau2-{cell}-results.json per-task rewards ([1,0,0,1,1] R0; [1,1,1,0,1] R1-R4) match matrix-table.tsv and matrix-summary.md exactly; fingerprint.txt matches cfg-0164/cfg-0172 lineage (binary e121a4d3, model 360caf73, build "version: 80 (e1da26bb8)", gdn_output_fallback=48); guard-waiver.txt rationale recorded. All 5 energy windows joined from HA history same-day (well within retention; cutoff ~2026-08-28 not approached) — no unjoined_pending_ingestion records needed. [2026-08-24 COMPARABILITY ANNOTATION — HK-V12-ANNOTATE, per t_b05473c9] All five matrix cells ran on the v1.1 Kairic Edge runtime, whose native IU4 M65 verifier could flip a greedy target token on low-margin reproduced cases; throughput figures are UNAFFECTED (GGUF/sidecars unchanged), but the correctness of recorded tau2 outputs carries this caveat — the reasoning-ON vs OFF smoke deltas could shift under the v1.2 strict-compact verifier. Capability-comparability re-runs (the matrix and the full-26 confirmation at the reviewer's R2-B8k pick) are justified on v1.2. All future Kairic Edge runs MUST use runtime tag kairic-edge-qwen38-27b-v1.2 @ 205a3e5f (strict-compact M65 verifier default; unsafe native path only via KAIRIC_UNSAFE_NATIVE_M65_VERIFY=1).

Cited by — computed at build time, never stored

claims clm-0127
candidate gate history qwen38-27b