Home › Evidence › Records › clm-0100
⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.

clm-0100

measured-heremed ●●○
citable URL: https://halobench.com/records/clm-0100/ — this address never moves; the anchor /records/#clm-0100 keeps resolving

HO-011 measured Ornith-1.5-35B Q4_K_M screening and full-26 tau2 capability on the stock ROCm0 3653e6d runtime at c32768 f16 KV. Screening passed guard 4/4, smoke 4/5 (mean 0.800, 74 tool-call messages, 0 empty turns, 0 infra errors), and throughput at d262144 of 142.38 t/s prefill and 20.99 t/s decode (N=5 fresh-process reps). The full-26 tau2 airline run scored 8/25 evaluated with mean_reward 0.320, 1 infra error (task 20 JSON parse, 4 retries), 5 TOO_MANY_ERRORS terminations, 2 MAX_STEPS terminations, and 8 context-overflow errors from conversation buildup exceeding c=32768. This is a sharp regression from Ornith-1.0-35B UD-Q4_K_XL (22/26, 0.846) — a version-replacement step-back on the standard suite.

verified 2026-08-21 · volatility medium
evidence run-0545 run-0546 run-0547 run-0548 run-0549 run-0550

Note — the record's own working

Scope boundary: HO-011 reviewer-admitted from /home/aihydra/bench-results/ho011-ornith15-q4km-preflight-r6 (screening) and /home/aihydra/bench-results/ho011-ornith15-q4km-full26-r1 (full-26 tau2). This is Q4_K_M evidence on its own boundary; it does not inherit Ornith-1.0 Q8_0 or UD-Q4_K_XL capability, does not establish 1.0/1.5 equivalence, and does not create an MTP/speculation, energy-ranking, or production-recommendation claim. The screening result is a sharp regression and does not justify promotion to full bench. Energy windows are recorded but not joined — this hbreviewer profile lacks Home Assistant credentials; run-meta timestamps are preserved for later batch join. Failed preflight attempts r1 through r5 and their bak001-bak005 backups are operational history only, not admitted metrics. Compare: Ornith-1.0 UD-Q4_K_XL (run-0499, clm-0093) measured 22/26 mean_reward 0.846 on the same tau2 airline suite.

Cited by — computed at build time, never stored

candidate gate history ornith-15-35b