clm-0093
measured-heremed ●●○
citable URL: https://halobench.com/records/clm-0093/ — this address never moves; the anchor /records/#clm-0093 keeps resolving
Ornith-1.0-35B UD-Q4_K_XL has its own measured HO-005 full-26 tau2 airline capability record on the stock 3653e6d6d Vulkan/f16-KV boundary: after a fresh schema-visible guard passed 4/4 (run-0498), the full task-id 0..25 tau2 airline run completed rc=0 at 22/26 with mean_reward 0.8461538461538461 (run-0499), 185 tool-call messages, 0 empty-argument tool calls, 0 empty assistant turns, 0 infrastructure errors, and 0 max-step cuts. This is measured Q4_K_XL evidence, not inherited Q8_0 capability and not a Q4/Q8 equivalence claim; the earlier bounded preflight/smoke claim clm-0092 remains bounded historical evidence.
Note — the record's own working
METHOD — HO-005 reviewer-admitted full-26 r1 evidence from /home/aihydra/bench-results/ho005-ornith-q4kxl-vulkan-full26-r1. Artifact: /home/aihydra/models/ornith-35b/Ornith-1.0-35B-UD-Q4_K_XL.gguf, 22,324,804,000 bytes, sha256 67081ae4a1a291bd6c72834094ea056332cb3cb5fa15e88536ec7f233a475b71, sidecar/local sha match, source unsloth/Ornith-1.0-35B-GGUF per the candidate record. Build/backend: /home/aihydra/src/llama.cpp-vk3653e6d HEAD 3653e6d6d547ec763317d9ecd0ace334a7e21359, build_commit 3653e6d6d, Vulkan0 Radeon 8060S, f16/f16 KV, command flags exactly -ngl 999 -fa 1 -b 2048 -ub 512 -c 32768 -dev Vulkan0 -ctk f16 -ctv f16 -t 16 --load-mode none --parallel 1 --jinja --reasoning-format deepseek -n 4096 --slots.
Guard: guard-c32768-depth8000 immediately before tau2, 2026-08-20T03:21:15Z..2026-08-20T03:21:24Z, rc=0, 4/4; guard-output.txt shows coherent generation, native get_booking tool call with non-empty arguments, retrieval at depth 8000, and isolation skipped only because --parallel 1/-np 1.
Tau2: suite tau2-bench-airline@668d3bc / tau2 head 668d3bcd135c02aa3438f987ef45735b7c163ee3, domain airline, task ids exactly 0..25, seed 42, num_trials 1, max_concurrency 1, max_steps 200, agent temperature 0 and max_tokens 4096, user simulator openrouter/anthropic/claude-haiku-4.5 temperature 0. Deterministic scoring and judge null match the admitted Q8 baseline/run-0345 practice; tau2-judge-note.txt records the CLI/judge rationale. Raw tau2-full26-results.json and tau2-full26-summary.json report failed task ids 7, 21, 23, 24, all termination_reason user_stop.
ENERGY STATUS — hborchestrator joined the wall-meter energy while raw Home Assistant history was still available. run-0498 is wired to eng-0202: 0.1888 Wh over the 9 s guard window, with no correct-answer denominator. run-0499 is wired to eng-0203: counter interpolated from 27.780999128 to 27.927536928 kWh across 2026-08-20T03:21:24Z..2026-08-20T04:31:09Z, 146.5378 Wh total, 6.6608 Wh/correct over 22 correct tasks, 2.0182 pence/correct at the standing 30.3 p/kWh tariff. This claim does not make a Q4/Q8 equivalence or energy-ranking claim.
CUSTODY/HEALTH — cleanup.log claim string was "claimed by HO-005 Ornith Q4 full-26 tau2 t_210dd8e6, started 2026-08-20T03:19:50Z, no fixed deadline"; cleanup left BOX_AFTER=ABSENT and no residual llama. postrun-validation.json reports final_state success, box_after ABSENT, residual_llama empty, and operational_fault_marker_hits empty. supervise-preclaim-proof showed GPU use 0% and llama-server ABSENT before claim; no reviewed DeviceLost/OOM/SVM/resident- limit/kernel-fault marker appears in operational logs.
COMPARISON BOUNDARY — run-0499 may be compared to the Q8_0 full-26 baseline run-0345 only as context: Q8_0 measured 23/26 and mean 0.8846 there, while this lossy UD-Q4_K_XL arm measured 22/26 and mean 0.8461538461538461 here. The two measurements do not establish quant equivalence or permit inheriting Q8_0 capability/energy claims onto Q4_K_XL.
Cited by — computed at build time, never stored
model pages ornith-35b
candidate gate history ornith-35b