citable URL: https://halobench.com/records/clm-0068/ — this address never moves; the anchor /records/#clm-0068 keeps resolving
Ornith-1.0-35B (ornith-ai/deepreinforce-ai, arch qwen35moe, A3B-class MoE, MIT) passed the house standard screen on release-adjacent acquisition, stock 3653e6d ROCm, Q8_0, -np 1, c=32768: FIT loaded in 12s at 35,607 MiB GTT (comfortable against the 122,880 MiB boot window), GUARD 4/4 (coherence, tool call, needle at 8000 tokens, isolation skipped under -np 1), SMOKE 5/5 with mean_reward 1.000 on the 5-task tau2 airline smoke (21 tool-call messages, 0 empty assistant turns, 0 infra errors -- VALID per the smoke gate; run-0330, run-0331). This is a joint-strongest smoke result on this box to date, and directionally consistent with the vendor's own agentic-coding benchmark table (Terminal-Bench 2.1 64.2 vs Qwen3.6-35B's 52.5, SWE-bench Verified 75.6 vs 73.4).
ACQUISITION: Q8_0 from the OFFICIAL repo (ornith-ai/Ornith-1.0-35B-GGUF, 36,903,138,880 bytes, sha256 cbc992bca07901c1a51f33e65e6fc5d687de179c852a772dfd 15e4c3261dbf5c -- verified exact against the HF LFS oid post-download) as the primary screening quant; UD-Q4_K_XL and mmproj-F16 from unsloth/Ornith-1.0-35B- GGUF (22,324,804,000 bytes / 899,283,680 bytes, both sha256-verified exact against HF) as the comparability arm and vision file. Q8_0 is not in protocol.json quants.expected -- explicit PROTOCOL_OVERRIDE carried on every run (same rationale as qwen38-27b's cfg-0066: removes the quant confound, matches the candidate's own designated primary quant).
VISION CLAIM IS UNCONFIRMED FROM THE PRIMARY SOURCE. The task brief that motivated this acquisition described Ornith-1.0-35B as "vision-capable," but the official README's own highlights section describes the family purely as agentic-coding models (9B-Dense / 31B-Dense / 35B-MoE / 397B-MoE) with zero mention of vision, multimodal input, or an mmproj artifact anywhere in the page -- and the OFFICIAL repo ships NO mmproj file at all (siblings: only the five weight quants + README + two logo assets). unsloth's mirror DOES ship mmproj-{BF16,F16,F32}.gguf, but unsloth mechanically derives mmproj files from whatever vision_config is present in the base checkpoint's config -- this is consistent with "post-trained on top of ... Qwen 3.5" (the README's own framing) carrying over Qwen3.5's vision tower structurally without the coding-focused fine-tune having trained or evaluated vision capability at all. Acquired per the job brief regardless (the file is small, 899 MiB, and useful library context either way) but NOT screened -- the NPU/GPU screening tiers do not currently define a vision screen (same caveat qwen38-27b's mmproj-F16 carries) -- and the "vision-capable" framing should not be repeated as a confirmed fact without an actual vision-path test against the unsloth mmproj.
SCREEN METHOD NOTE: the first screen attempt (not ingested) omitted -np 1, defaulting to n_slots=4 / kv_unified=true -- a deviation from this job's stated house rule (--parallel 1 for every measured arm), inherited by copying the nemotron3-super-q4km-screen.sh template which has the same gap. Caught before ingestion; re-run clean with -np 1 explicit (n_slots=1, kv_unified=false confirmed in the server log) and BOTH runs scored identically (5/5, 1.000, tool-call counts 21 vs 25 -- both VALID, well within normal task-to-task variance for an agentic simulator), so the deviation did not appear to affect the result in this instance, but the ingested numbers (run-0330/run-0331) are from the corrected -np 1 run only.
Energy: both runs carry energy: null -- run-meta.jsonl (aihydra) has the started_at/ended_at pair for a follow-up batch join before the ~10-day HA retention window closes (protocol §11.2); not joined within this job's bounded time budget.
Not yet run: the throughput/depth matrix (qwen38-27b's queue-qwen38-screen.sh shape -- backend x depth x quant grid) and any MTP/speculative arm. The GGUF metadata was not inspected for MTP tensors in this job; if present, the qwen35moe family's own EOS-cliff precedent (clm-0055, on the architecture sibling qwen38-27b) is the first thing any MTP arm on this model should check for before trusting a spec-draft-n-max >= 4 result.