⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.
clm-0085
measured-heremed ●●○
citable URL: https://halobench.com/records/clm-0085/ — this address never moves; the anchor /records/#clm-0085 keeps resolving
Gemma 4 26B-A4B UD-Q4_K_M did not earn a reflex-tier promotion on HG-001: the corrected ROCm stock run passed coherence, native tool-call arguments and the parallel-1 isolation skip, but failed the house needle retrieval at the first declared depth, 2048 tokens, for a 3/4 guard result. The predeclared stop gate fired before tau2 smoke, throughput, Vulkan testing or energy measurement, so no performance number or aihydra capability claim is admissible for this candidate. Existing evidence does not isolate a runtime crash or backend/device fault: the model loaded and answered the other guard probes, and the failing artifact is a retrieval/probe check returning tool-call text instead of the needle. Treat the rejection as a measured retrieval-envelope/probe failure, not as a general model-quality verdict.
Evidence source: HG-001 job card and aihydra ~/bench-results/gemma4-26b-promotion-20260818T220904Z/{status.tsv,queue.log, guard-rocm-d2048.txt}. Status shows artifact/build preflight OK (bytes=16947541728, sha256=f2c28b3dc4776931ac6f879e11f203dec637ea0f14267a86ec8f6165f63f293f, ROCm 3653e6d, Vulkan 3653e6d6d), ROCm serve OK, guard rc=1, ALL FAIL safe_depth=none. The earlier 20260818T220644Z attempt stopped during preflight before model load and is not benchmark evidence.