Home › Evidence › Records › clm-0089
⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.

clm-0089

measured-herehigh ●●●
citable URL: https://halobench.com/records/clm-0089/ — this address never moves; the anchor /records/#clm-0089 keeps resolving

Under HG-003's corrected tokenizer-measured protocol, LFM2-24B-A2B Q4_K_M failed retrieval at the minimum tested depth of 2048 tokens with the needle planted early at token position 512. The model returned plain placeholder content with non-empty token and position fields, but neither made a native tool call nor reproduced the planted token. The pre-registered bounded stop therefore fired before late placement and all Stage B guard or performance work. Safe depth is below the practical 4096-token promotion threshold, so this candidate earns no promotion and no performance claim. This is narrow retrieval-capability evidence, not a broad model-quality judgment or an infrastructure fault.

verified 2026-08-19 · volatility medium
evidence run-0473 eng-0197

Note — the record's own working

Evidence root: /home/aihydra/bench-results/lfm2-24b-retrieval-promotion-scientific-20260819T141200Z-supervisor-2edef59c. Exact model sha256 eb4d2d4d4e61b795726c2f526c4434ca6bc725ad7a783691b58681f025cf58f2; ROCm binary build commit 3653e6d (binary sha256 17f55ef427a56ebd98aabf0d3c48e43a93e384e1db5075c5c7bed5ed6a0aad01), with clean source worktree at ce7689f; prompt sha256 d68e71389f75666f2bcc63f55224e4de85a4dafe5a946d99e0a57f81e790ac8c.

Cited by — computed at build time, never stored

candidate gate history lfm2-24b