Home › Evidence › Records › clm-0104
⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.

clm-0104

measured-heremed ●●○
citable URL: https://halobench.com/records/clm-0104/ — this address never moves; the anchor /records/#clm-0104 keeps resolving

ggml_flash_attn_ext on this HIP/gfx1151 build does not fully honour an additive F16 block-causal mask: a committed prefix's per-layer K/V changed by several logits when a later masked block entered the attention window (compounding through layers), while the manual ggml_soft_max_ext path with the identical mask did not. The boundary tracks 64 keys (= 2x the Wave Size 32), pointing at a tiling/reduction issue in the flash kernel's handling of masked lanes. Inferred so far via a downstream KV-cache oracle; an independent generation-free minimal repro is drafted but NOT yet run, so this is not yet an upstream-filed bug.

verified 2026-08-22 · volatility medium
evidence run-0553

Cited by — computed at build time, never stored

model pages llada22-flash
candidate gate history qwen38-27b
docs benchmark-protocol