Home › Evidence › Records › con-0005
⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.

con-0005

carrying
citable URL: https://halobench.com/records/con-0005/ — this address never moves; the anchor /records/#con-0005 keeps resolving
kind fork · upstream Akicou/diffuse-cpp
opened 2026-08-22 · https://github.com/headbouyJB/diffuse-cpp
carrying — reproducibility hazard: in use but not upstream. Every config listing this id in its divergence runs a fork.
problem — clm-0103
Makes inclusionAI LLaDA2.2-flash (100B-A13B block-diffusion MoE) run GPU-accelerated on AMD Strix Halo (gfx1151) and benchmarkable as an agent. Over the Akicou base (b799157): GPU offload of the MoE forward to the HIP/ROCm backend; OpenAI tool-calling in diffuse-server; a manual soft_max attention path that works around a gfx1151 flash_attn_ext mask leak [clm-0104]; a resurrected + GPU-resident inter-step KV cache with in-place flash decode and cross-turn prompt reuse (the key speedup, ~1 -> ~5-12 tok/s at long context) [clm-0103]; and a faithful port of the model's Levenshtein editing (M2T + T2T + DELETE/SPLIT + anti-loop + post-steps) restoring its self-correction. Fork to be published at headbouyJB/diffuse-cpp; patch currently held on aihydra. Base runtime AND the diffuse.* GGUF are a matched pair, so this carries the Akicou base rather than rebasing.

Cited by — computed at build time, never stored

claims clm-0103
model pages llada22-flash
candidate gate history gemma4-26b