clm-0103
measured-herehigh ●●●
citable URL: https://halobench.com/records/clm-0103/ — this address never moves; the anchor /records/#clm-0103 keeps resolving
The headbouyJB/diffuse-cpp fork makes a scaled block-diffusion LLM benchmarkable on portable AMD hardware: GPU offload of the MoE forward (gfx1151), OpenAI tool-calling, and a GPU-resident inter-step KV cache with cross-turn prompt reuse together take decode-at-long-context from ~1 tok/s (host-array cache) to ~5-12 tok/s, turning a multi-turn diffusion agent from impractical (~tens of seconds to minutes per turn) into a runnable tau2 screen (~1-3 min/task).
Cited by — computed at build time, never stored
model pages llada22-flash
docs benchmark-protocol