Home › Evidence › Records › clm-0103
⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.

clm-0103

measured-herehigh ●●●
citable URL: https://halobench.com/records/clm-0103/ — this address never moves; the anchor /records/#clm-0103 keeps resolving

The headbouyJB/diffuse-cpp fork makes a scaled block-diffusion LLM benchmarkable on portable AMD hardware: GPU offload of the MoE forward (gfx1151), OpenAI tool-calling, and a GPU-resident inter-step KV cache with cross-turn prompt reuse together take decode-at-long-context from ~1 tok/s (host-array cache) to ~5-12 tok/s, turning a multi-turn diffusion agent from impractical (~tens of seconds to minutes per turn) into a runnable tau2 screen (~1-3 min/task).

verified 2026-08-22 · volatility low
evidence run-0553 con-0005

Cited by — computed at build time, never stored