LLaDA2.2-flashscreenedno guard
②Verdict
LLaDA2.2-flash is the first non-autoregressive model on the field: a 100B-A13B block-diffusion MoE that lays out a masked template and refines it block by block, with Levenshtein DELETE/INSERT edits to correct its own drafts. Getting it to run as an agent at all required a purpose-built runtime — the headbouyJB/diffuse-cpp fork — because the stock diffusion runtime re-prefills the whole prompt on every denoising step (tens of seconds to minutes per turn) and has no GPU path for this hardware clm-0103.
On the 5-task tau2 airline SCREEN it scored 5/5, mean reward 1.000, with real multi-turn tool use and every task converging in 8-10 turns clm-0102. That is a strong screen, but it is a screen: N=5 on the airline split, not the full 26-task arm; energy-joined at 62.10 Wh for the window, 12.42 Wh per correct answer eng-0226. Read it as "the paradigm works here," not a settled capability number.
Two things must be said plainly. First, diffusion throughput is not comparable to the autoregressive field: the model does dozens of parallel-refinement forwards per 32-token block, so its ~6 tok/s on this workload is a different quantity from an AR decode rate — time-per-correct and Wh-per-correct are the honest cross-paradigm axes -- and on this screen the diffusion path spent 12.42 Wh per correct answer, roughly 2.3x Ling-3.0-flash's 5.32 Wh per correct on its full autoregressive run eng-0226, an early indicator, not a settled law (different N and task mix). Second, the run carries a fork and a hardware workaround: flash attention leaks the block-causal mask on gfx1151 clm-0104, so attention runs the slower manual path whenever the cache is on. The full run, the energy join, and how a diffusion model should be represented against an AR field are open editorial questions, not settled here.
③Best configuration
| model | Akicou/inclusionAI_LLaDA2.2-flash-GGUF/LLaDA2.2-flash-Q4_K_S.gguf · Q4_K_S |
| engine | Akicou/diffuse-cpp b799157 · rocm · host aihydra (igpu) |
| flags | -ngl 99 -t 8 -s 32 --cache --cache-ctx 16384 --host 127.0.0.1 --port 8080 |
| sampling | temp 0 · block_length 32 · threshold 0.95 · editing_threshold 0 · max_post_steps 16 |
| template | not recorded at test time |
| tree | fork — carrying con-0005 |
config record cfg-0143
carrying Runs on the headbouyJB/diffuse-cpp fork: GPU offload of the MoE forward (gfx1151), a GPU-resident inter-step KV cache with cross-turn prompt reuse, a manual soft_max attention path (flash leaks the block-causal mask on gfx1151), and a faithful port of the model's Levenshtein editing.
con-0005 clm-0103 clm-0104
every cell generated from the record at build time · throughput cells from cfg-0143 (same build b799157, fa on, f16 KV) · provenance: measured-here throughout
⑤Open questions
- The full 26-task airline arm and so its capability is comparable to the rest of the field (the 5-task screen is energy-joined at eng-0226; the full arm is not yet run).
- How a diffusion model should be surfaced against an autoregressive field — the cross-paradigm protocol (t/s is not comparable; time-per-correct and Wh-per-correct are). An editorial decision.
- An independent, generation-free minimal repro of the gfx1151 flash_attn_ext mask leak clm-0104, and whether it is worth an upstream ggml bug report.
- Whether the diffusion decode can be sped up (parallel-decode threshold, fewer post-steps, step distillation); block geometry is fixed by the model and is not a throughput lever.