Models › llada22-flash
⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.

LLaDA2.2-flashscreenedno guard

100B / 13B activeLLaDA2 block-diffusion MoE (100B / 13B active) + Levenshtein editing — the field's first diffusion modelquant held: Q4_K_S (Akicou diffuse.* GGUF; loads only under diffuse-cpp)first measured 2026-08-22latest run 2026-08-22

②Verdict

LLaDA2.2-flash is the first non-autoregressive model on the field: a 100B-A13B block-diffusion MoE that lays out a masked template and refines it block by block, with Levenshtein DELETE/INSERT edits to correct its own drafts. Getting it to run as an agent at all required a purpose-built runtime — the headbouyJB/diffuse-cpp fork — because the stock diffusion runtime re-prefills the whole prompt on every denoising step (tens of seconds to minutes per turn) and has no GPU path for this hardware clm-0103.

On the 5-task tau2 airline SCREEN it scored 5/5, mean reward 1.000, with real multi-turn tool use and every task converging in 8-10 turns clm-0102. That is a strong screen, but it is a screen: N=5 on the airline split, not the full 26-task arm; energy-joined at 62.10 Wh for the window, 12.42 Wh per correct answer eng-0226. Read it as "the paradigm works here," not a settled capability number.

Two things must be said plainly. First, diffusion throughput is not comparable to the autoregressive field: the model does dozens of parallel-refinement forwards per 32-token block, so its ~6 tok/s on this workload is a different quantity from an AR decode rate — time-per-correct and Wh-per-correct are the honest cross-paradigm axes -- and on this screen the diffusion path spent 12.42 Wh per correct answer, roughly 2.3x Ling-3.0-flash's 5.32 Wh per correct on its full autoregressive run eng-0226, an early indicator, not a settled law (different N and task mix). Second, the run carries a fork and a hardware workaround: flash attention leaks the block-causal mask on gfx1151 clm-0104, so attention runs the slower manual path whenever the cache is on. The full run, the energy join, and how a diffusion model should be represented against an AR field are open editorial questions, not settled here.

verdict written 2026-08-22 · every number above stands next to the claim chip that carries it

③Best configuration

modelAkicou/inclusionAI_LLaDA2.2-flash-GGUF/LLaDA2.2-flash-Q4_K_S.gguf · Q4_K_S
engineAkicou/diffuse-cpp b799157 · rocm · host aihydra (igpu)
flags-ngl 99 -t 8 -s 32 --cache --cache-ctx 16384 --host 127.0.0.1 --port 8080
samplingtemp 0 · block_length 32 · threshold 0.95 · editing_threshold 0 · max_post_steps 16
templatenot recorded at test time
treefork — carrying con-0005

config record cfg-0143
carrying Runs on the headbouyJB/diffuse-cpp fork: GPU offload of the MoE forward (gfx1151), a GPU-resident inter-step KV cache with cross-turn prompt reuse, a manual soft_max attention path (flash leaks the block-causal mask on gfx1151), and a faithful port of the model's Levenshtein editing. con-0005 clm-0103 clm-0104

τ² airline · thinking off
1.000 n=5 smoke
run-0553 · passed 5/5
Wh per correct answer (τ², thinking off)
12.42 Wh n=5 smoke
eng-0226 run-0553 · 0.38 p per answer

every cell generated from the record at build time · throughput cells from cfg-0143 (same build b799157, fa on, f16 KV) · provenance: measured-here throughout

⑤Open questions

⑥Provenance

bench host aihydra · rocm · Akicou/diffuse-cpp b799157
discipline no chart-cell CV aggregate · guard or written waiver on every performance series
window 2026-08-22 → 2026-08-22