Models › ling-30-flash
⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.

Ling-3.0-flashbenchedguard 4/4

124B / 5.1B activebailingmoe3 hybrid KDA + gated MLA MoE — MTP activated, no speed winquant held: Q4_K_M (corrected post-PR-26608 reference conversion)first measured 2026-08-13latest run 2026-08-19

②Verdict

Ling-3.0-flash finally became measurable once both halves of the blocker were fixed: llama.cpp PR #26608 supplied the bailingmoe3 architecture, and the corrected bloomer010 GGUF supplied the `ssm_f_a` tensors missing from the older AtomicChat stock artifact. On that corrected path it is stable enough to run: FIT passed, the house guard passed 4/4, and the full 26-task tau2 airline arm completed cleanly with real tool use clm-0083.

Capability is weaker than the controlled-comparison motivation hoped: 13/26 = 0.500 mean reward, below the incumbent Qwen3.5-122B and below this lab's stronger full-bench records cited in the comparison table. The result is still scientifically useful because Ling has almost the same total footprint as the 122B while activating roughly half as many parameters per token, making it a direct active-parameter-count probe for the 122B comparison clm-0015 clm-0083.

Its efficiency result is the surprise: 69.218 Wh for the whole tau2 window and 5.32 Wh per correct answer, because the 30-minute run was much shorter than the slower large-model arms despite the lower score. Treat that as a same-protocol headline, not a settled architectural law: the run is ROCm-only, plain decode only, and no speculative/MTP path was exercised clm-0083.

HO-002 closes the MTP activation question: build 7077abb really loads the shipped NextN head, all n_max=1..3 arms had nonzero acceptance and passed the 4/4 guard, but their matched median ratios were only 1.007x, 0.968x and 0.850x. Plain decode therefore remains the production recommendation clm-0086.

verdict written 2026-08-18 · every number above stands next to the claim chip that carries it

③Best configuration

modelLing-3.0-flash-Q4_K_M.gguf · Q4_K_M
engineggml-org/llama.cpp 7077abb · rocm · host aihydra (igpu)
flags-ngl 999 -fa on -c 32768 --parallel 1 --load-mode none --host 127.0.0.1 --port 8090 --jinja -rea off
templatenot recorded at test time
treeupstream — stock

config record cfg-0115

decode @ 0
36.08 t/s
run-0436 · CV 0.2%
decode @ 32k
32.89 t/s
run-0438 · CV 0.2%
prefill @ 0
409.55 t/s
run-0435 · CV 0.4%
prefill @ 32k
229.06 t/s
run-0437 · CV 0.2%
τ² airline · thinking off
0.500 ±0.192
run-0439 · passed 13/26
Wh per correct answer (τ², thinking off)
5.32 Wh
eng-0173 run-0439 · 0.16 p per answer

every cell generated from the record at build time · throughput cells from cfg-0114 (same build 7077abb, fa on, f16 KV) · provenance: measured-here throughout

③aDecode against context depth

010203040032k65k131k204.8kdecode t/scontext depth (tokens)serving context 32k36.08 t/s @ depth 0 · run-0436 · CV 0.2% · N=332.89 t/s @ depth 32k · run-0438 · CV 0.2% · N=3ROCm · Q4_K_M · plain decode · 32.89
3 reps per cell · max CV 0.2% · build 7077abb · Initial full bench covers ROCm plain decode at d0 and d32768 only. Ling's KDA-heavy architecture makes deeper context particularly interesting, but that ladder is deliberately left as a follow-up rather than mixed into the first publishable full run. · records: run-0436 run-0438

⑤Open questions

⑥Provenance

bench host aihydra · rocm · ggml-org/llama.cpp 7077abb
discipline 3 reps per throughput cell · scatter published per cell (max CV 0.4%) · guard or written waiver on every performance series
window 2026-08-13 → 2026-08-19