Ling-3.0-flashbenchedguard 4/4
②Verdict
Ling-3.0-flash finally became measurable once both halves of the blocker were fixed: llama.cpp PR #26608 supplied the bailingmoe3 architecture, and the corrected bloomer010 GGUF supplied the `ssm_f_a` tensors missing from the older AtomicChat stock artifact. On that corrected path it is stable enough to run: FIT passed, the house guard passed 4/4, and the full 26-task tau2 airline arm completed cleanly with real tool use clm-0083.
Capability is weaker than the controlled-comparison motivation hoped: 13/26 = 0.500 mean reward, below the incumbent Qwen3.5-122B and below this lab's stronger full-bench records cited in the comparison table. The result is still scientifically useful because Ling has almost the same total footprint as the 122B while activating roughly half as many parameters per token, making it a direct active-parameter-count probe for the 122B comparison clm-0015 clm-0083.
Its efficiency result is the surprise: 69.218 Wh for the whole tau2 window and 5.32 Wh per correct answer, because the 30-minute run was much shorter than the slower large-model arms despite the lower score. Treat that as a same-protocol headline, not a settled architectural law: the run is ROCm-only, plain decode only, and no speculative/MTP path was exercised clm-0083.
HO-002 closes the MTP activation question: build 7077abb really loads the shipped NextN head, all n_max=1..3 arms had nonzero acceptance and passed the 4/4 guard, but their matched median ratios were only 1.007x, 0.968x and 0.850x. Plain decode therefore remains the production recommendation clm-0086.
③Best configuration
| model | Ling-3.0-flash-Q4_K_M.gguf · Q4_K_M |
| engine | ggml-org/llama.cpp 7077abb · rocm · host aihydra (igpu) |
| flags | -ngl 999 -fa on -c 32768 --parallel 1 --load-mode none --host 127.0.0.1 --port 8090 --jinja -rea off |
| template | not recorded at test time |
| tree | upstream — stock |
config record cfg-0115
every cell generated from the record at build time · throughput cells from cfg-0114 (same build 7077abb, fa on, f16 KV) · provenance: measured-here throughout
③aDecode against context depth
⑤Open questions
- Deeper-context scaling beyond d32768. The model's 35 KDA layers should make long-context behaviour unusually interesting, but this first bench only covers the standard serving context.
- Vulkan comparison on the corrected artifact. ROCm was chosen for the first publishable run to minimise surface area after the artifact-format blocker.