Models › deepseek-v4-flash
⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.

DeepSeek-V4-Flash-0731benchedguard 4/4

256 experts / 6 active + 1 shared (deepseek4 arch)MoE · 256 experts / 6 active + 1 shared (deepseek4) · MLA-style compressed KVquant held: UD-IQ3_XXS (unsloth) · sole staged quant, forced by the ~104 GB weight footprintfirst measured 2026-08-15latest run 2026-08-18

②Verdict

A 256-expert (6 active + 1 shared) MoE served at UD-IQ3_XXS — the only quant this candidate's ~104 GB weight footprint let onto the box — whose KV cache turns out to be roughly a ninth the size, per token, of a conventional hybrid-attention model's clm-0060. That is the largest single finding of this bench: the model is nowhere near memory-bound on this hardware even at its full 1M-token native context, and depth in this matrix (through d262144, the deepest cell measured in this programme so far) was chosen for wall-clock usefulness, not because anything deeper risked OOM.

Backend is build-specific. Stock 3653e6d ROCm swept the original matrix while stock Vulkan lost the GPU device on both allowed attempts at every depth from 32,768 tokens on clm-0059. A later carried Vulkan fork at baf6360be passed the house guard and completed cleanly through d262144, with no device loss clm-0084. Because that payload bundles multiple changes and has not run the full capability suite, ROCm remains the measured serving recommendation; the new Vulkan line is promising fork-specific performance evidence, not a stock-backend acquittal. Decode is plain: this candidate carries no native draft/MTP head in the stock build, so its 15.30 t/s floor at d0 falls to 6.86 t/s at d262144 with nothing to speed it up clm-0059.

Capability more than clears the bar the screen set a day earlier: the FULL standard 26-task τ² airline run completed end to end — not a wall-bound cut — at 0.846 mean reward (22/26), and the first five tasks reproduced the screen's own 1.000 exactly, on the same seed and task ids clm-0061. Energy favours this model over the programme's other full-bench comparator on the same denominator: 24.29 Wh per correct answer, 0.74 pence at the standing tariff, roughly a third less than qwen38-27b's 38.29 Wh despite running with no speculation at all — a small-active-parameter MoE apparently costs less energy per correct answer than a fast dense model decoding every parameter every step, even before accounting for decode speed clm-0061.

The one hazard this bench deliberately did not touch: an unverified community report of a K-cache quantisation bug specific to this model, so every number on this page assumes f16 KV and there is no quantised-KV arm to compare it against.

verdict written 2026-08-18 · every number above stands next to the claim chip that carries it

③Best configuration

modelDeepSeek-V4-Flash-0731-UD-IQ3_XXS-00001-of-00004.gguf · UD-IQ3_XXS
engineggml-org/llama.cpp 3653e6d · rocm · host aihydra (igpu)
flags-ngl 999 -fa on -c 32768 -dev ROCm0 -ctk f16 -ctv f16 --jinja --reasoning-format deepseek -n 4096 --parallel 1 --load-mode none --slots
templatenot recorded at test time
treeupstream — stock

config record cfg-0084

decode @ 0
15.30 t/s
run-0289 · CV 0%
decode @ 32k
12.35 t/s
run-0291 · CV 0%
prefill @ 0
140.80 t/s
run-0288 · CV 0.1%
prefill @ 32k
81.28 t/s
run-0290 · CV 0.1%
τ² airline · thinking off
0.846 ±0.139
run-0287 · passed 22/26
turns to done · median (all tasks)
24
run-0287 · successes only: 23.5
wall-clock to done · median (successes only)
5.0 min
run-0287 · failures excluded — they have no done
Wh per correct answer (τ², thinking off)
24.29 Wh
eng-0107 run-0287 · 0.74 p per answer

every cell generated from the record at build time · throughput cells from cfg-0082 (same build 3653e6d, fa on, f16 KV) · provenance: measured-here throughout

③aDecode against context depth

ROCmVulkanVulkan · v0.6.4 fork
05101520065k131k204.8k262.1kdecode t/scontext depth (tokens)serving context 32k15.30 t/s @ depth 0 · run-0289 · CV 0% · N=312.35 t/s @ depth 32k · run-0291 · CV 0% · N=310.50 t/s @ depth 65k · run-0293 · CV 0% · N=18.86 t/s @ depth 131k · run-0295 · CV 0% · N=16.86 t/s @ depth 262.1k · run-0297 · CV 0% · N=112.44 t/s @ depth 0 · run-0299 · CV 0% · N=313.32 t/s @ depth 0 · run-045014.24 t/s @ depth 0 · run-044815.41 t/s @ depth 0 · run-044617.83 t/s @ depth 0 · run-044419.54 t/s @ depth 0 · run-0442Vulkan · v0.6.4 fork · 19.54Vulkan · 12.44ROCm · 6.86
1/3 reps per cell · max CV 0.0% · build 3653e6d / 3653e6d / baf6360 · The stock Vulkan line ends at d0 because every deeper cell lost the GPU device on both allowed attempts. The carried baf6360be fork is a separate fingerprint and completed cleanly through d262144. · records: run-0289 run-0291 run-0293 run-0295 run-0297 run-0299 run-0450 run-0448 run-0446 run-0444 run-0442

④Other configurations tested — each as a delta against best

variantΔ decodeΔ turns (paired tasks)Δ energynoterecords
Stock Vulkan backend · same quant-19% @0—-33%On stock 3653e6d, decode is already ahead on ROCm at the one depth Vulkan survives, and prefill is close but ROCm-favoured — the gap that matters is not in this row. Vulkan cannot complete a single cell past d0; see the fit-ceiling row below for why this model can be served far deeper than d0 on the backend that actually holds. clm-0059
Carried Vulkan fork · same quant———The carried baf6360be Vulkan payload completes the full bounded depth series and is substantially faster than the old stock Vulkan d0 point. This is an observed configuration delta, not a one-patch attribution: the payload bundles Vulkan tuning, DeepSeek work and newer upstream changes, and it has only a house guard rather than a full capability run. clm-0084
KV cache cost (fit)———This model's KV cache costs roughly a ninth of a conventional hybrid-attention model's per token — small enough that its full declared native context fits the GTT window with room to spare. Depth was never a fit risk in this bench; wall-clock was the only real ceiling, and the matrix was sized around that. clm-0060
KV cache quantisation hazard—A community field report describes an open upstream incoherence-rotation bug that quantising the K cache trips specifically on this model, with a measured throughput cost in the wrong direction on top of the correctness risk. The original production bench avoided quantised KV by design; a later fork-specific v0.6.6 investigation admitted q8_0/q4_0 sparse-path proof and guarded throughput cells only, not quality equivalence or capability. run-0523 through run-0540 record the v0.6.6 portable Vulkan payload's matching f16/f16, q8_0/q8_0 and q4_0/q4_0 K/V cache performance through d262144 with sparse-path proof for q8_0/q4_0. This is not a full capability result and must not be read as q8/q4 quality-equivalence. clm-0096

deltas computed at build time from the named runs, paired-task traces and metered windows — a hazard row shows words where a delta would mislead

⑤Open questions

⑥Provenance

bench host aihydra · rocm · ggml-org/llama.cpp 3653e6d
discipline 3/1 reps per throughput cell · scatter published per cell (max CV 0.1%) · guard or written waiver on every performance series
window 2026-08-15 → 2026-08-18