DeepSeek-V4-Flash-0731benchedguard 4/4
②Verdict
A 256-expert (6 active + 1 shared) MoE served at UD-IQ3_XXS — the only quant this candidate's ~104 GB weight footprint let onto the box — whose KV cache turns out to be roughly a ninth the size, per token, of a conventional hybrid-attention model's clm-0060. That is the largest single finding of this bench: the model is nowhere near memory-bound on this hardware even at its full 1M-token native context, and depth in this matrix (through d262144, the deepest cell measured in this programme so far) was chosen for wall-clock usefulness, not because anything deeper risked OOM.
Backend is build-specific. Stock 3653e6d ROCm swept the original matrix while stock Vulkan lost the GPU device on both allowed attempts at every depth from 32,768 tokens on clm-0059. A later carried Vulkan fork at baf6360be passed the house guard and completed cleanly through d262144, with no device loss clm-0084. Because that payload bundles multiple changes and has not run the full capability suite, ROCm remains the measured serving recommendation; the new Vulkan line is promising fork-specific performance evidence, not a stock-backend acquittal. Decode is plain: this candidate carries no native draft/MTP head in the stock build, so its 15.30 t/s floor at d0 falls to 6.86 t/s at d262144 with nothing to speed it up clm-0059.
Capability more than clears the bar the screen set a day earlier: the FULL standard 26-task τ² airline run completed end to end — not a wall-bound cut — at 0.846 mean reward (22/26), and the first five tasks reproduced the screen's own 1.000 exactly, on the same seed and task ids clm-0061. Energy favours this model over the programme's other full-bench comparator on the same denominator: 24.29 Wh per correct answer, 0.74 pence at the standing tariff, roughly a third less than qwen38-27b's 38.29 Wh despite running with no speculation at all — a small-active-parameter MoE apparently costs less energy per correct answer than a fast dense model decoding every parameter every step, even before accounting for decode speed clm-0061.
The one hazard this bench deliberately did not touch: an unverified community report of a K-cache quantisation bug specific to this model, so every number on this page assumes f16 KV and there is no quantised-KV arm to compare it against.
③Best configuration
| model | DeepSeek-V4-Flash-0731-UD-IQ3_XXS-00001-of-00004.gguf · UD-IQ3_XXS |
| engine | ggml-org/llama.cpp 3653e6d · rocm · host aihydra (igpu) |
| flags | -ngl 999 -fa on -c 32768 -dev ROCm0 -ctk f16 -ctv f16 --jinja --reasoning-format deepseek -n 4096 --parallel 1 --load-mode none --slots |
| template | not recorded at test time |
| tree | upstream — stock |
config record cfg-0084
every cell generated from the record at build time · throughput cells from cfg-0082 (same build 3653e6d, fa on, f16 KV) · provenance: measured-here throughout
③aDecode against context depth
④Other configurations tested — each as a delta against best
| variant | Δ decode | Δ turns (paired tasks) | Δ energy | note | records |
|---|---|---|---|---|---|
| Stock Vulkan backend · same quant | -19% @0 | — | -33% | On stock 3653e6d, decode is already ahead on ROCm at the one depth Vulkan survives, and prefill is close but ROCm-favoured — the gap that matters is not in this row. Vulkan cannot complete a single cell past d0; see the fit-ceiling row below for why this model can be served far deeper than d0 on the backend that actually holds. | clm-0059 |
| Carried Vulkan fork · same quant | — | — | — | The carried baf6360be Vulkan payload completes the full bounded depth series and is substantially faster than the old stock Vulkan d0 point. This is an observed configuration delta, not a one-patch attribution: the payload bundles Vulkan tuning, DeepSeek work and newer upstream changes, and it has only a house guard rather than a full capability run. | clm-0084 |
| KV cache cost (fit) | — | — | — | This model's KV cache costs roughly a ninth of a conventional hybrid-attention model's per token — small enough that its full declared native context fits the GTT window with room to spare. Depth was never a fit risk in this bench; wall-clock was the only real ceiling, and the matrix was sized around that. | clm-0060 |
| KV cache quantisation hazard | — | A community field report describes an open upstream incoherence-rotation bug that quantising the K cache trips specifically on this model, with a measured throughput cost in the wrong direction on top of the correctness risk. The original production bench avoided quantised KV by design; a later fork-specific v0.6.6 investigation admitted q8_0/q4_0 sparse-path proof and guarded throughput cells only, not quality equivalence or capability. | run-0523 through run-0540 record the v0.6.6 portable Vulkan payload's matching f16/f16, q8_0/q8_0 and q4_0/q4_0 K/V cache performance through d262144 with sparse-path proof for q8_0/q4_0. This is not a full capability result and must not be read as q8/q4 quality-equivalence. | clm-0096 | |
deltas computed at build time from the named runs, paired-task traces and metered windows — a hazard row shows words where a delta would mislead
⑤Open questions
- Speculative decoding — the community's DSpark drafter (bf16, ~11 GB) claims a 1.46x decode lift on this exact silicon, and a quantised-draft variant (Q2K, ~6.5 GB) that reportedly crashes on a known llama.cpp bug with an unconfirmed fix. Out of scope for this job by design (no native MTP/draft in the stock setup used here); a natural next bench for this candidate.
- K-only KV quantisation isolation — the community field report that motivated the f16-mandatory hazard could not isolate K from V (their build refused mixed cache types), so whether the incoherence bug is K-specific or a V-cache artefact too remains untested anywhere, including here.
- Whether decode continues degrading roughly linearly with depth past d262144, or hits a cliff before this model's declared 1,048,576-token native ceiling — the fit arithmetic says it would load, but nothing past d262144 was measured.
- Whether the v0.6.6 q8_0/q4_0 sparse-path performance cells preserve agentic capability or output quality. HO-001 v0.6.6 only admitted sparse-path proof, guards and llama-bench throughput, not a capability or quality-equivalence result.
- The standing community question this candidate was partly acquired to settle — antirez's Q2-Q4 quants vs this Unsloth UD-IQ3_XXS on quality — is still open; only the Unsloth quant was staged and screened here.
- Whether a higher quant than IQ3_XXS (UD-Q4_K_M or similar, in protocol.json's expected list) could be staged at all against the ~104 GB weight footprint this quant already carries — unresourced in this job; every run here used the protocol override this candidate has carried since its screen.