clm-0117
On Qwen3.6-35B-A3B-MTP UD-Q4_K_M (build 2586f6edd strix-halo v0.6.10 fork/Vulkan RADV, capability c32768, plain mode -rea off, no MTP, f16/f16 KV baseline, temp 0 seed 42 max_tokens 4096) the production-optimisation matrix interior conclusions hold (internal, same model/quant/build): (1) interactive first-token latency (ttft) is minimally sensitive to batch once ubatch=512 — b1024/ub512 1353.1 ms (CV 1.46%), b2048/ub512 1357.4 (CV 2.01%), b512/ub512 1364.5 (CV 1.13%), all N=5 d131072 — and degrades when ubatch drops below 512 (ub256 ttft ~1538-1545 ms, ub128 2320 ms); tpot stays flat ~28.4-28.5 ms/tok across all six arms, so b1024/ub512 is the interactive/ttft production choice (35.1 t/s decode). (2) Depth is viable through d262144 (no OOM) at b1024/ub512: pp_tps 356.8/238.8/167.4 and tg_tps 35.0/28.8/24.9 at d131072/d200000/d262144 (N=3). CAVEAT required (house scatter rule, reviewer-required): the d262144 pp_tps 167.4 figure carries pp_cv 8.22% (>3% scatter rule) so it is a COARSE survival figure reported WITH its CV, NOT a clean mean; tg_tps 24.9 (CV 0.32%) is clean; the production config is d131072 so the recommendation is unaffected. The >=200k production-depth target for the agentic stack is satisfied. (3) KV q4_0 is a CANDIDATE production setting at the production config: decode ~24% faster (tpot 21.670 vs 28.503 ms, ttft within noise 1334.2 vs 1357.2 ms, N=5), mean KLD 0.0109±0.0003, PPL ratio 1.0005, same-top agreement 95.33% (control 100%) — small KLD, no cliff, no resolved regression, not rejected. CAVEAT (hybrid-arch, reviewer-required): this is a hybrid/recurrent model — only 11 of 40 layers expose a quantisable full-attention KV (G3 probe blk 3..40 step 4), so the small delta is partly ARCHITECTURAL; the recommendation is bounded to THIS exact fingerprint and is not a claim that q4_0 KV is lossless on a dense-attention or long-KV model. Boundary: internal production matrix only — same model, same quant, same build, same backend; only batch/ubatch, context depth, and KV quant vary. NO cross-model, cross-build, cross-backend-equivalence, MTP n2/n4, or reasoning-ON claim is made or inherited (HO-013 closed n_max flat + reasoning-ON rejected).