Home › Evidence › Records › clm-0117
⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.

clm-0117

measured-herehigh ●●●
citable URL: https://halobench.com/records/clm-0117/ — this address never moves; the anchor /records/#clm-0117 keeps resolving

On Qwen3.6-35B-A3B-MTP UD-Q4_K_M (build 2586f6edd strix-halo v0.6.10 fork/Vulkan RADV, capability c32768, plain mode -rea off, no MTP, f16/f16 KV baseline, temp 0 seed 42 max_tokens 4096) the production-optimisation matrix interior conclusions hold (internal, same model/quant/build): (1) interactive first-token latency (ttft) is minimally sensitive to batch once ubatch=512 — b1024/ub512 1353.1 ms (CV 1.46%), b2048/ub512 1357.4 (CV 2.01%), b512/ub512 1364.5 (CV 1.13%), all N=5 d131072 — and degrades when ubatch drops below 512 (ub256 ttft ~1538-1545 ms, ub128 2320 ms); tpot stays flat ~28.4-28.5 ms/tok across all six arms, so b1024/ub512 is the interactive/ttft production choice (35.1 t/s decode). (2) Depth is viable through d262144 (no OOM) at b1024/ub512: pp_tps 356.8/238.8/167.4 and tg_tps 35.0/28.8/24.9 at d131072/d200000/d262144 (N=3). CAVEAT required (house scatter rule, reviewer-required): the d262144 pp_tps 167.4 figure carries pp_cv 8.22% (>3% scatter rule) so it is a COARSE survival figure reported WITH its CV, NOT a clean mean; tg_tps 24.9 (CV 0.32%) is clean; the production config is d131072 so the recommendation is unaffected. The >=200k production-depth target for the agentic stack is satisfied. (3) KV q4_0 is a CANDIDATE production setting at the production config: decode ~24% faster (tpot 21.670 vs 28.503 ms, ttft within noise 1334.2 vs 1357.2 ms, N=5), mean KLD 0.0109±0.0003, PPL ratio 1.0005, same-top agreement 95.33% (control 100%) — small KLD, no cliff, no resolved regression, not rejected. CAVEAT (hybrid-arch, reviewer-required): this is a hybrid/recurrent model — only 11 of 40 layers expose a quantisable full-attention KV (G3 probe blk 3..40 step 4), so the small delta is partly ARCHITECTURAL; the recommendation is bounded to THIS exact fingerprint and is not a claim that q4_0 KV is lossless on a dense-attention or long-KV model. Boundary: internal production matrix only — same model, same quant, same build, same backend; only batch/ubatch, context depth, and KV quant vary. NO cross-model, cross-build, cross-backend-equivalence, MTP n2/n4, or reasoning-ON claim is made or inherited (HO-013 closed n_max flat + reasoning-ON rejected).

Note — the record's own working

Ingested from the combination of HO-014A-FILL Cell-1/2 (t_315f0104) and HO-014-RUN Cell-3 (t_677f1b47) via the HO-014-REVIEW ADMIT verdict (t_dc235dcd). All reviewer-required corrections applied verbatim: (1) d262144 pp_tps 167.4 labelled a coarse survival figure with its pp_cv 8.22%, never a clean mean; (2) energy join is an explicit HA counter-diff (eng-0256..0266), never a silent null; (3) production recommendation bounded to this fingerprint WITH the hybrid-arch caveat (11/40 layers quantisable). Values are RECORD values re-verified against summary.tsv / cell2-summary.tsv / cell3-summary.tsv and cell3-verdict.md. The Cell-3A KLD earlyoom at ctx=131072 was an instrument logits-buffer wall (~96 GiB RSS, not model OOM) and was repaired at capability c32768 with both arms identical (recorded in cfg-0171/run-0617). q4_0 decode advantage and safe-KV verdict establish a candidate production KV setting that only a full cross-model capability Δ could overturn — none is made here. This is a production-optimisation claim for the local-agent stack; it does not change the capability verdict (clm-0113) and is interior to this fingerprint.

Cited by — computed at build time, never stored

model pages qwen36-35b