Home › Evidence › Records › cfg-0134
⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.

cfg-0134

backfilled
citable URL: https://halobench.com/records/cfg-0134/ — this address never moves; the anchor /records/#cfg-0134 keeps resolving
modelQwen3.6-35B-A3B-MTP-UD-Q4_K_M.gguf · UD-Q4_K_M · rev 0b21525e972670ed59e1812e170b27c26355381f0656ecc4e25617ece7dac58b
engineggml-org/llama.cpp 7077abbe14c510cb829c93a1328c2815b5805ebd · ROCm0/gfx1151 · host aihydra (igpu)
flagsHO-009 promoted MTP n_max=2 fingerprint: same model/runtime/backend/KV/batch as cfg-0132 plus native model MTP/NextN path with --spec-draft-n-max 2. Depth performance uses raw llama-server /completion timing because this runtime's llama-bench exposes no MTP/spec flags; no ngram/external drafter, no KV quant, no proxy/llama-swap, no n_max>=4.
templatenot recorded at test time
treeupstream — stock
aged evidence — reconstructed from the archive. Reviewer-admitted best clean MTP arm for HO-009 through d204800 only. n2 d262144 server-perf is preserved as a dropped HTTP 400 context-fit/refusal cell and must not be substituted.

Capability basis

measured on this config

Runs on this config (8)

rundatekindsuitemetrics
run-05262026-08-20performancestageA-c32768decode_tps=76.07163244267113
run-05292026-08-20guardguard-c32768-depth8000tasks_passed=4 · tasks_total=4
run-05352026-08-20performancellama-server-completion@1decode_tps=64.94594427829692
run-05362026-08-20performancellama-server-completion@1decode_tps=65.06434302281384
run-05372026-08-20performancellama-server-completion@1decode_tps=43.957285713081795
run-05382026-08-20performancellama-server-completion@1decode_tps=36.20958105514719
run-05392026-08-20performancellama-server-completion@1runner_invalid=true
run-05402026-08-20performancellama-bench@1decode_tps=15.063597 · prefill_tps=113.604116