cfg-0134
backfilled
citable URL: https://halobench.com/records/cfg-0134/ — this address never moves; the anchor /records/#cfg-0134 keeps resolving
| model | Qwen3.6-35B-A3B-MTP-UD-Q4_K_M.gguf · UD-Q4_K_M · rev 0b21525e972670ed59e1812e170b27c26355381f0656ecc4e25617ece7dac58b |
| engine | ggml-org/llama.cpp 7077abbe14c510cb829c93a1328c2815b5805ebd · ROCm0/gfx1151 · host aihydra (igpu) |
| flags | HO-009 promoted MTP n_max=2 fingerprint: same model/runtime/backend/KV/batch as cfg-0132 plus native model MTP/NextN path with --spec-draft-n-max 2. Depth performance uses raw llama-server /completion timing because this runtime's llama-bench exposes no MTP/spec flags; no ngram/external drafter, no KV quant, no proxy/llama-swap, no n_max>=4. |
| template | not recorded at test time |
| tree | upstream — stock |
aged evidence — reconstructed from the archive. Reviewer-admitted best clean MTP arm for HO-009 through d204800 only. n2 d262144 server-perf is preserved as a dropped HTTP 400 context-fit/refusal cell and must not be substituted.
Capability basis
measured on this config
Runs on this config (8)
| run | date | kind | suite | metrics |
|---|---|---|---|---|
| run-0526 | 2026-08-20 | performance | stageA-c32768 | decode_tps=76.07163244267113 |
| run-0529 | 2026-08-20 | guard | guard-c32768-depth8000 | tasks_passed=4 · tasks_total=4 |
| run-0535 | 2026-08-20 | performance | llama-server-completion@1 | decode_tps=64.94594427829692 |
| run-0536 | 2026-08-20 | performance | llama-server-completion@1 | decode_tps=65.06434302281384 |
| run-0537 | 2026-08-20 | performance | llama-server-completion@1 | decode_tps=43.957285713081795 |
| run-0538 | 2026-08-20 | performance | llama-server-completion@1 | decode_tps=36.20958105514719 |
| run-0539 | 2026-08-20 | performance | llama-server-completion@1 | runner_invalid=true |
| run-0540 | 2026-08-20 | performance | llama-bench@1 | decode_tps=15.063597 · prefill_tps=113.604116 |