cfg-0131
backfilled
citable URL: https://halobench.com/records/cfg-0131/ — this address never moves; the anchor /records/#cfg-0131 keeps resolving
| model | Ornith-1.0-35B-UD-Q4_K_XL.gguf · UD-Q4_K_XL |
| engine | ggml-org/llama.cpp 3653e6d6d · vulkan · host aihydra (igpu) |
| flags | Throughput/guard boundary: -ngl 999 -fa 1 -b 2048 -ub 512 --load-mode none -dev Vulkan0 -ctk q8_0 -ctv q8_0 -t 16; llama-bench rows add pp1024/tg256 depth flags. Guard used c32768/depth8000 with raw llama-server, --parallel 1 / -np 1 semantics, no speculation, no MTP, no drafter, no proxy. |
| template | not recorded at test time |
| tree | upstream — stock |
aged evidence — reconstructed from the archive. HO-005 reviewer-admitted q8_0/q8_0 KV production-depth performance probe from /home/aihydra/bench-results/ho005-ornith-q4q8-tuning-matrix-r2, using the same UD-Q4_K_XL weight artifact as cfg-0129 and Vulkan build_commit 3653e6d6d. This is a KV-policy performance probe only. It does not establish lossy-KV quality, Q4/Q8 capability equivalence, q4_0 KV behavior, or a broad production recommendation.
Capability basis
measured on this config
Runs on this config (7)
| run | date | kind | suite | metrics |
|---|---|---|---|---|
| run-0502 | 2026-08-20 | guard | guard-c32768-depth8000 | tasks_passed=4 · tasks_total=4 |
| run-0517 | 2026-08-20 | performance | llama-bench@1 | prefill_tps=515.418847 |
| run-0518 | 2026-08-20 | performance | llama-bench@1 | decode_tps=53.909855 |
| run-0519 | 2026-08-20 | performance | llama-bench@1 | prefill_tps=134.867384 |
| run-0520 | 2026-08-20 | performance | llama-bench@1 | decode_tps=34.199869 |
| run-0521 | 2026-08-20 | performance | llama-bench@1 | prefill_tps=108.102044 |
| run-0522 | 2026-08-20 | performance | llama-bench@1 | decode_tps=30.617346 |