Home › Evidence › Records › cfg-0131

cfg-0131

backfilled
citable URL: https://halobench.com/records/cfg-0131/ — this address never moves; the anchor /records/#cfg-0131 keeps resolving
modelOrnith-1.0-35B-UD-Q4_K_XL.gguf · UD-Q4_K_XL
engineggml-org/llama.cpp 3653e6d6d · vulkan · host aihydra (igpu)
flagsThroughput/guard boundary: -ngl 999 -fa 1 -b 2048 -ub 512 --load-mode none -dev Vulkan0 -ctk q8_0 -ctv q8_0 -t 16; llama-bench rows add pp1024/tg256 depth flags. Guard used c32768/depth8000 with raw llama-server, --parallel 1 / -np 1 semantics, no speculation, no MTP, no drafter, no proxy.
templatenot recorded at test time
treeupstream — stock
aged evidence — reconstructed from the archive. HO-005 reviewer-admitted q8_0/q8_0 KV production-depth performance probe from /home/aihydra/bench-results/ho005-ornith-q4q8-tuning-matrix-r2, using the same UD-Q4_K_XL weight artifact as cfg-0129 and Vulkan build_commit 3653e6d6d. This is a KV-policy performance probe only. It does not establish lossy-KV quality, Q4/Q8 capability equivalence, q4_0 KV behavior, or a broad production recommendation.

Capability basis

measured on this config

Runs on this config (7)

rundatekindsuitemetrics
run-05022026-08-20guardguard-c32768-depth8000tasks_passed=4 · tasks_total=4
run-05172026-08-20performancellama-bench@1prefill_tps=515.418847
run-05182026-08-20performancellama-bench@1decode_tps=53.909855
run-05192026-08-20performancellama-bench@1prefill_tps=134.867384
run-05202026-08-20performancellama-bench@1decode_tps=34.199869
run-05212026-08-20performancellama-bench@1prefill_tps=108.102044
run-05222026-08-20performancellama-bench@1decode_tps=30.617346