cfg-0178
citable URL: https://halobench.com/records/cfg-0178/ — this address never moves; the anchor /records/#cfg-0178 keeps resolving
| model | Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf · UD-Q4_K_XL · rev 4448186216b3af4cc558bbce2c3213f01608f8f8b2e5267a9767971dd3ec8082 |
| engine | unslothai/llama.cpp 250b61446efc91e3a179c8677956f2667c8fbda0 · vulkan · host aihydra (igpu) |
| flags | llama-bench decode-depth ladder: -ngl 999 -fa 1 -ctk f16 -ctv f16 -dev Vulkan0 -lm none -b 2048 -ub 512 -p 512 -n 128 -d {0,32768,131072} -r 3 in one matrix invocation, plus d204800 -r 3 as a separate clean run. Fresh-process per cell. NOTE: d245760 failed as a cell in the 0/32768/131072/245760 matrix invocation (opaque — rc=0, truncated output, no stderr error) but d204800 (>200k production depth) ran clean standalone; 262144 is omitted (Vulkan workgroup-count GGML_ASSERT at ggml-vulkan.cpp fires > ~262140). |
| template | not recorded at test time |
| tree | upstream — stock |
Capability basis
inherited — chain: cfg-0178 ← cfg-0177 · inheritance is legal across neutral levers only (protocol §1a)
Runs on this config (4)
| run | date | kind | suite | metrics |
|---|---|---|---|---|
| run-0633 | 2026-08-27 | performance | llama-bench-decode-depth | decode_tps=22.82 · prefill_tps=401.8 |
| run-0634 | 2026-08-27 | performance | llama-bench-decode-depth | decode_tps=14.09 · prefill_tps=201.3 |
| run-0635 | 2026-08-27 | performance | llama-bench-decode-depth | decode_tps=7.09 · prefill_tps=86.9 |
| run-0636 | 2026-08-27 | performance | llama-bench-decode-depth | decode_tps=5.73 · prefill_tps=63.9 |
Cited by — computed at build time, never stored
claims clm-0125
model pages qwen38-flash-next