cfg-0177
citable URL: https://halobench.com/records/cfg-0177/ — this address never moves; the anchor /records/#cfg-0177 keeps resolving
| model | Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf · UD-Q4_K_XL · rev 4448186216b3af4cc558bbce2c3213f01608f8f8b2e5267a9767971dd3ec8082 |
| engine | unslothai/llama.cpp 250b61446efc91e3a179c8677956f2667c8fbda0 · vulkan · host aihydra (igpu) |
| flags | llama-server -m <UD-Q4_K_XL 4-shard set> -ngl 999 -fa 1 -b 2048 -ub 512 -c 32768 -dev Vulkan0 -ctk f16 -ctv f16 -t 16 --load-mode none --parallel 1 --jinja --reasoning-format deepseek --chat-template-kwargs '{"reasoning_effort":"low"}' -n 4096. The n-gram/PLE table tensor (per_layer_token_embd, 26.8 GiB iq4_nl) is placed on CPU AUTOMATICALLY by the Vulkan backend ("cannot be used with preferred buffer type Vulkan_Host, using CPU instead") — so -ot per_layer_token_embd=CPU is a NO-OP and is omitted; GTT-resident footprint = 76.7 GiB with ~43 GiB free for KV. tau2 agent args temperature=0.0, max_tokens=4096; airline domain, seed 42, tasks 0-25, claude-haiku-4.5 simulator (sim_self_play=false), max-steps 200. |
| template | not recorded at test time |
| tree | upstream — stock |
Capability basis
measured on this config
Runs on this config (1)
| run | date | kind | suite | metrics |
|---|---|---|---|---|
| run-0632 | 2026-08-27 | capability | tau2-airline-full26 | tasks_passed=24 · tasks_total=26 |