cfg-0098
backfilled
citable URL: https://halobench.com/records/cfg-0098/ — this address never moves; the anchor /records/#cfg-0098 keeps resolving
| model | Qwen3.6-27B-Q4_K_M.gguf · Q4_K_M |
| engine | ggml-org/llama.cpp 3653e6d6d · vulkan · host aihydra (igpu) |
| flags | -ngl 999 -fa 1 -b 2048 -ub 512 -ctk f16 -ctv f16 -t 16 |
| template | not recorded at test time |
| tree | upstream — stock |
aged evidence — reconstructed from the archive. Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Stock 3653e6d6d Vulkan binary (~/src/llama.cpp-vk3653e6d/build/bin; house convention: reports build_commit 3653e6d6d, prefix-matches 3653e6d in protocol.json known_builds). Depths 0/32768 ran N=3 fresh-process reps and completed cleanly (max CV 3.0%, d32768 prefill — at the house's flag threshold, worth naming though decode at the same depth was clean at 0.03%, no defective-path pattern beyond ordinary prefill scatter). PARTIAL SERIES: d131072 is NOT in this config's throughput record — Vulkan device-lost on BOTH allowed attempts (2-attempt cap; vk::DeviceLostError, GPU wedged, recovered via kernel amdgpu ring reset both times), matching the laguna-s-21/deepseek-v4-flash device-loss precedent at comparable depth (ornith-35b broke that streak once; this candidate does not). No MTP/speculation: same from-scratch GGUF parse as cfg-0097, zero nextn/mtp/eagle/medusa/draft tensors; a live --spec-type draft-mtp load attempt on this exact Vulkan binary also refuses to start with the identical "model doesn't contain MTP layers" error. Output sanity gate passed via the chat endpoint before recording any throughput.
Capability basis
unmeasured — performance numbers on this config stand on a guard alone, not an established capability
Runs on this config (5)
| run | date | kind | suite | metrics |
|---|---|---|---|---|
| run-0352 | 2026-08-16 | performance | llama-bench@1 | prefill_tps=302.724 |
| run-0353 | 2026-08-16 | performance | llama-bench@1 | decode_tps=12.8337 |
| run-0354 | 2026-08-16 | performance | llama-bench@1 | prefill_tps=95.7181 |
| run-0355 | 2026-08-16 | performance | llama-bench@1 | decode_tps=11.331 |
| run-0356 | 2026-08-16 | performance | llama-bench@1 | device_lost=true |
Cited by — computed at build time, never stored
model pages qwen36-27b-mtp