Home › Evidence › Records
Records
Values marked derived are computed at build time and never stored in YAML — rendering them makes every build a check. Every record id also has a stable per-record page — /records/<id>/, linked from each heading — carrying the full record and its computed backlinks; the #<id> anchors on this page keep resolving.
Nodes (6)
gpt-oss-120b — gpt-oss-120b @ Q4_K_M
- kind model · engine igpu · tier candidate · runs_on aibeast
- config → cfg-0003
- derived current_state: rejected
- derived days_in_production: 0
- lifecycle: candidate 2026-06-27 → rejected 2026-06-28
ling-30-flash — Ling-3.0-flash @ Q4_K_M corrected GGUF
- kind model · engine igpu · tier candidate · runs_on aihydra
- config → cfg-0115
- derived current_state: candidate
- derived days_in_production: 0
- lifecycle: candidate 2026-08-18
nemotron3-super — Nemotron 3 Super @ IQ4_XS (not quant-equivalent to peers)
- kind model · engine igpu · tier candidate · runs_on aibeast
- config → cfg-0004
- derived current_state: rejected
- derived days_in_production: 0
- lifecycle: candidate 2026-06-27 → rejected 2026-06-28
qwen35-122b-a10b — Qwen3.5-122B-A10B-MTP @ UD-Q4_K_M, single slot
- kind model · engine igpu · tier daily-driver · runs_on aibeast
- config → cfg-0002
- derived current_state: retired
- derived days_in_production: 66
- lifecycle: candidate 2026-06-27 → live 2026-06-28 → retired 2026-09-02
qwen36-35b-a3b — Qwen3.6-35B-A3B @ UD-Q4_K_XL, reasoning off
- kind model · engine igpu · tier daily-driver · runs_on aibeast
- config → cfg-0001
- derived current_state: retired
- derived days_in_production: 58
- lifecycle: live 2026-05-01 → retired 2026-06-28
qwen38-flash-next — Qwen3.8-Flash-Next @ agention ROCmFP4-FAST-v2, single slot + adaptive MTP
- kind model · engine igpu · tier daily-driver · runs_on aibeast
- config → cfg-0182
- derived current_state: live
- derived days_in_production: 28
- lifecycle: live 2026-09-02
Configs (187)
cfg-0001 — unsloth/Qwen3.6-35B-A3B-GGUF @ UD-Q4_K_XL
- host aibeast · engine igpu · runtime ggml-org/llama.cpp@unknown-2026-06 (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; llama.cpp commit, chat-template hash and the memory block were not recorded at test time. Quant, backend, context, KV quant, serving flags and the reasoning setting are from the Phase 3C write-up.
- memory block not recorded at test time
- flags:
-ngl 999 --no-mmap --parallel 1 -fa 1 --jinja -c 131072 (reasoning OFF / enable_thinking:false; q8_0 KV; no MTP — this GGUF ships no heads) - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0002 — unsloth/Qwen3.5-122B-A10B-MTP-GGUF @ UD-Q4_K_M
- host aibeast · engine igpu · runtime ggml-org/llama.cpp@b6142 (rocm)
- derived runtime.tree: fork — carrying con-0001
- derived ctx_total: 200,000(200,000 × 1 slot(s))
- derived total_gb: 80.4 · headroom_gb: 15.6 · pool 96
- derived pool_utilisation (static estimate): 83.8%
- derived risk_utilisation: 83.8% — basis estimate (no observed peak collected; see clm-0001 — the estimate has never predicted an OOM) ⚠ OOM RISK
- acknowledged (ack-0001, review 2026-09-15): Runs at 83.8% of pool and has been the daily driver since 2026-06-28. Accepted deliberately: the 122B does not fit under 80% at any usable context, and the alternative is a materially weaker model. Mitigations after the 2026-07-21 OOM panic: --ctx-checkpoints capped at 4, desktop stack removed, swap raised to 16G. Revisit when aihydra arrives and the load can move.
- flags:
-ngl 999 --no-mmap --parallel 1 -fa 1 --jinja --spec-type draft-mtp --spec-draft-n-max 6 --ctx-checkpoints 4 - chat template: 7b21e9c4a5f60d18
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0003 — openai/gpt-oss-120b-GGUF @ Q4_K_M
- host aibeast · engine igpu · runtime ggml-org/llama.cpp@unknown-2026-06 (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; llama.cpp commit, chat-template hash and the per-model memory block were not recorded at test time. Quant, backend, context, KV quant, serving flags and the reasoning setting are from the Phase 3C write-up.
- memory block not recorded at test time
- flags:
-ngl 999 --no-mmap --parallel 1 -fa 1 --jinja -c 131072 (reasoning OFF / enable_thinking:false; q8_0 KV) - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0004 — nvidia/Nemotron-3-Super-GGUF @ IQ4_XS
- host aibeast · engine igpu · runtime ggml-org/llama.cpp@unknown-2026-06 (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; llama.cpp commit, chat-template hash and the per-model memory block were not recorded at test time. Quant, backend, context, KV quant, serving flags and the reasoning setting are from the Phase 3C write-up.
- memory block not recorded at test time
- flags:
-ngl 999 --no-mmap --parallel 1 -fa 1 --jinja -c 131072 (reasoning OFF / enable_thinking:false; q8_0 KV) - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0005 — unsloth/Qwen3.6-27B-GGUF @ Q4_K_M
- host aibeast · engine igpu · runtime ggml-org/llama.cpp@unknown-2026-06 (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; llama.cpp commit, chat-template hash and the per-model memory block were not recorded at test time. Quant, backend, context, KV quant, serving flags and the reasoning setting are from the Phase 3C write-up.
- memory block not recorded at test time
- flags:
-ngl 999 --no-mmap --parallel 1 -fa 1 --jinja -c 131072 (reasoning OFF / enable_thinking:false; q8_0 KV) - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0006 — unsloth/Qwen3.5-122B-A10B-MTP-GGUF @ UD-Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d6d547ec763317d9ecd0ace334a7e21359 (rocm)
- derived runtime.tree: upstream
- derived ctx_total: 16,384(16,384 × 1 slot(s))
- derived total_gb: 76.98 · headroom_gb: 28.02 · pool 105
- derived pool_utilisation (static estimate): 73.3%
- derived risk_utilisation: 73.3% — basis estimate (no observed peak collected; see clm-0001 — the estimate has never predicted an OOM)
- flags:
-ngl 999 -fa on -c 16384 --parallel 1 --load-mode none --jinja -ctk f16 -ctv f16 - chat template: not recorded
- capability: measured
cfg-0007 — unsloth/Qwen3.5-122B-A10B-MTP-GGUF @ UD-Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d6d547ec763317d9ecd0ace334a7e21359 (rocm)
- derived runtime.tree: upstream
- derived ctx_total: 16,384(16,384 × 1 slot(s))
- derived total_gb: 76.98 · headroom_gb: 28.02 · pool 105
- derived pool_utilisation (static estimate): 73.3%
- derived risk_utilisation: 73.3% — basis estimate (no observed peak collected; see clm-0001 — the estimate has never predicted an OOM)
- flags:
-ngl 999 -fa on -c 16384 --parallel 1 --load-mode none --jinja -ctk f16 -ctv f16 --spec-type draft-mtp --spec-draft-n-max 3 - chat template: not recorded
- capability: inherited from cfg-0006 across neutral lever(s): speculation
cfg-0008 — Qwen3.5-122B-A10B-UD-Q4_K_M-00001-of-00003.gguf @ UD-Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this tool. Everything present here — build commit, backend, KV quant, batch sizes, flash attention — is machine-reported.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 0 -b 2048 -ub 512 -ctk f16 -ctv f16 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0009 — Qwen3.5-122B-A10B-UD-Q4_K_M-00001-of-00003.gguf @ UD-Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this tool. Everything present here — build commit, backend, KV quant, batch sizes, flash attention — is machine-reported.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk f16 -ctv f16 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0010 — Qwen3.5-122B-A10B-UD-Q4_K_M-00001-of-00003.gguf @ UD-Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this tool. Everything present here — build commit, backend, KV quant, batch sizes, flash attention — is machine-reported.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk q8_0 -ctv q8_0 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0011 — Qwen3.5-122B-A10B-UD-Q4_K_M-00001-of-00003.gguf @ UD-Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@ce7689f (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this tool. Everything present here — build commit, backend, KV quant, batch sizes, flash attention — is machine-reported.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk q8_0 -ctv q8_0 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0012 — Qwen3.6-35B-A3B-UD-Q4_K_XL.gguf @ UD-Q4_K_XL
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this tool. Everything present here — build commit, backend, KV quant, batch sizes, flash attention — is machine-reported.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk f16 -ctv f16 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0013 — Qwen3.5-122B-A10B-UD-Q4_K_M-00001-of-00003.gguf @ UD-Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@ce7689f (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this tool. Everything present here — build commit, backend, KV quant, batch sizes, flash attention — is machine-reported.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk q8_0 -ctv q8_0 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0014 — Qwen3.5-122B-A10B-UD-Q4_K_M-00001-of-00003.gguf @ UD-Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@ce7689f (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this tool. Everything present here — build commit, backend, KV quant, batch sizes, flash attention — is machine-reported.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk q8_0 -ctv q8_0 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0015 — NVIDIA-Nemotron-3-Super-120B-A12B-UD-IQ4_XS-00001-of-00003.gguf @ UD-IQ4_XS
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this tool. Everything present here — build commit, backend, KV quant, batch sizes, flash attention — is machine-reported.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk f16 -ctv f16 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0016 — gpt-oss-120b-UD-Q4_K_XL-00001-of-00002.gguf @ UD-Q4_K_XL
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this tool. Everything present here — build commit, backend, KV quant, batch sizes, flash attention — is machine-reported.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk f16 -ctv f16 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0017 — gpt-oss-120b-UD-Q4_K_XL-00001-of-00002.gguf @ UD-Q4_K_XL
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this tool. Everything present here — build commit, backend, KV quant, batch sizes, flash attention — is machine-reported.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk f16 -ctv f16 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0018 — gpt-oss-120b-UD-Q4_K_XL-00001-of-00002.gguf @ UD-Q4_K_XL
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this tool. Everything present here — build commit, backend, KV quant, batch sizes, flash attention — is machine-reported.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk q8_0 -ctv q8_0 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0019 — NVIDIA-Nemotron-3-Super-120B-A12B-UD-IQ4_XS-00001-of-00003.gguf @ UD-IQ4_XS
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this tool. Everything present here — build commit, backend, KV quant, batch sizes, flash attention — is machine-reported.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 0 -b 2048 -ub 512 -ctk f16 -ctv f16 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0020 — NVIDIA-Nemotron-3-Super-120B-A12B-UD-IQ4_XS-00001-of-00003.gguf @ UD-IQ4_XS
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this tool. Everything present here — build commit, backend, KV quant, batch sizes, flash attention — is machine-reported.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk f16 -ctv f16 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0021 — NVIDIA-Nemotron-3-Super-120B-A12B-UD-IQ4_XS-00001-of-00003.gguf @ UD-IQ4_XS
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this tool. Everything present here — build commit, backend, KV quant, batch sizes, flash attention — is machine-reported.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk q8_0 -ctv q8_0 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0022 — Qwen3.6-35B-A3B-UD-Q4_K_XL.gguf @ UD-Q4_K_XL
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this tool. Everything present here — build commit, backend, KV quant, batch sizes, flash attention — is machine-reported.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 0 -b 2048 -ub 512 -ctk f16 -ctv f16 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0023 — Qwen3.6-35B-A3B-UD-Q4_K_XL.gguf @ UD-Q4_K_XL
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this tool. Everything present here — build commit, backend, KV quant, batch sizes, flash attention — is machine-reported.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk f16 -ctv f16 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0024 — Qwen3.6-35B-A3B-UD-Q4_K_XL.gguf @ UD-Q4_K_XL
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this tool. Everything present here — build commit, backend, KV quant, batch sizes, flash attention — is machine-reported.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk q8_0 -ctv q8_0 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0025 — Qwen3.5-122B-A10B-UD-Q4_K_M-00001-of-00003.gguf @ UD-Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d6d547ec763317d9ecd0ace334a7e21359 (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; tau2 serving config for the Aug 10-11 protocol arms (0.545 headline, anchor pair, sim-sensitivity, task9-regression). Flags come from the runner scripts — powered.sh on-box for the Aug 9-10 arms, bench/queues/queue-2026-08-11.sh / queue-f / queue-g in this repo — with the build asserted fatal at launch, not trusted. Memory block, HF revision and chat-template hash were not recorded at test time; GGUF sha256 is ledgered in audits/environment.lock-aihydra-2026-08-12.json.
- memory block not recorded at test time
- flags:
-ngl 999 -fa on -c 32768 --parallel 1 --load-mode none --jinja -rea off -ctk f16 -ctv f16 - chat template: not recorded
- capability: measured
cfg-0026 — Qwen3.5-122B-A10B-UD-Q4_K_M-00001-of-00003.gguf @ UD-Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d6d547ec763317d9ecd0ace334a7e21359 (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; cfg-0025 with quantised KV — the arms differ ONLY in -ctk/-ctv, which is the whole design of the clm-0038 paired comparison. Same provenance and gaps as cfg-0025.
- memory block not recorded at test time
- flags:
-ngl 999 -fa on -c 32768 --parallel 1 --load-mode none --jinja -rea off -ctk q8_0 -ctv q8_0 - chat template: not recorded
- capability: measured
cfg-0027 — NVIDIA-Nemotron-3-Super-120B-A12B-UD-IQ4_XS-00001-of-00003.gguf @ UD-IQ4_XS
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d6d547ec763317d9ecd0ace334a7e21359 (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; tau2 serving config for the Nemotron IQ4_XS arms (Aug 10 comparison arm, Stage E2). Flags from powered.sh (on-box) and bench/queues/queue-e-nemotron-quant.sh; build asserted at launch. Memory block, HF revision and chat-template hash not recorded at test time; GGUF sha256 ledgered in audits/environment.lock-aihydra-2026-08-12.json.
- memory block not recorded at test time
- flags:
-ngl 999 -fa on -c 32768 --parallel 1 --load-mode none --jinja -rea off -ctk f16 -ctv f16 - chat template: not recorded
- capability: measured
cfg-0028 — NVIDIA-Nemotron-3-Super-120B-A12B-UD-Q4_K_M-00001-of-00003.gguf @ UD-Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d6d547ec763317d9ecd0ace334a7e21359 (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; cfg-0027 with the model at UD-Q4_K_M — Stage E1 of the quant-pair deconfound (clm-0047): the two arms differ only in the model file's quantisation. Flags from bench/queues/queue-e-nemotron-quant.sh; build asserted at launch. Same gaps as cfg-0027.
- memory block not recorded at test time
- flags:
-ngl 999 -fa on -c 32768 --parallel 1 --load-mode none --jinja -rea off -ctk f16 -ctv f16 - chat template: not recorded
- capability: measured
cfg-0029 — Qwen3.5-122B-A10B-UD-Q4_K_M-00001-of-00003.gguf @ UD-Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@ce7689f (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; cfg-0025 on the KV-dequant-patched build (clm-0022's lever, clm-0045's arms). ce7689f is a LOCAL commit object — the cherry-pick of Nathanw1014's 2a24abc onto 3653e6d, tree ~/src/llama.cpp-kvfix on aihydra; recipe recorded in the audit, exact SHA not re-derivable bit-for-bit off-box (same caveat as cfg-0011/13/14). Build self-report asserted fatal at launch by the runner. Same recording gaps as cfg-0025.
- memory block not recorded at test time
- flags:
-ngl 999 -fa on -c 32768 --parallel 1 --load-mode none --jinja -rea off -ctk f16 -ctv f16 - chat template: not recorded
- capability: measured
cfg-0030 — Qwen3.5-122B-A10B-UD-Q4_K_M-00001-of-00003.gguf @ UD-Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@ce7689f (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; cfg-0029 with quantised KV — the patched half of the clm-0045 paired design (arms differ only in -ctk/-ctv on the SAME patched build). Same provenance and gaps as cfg-0029.
- memory block not recorded at test time
- flags:
-ngl 999 -fa on -c 32768 --parallel 1 --load-mode none --jinja -rea off -ctk q8_0 -ctv q8_0 - chat template: not recorded
- capability: measured
cfg-0031 — Qwen3.5-122B-A10B-UD-Q4_K_M-00001-of-00003.gguf @ UD-Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@62bf73d (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; cfg-0025 on the fleet-candidate minimal build min-62bf73d — the NEW side of the Stage A anchor pair (clm-0046's build-delta calibration: +2.8% pp / +0.9% tg with identical tau2 capability). Build tree ~/src/llama.cpp-min62bf73d on aihydra, asserted at launch by bench/queues/queue-2026-08-11.sh. Same recording gaps as cfg-0025.
- memory block not recorded at test time
- flags:
-ngl 999 -fa on -c 32768 --parallel 1 --load-mode none --jinja -rea off -ctk f16 -ctv f16 - chat template: not recorded
- capability: measured
cfg-0032 — Qwen3.6-35B-A3B-UD-Q4_K_XL.gguf @ UD-Q4_K_XL
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this tool. Everything present here — build commit, backend, KV quant, batch sizes, flash attention — is machine-reported.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk f16 -ctv f16 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0033 — Qwen3.6-35B-A3B-UD-Q4_K_XL.gguf @ UD-Q4_K_XL
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this tool. Everything present here — build commit, backend, KV quant, batch sizes, flash attention — is machine-reported.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk q8_0 -ctv q8_0 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0034 — Qwen3.6-35B-A3B-UD-Q4_K_XL.gguf @ UD-Q4_K_XL
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d6d (vulkan)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this tool. Everything present here — build commit, backend, KV quant, batch sizes, flash attention — is machine-reported.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk f16 -ctv f16 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0035 — Qwen3.6-35B-A3B-UD-Q4_K_XL.gguf @ UD-Q4_K_XL
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d6d (vulkan)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this tool. Everything present here — build commit, backend, KV quant, batch sizes, flash attention — is machine-reported.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk q8_0 -ctv q8_0 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0036 — Qwen3.6-35B-A3B-UD-Q4_K_XL.gguf @ UD-Q4_K_XL
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3be50ccc2 (vulkan)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this tool. Everything present here — build commit, backend, KV quant, batch sizes, flash attention — is machine-reported.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk f16 -ctv f16 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0037 — Qwen3.6-35B-A3B-UD-Q4_K_XL.gguf @ UD-Q4_K_XL
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3be50ccc2 (vulkan)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this tool. Everything present here — build commit, backend, KV quant, batch sizes, flash attention — is machine-reported.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk q8_0 -ctv q8_0 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0038 — gpt-oss-120b-UD-Q4_K_XL-00001-of-00002.gguf @ UD-Q4_K_XL
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this tool. Everything present here — build commit, backend, KV quant, batch sizes, flash attention — is machine-reported.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 0 -b 2048 -ub 512 -ctk f16 -ctv f16 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0039 — gpt-oss-120b-UD-Q4_K_XL-00001-of-00002.gguf @ UD-Q4_K_XL
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this tool. Everything present here — build commit, backend, KV quant, batch sizes, flash attention — is machine-reported.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk f16 -ctv f16 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0040 — gpt-oss-120b-UD-Q4_K_XL-00001-of-00002.gguf @ UD-Q4_K_XL
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this tool. Everything present here — build commit, backend, KV quant, batch sizes, flash attention — is machine-reported.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk q8_0 -ctv q8_0 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0041 — Qwen3.6-35B-A3B-UD-Q4_K_XL.gguf @ UD-Q4_K_XL
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this tool. Everything present here — build commit, backend, KV quant, batch sizes, flash attention — is machine-reported.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 0 -b 2048 -ub 512 -ctk f16 -ctv f16 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0042 — Qwen3.6-35B-A3B-UD-Q4_K_XL.gguf @ UD-Q4_K_XL
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this tool. Everything present here — build commit, backend, KV quant, batch sizes, flash attention — is machine-reported.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk f16 -ctv f16 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0043 — Qwen3.6-35B-A3B-UD-Q4_K_XL.gguf @ UD-Q4_K_XL
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this tool. Everything present here — build commit, backend, KV quant, batch sizes, flash attention — is machine-reported.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk q8_0 -ctv q8_0 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0044 — gpt-oss-120b-UD-Q4_K_XL-00001-of-00002.gguf @ UD-Q4_K_XL
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d6d (vulkan)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this tool. Everything present here — build commit, backend, KV quant, batch sizes, flash attention — is machine-reported.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk f16 -ctv f16 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0045 — gpt-oss-120b-UD-Q4_K_XL-00001-of-00002.gguf @ UD-Q4_K_XL
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d6d (vulkan)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this tool. Everything present here — build commit, backend, KV quant, batch sizes, flash attention — is machine-reported.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk q8_0 -ctv q8_0 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0046 — NVIDIA-Nemotron-3-Super-120B-A12B-UD-Q4_K_M-00001-of-00003.gguf @ UD-Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d6d (vulkan)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this tool. Everything present here — build commit, backend, KV quant, batch sizes, flash attention — is machine-reported.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk f16 -ctv f16 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0047 — NVIDIA-Nemotron-3-Super-120B-A12B-UD-Q4_K_M-00001-of-00003.gguf @ UD-Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d6d (vulkan)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this tool. Everything present here — build commit, backend, KV quant, batch sizes, flash attention — is machine-reported.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk q8_0 -ctv q8_0 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0048 — Qwen3.6-35B-A3B-UD-Q4_K_XL.gguf @ UD-Q4_K_XL
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d6d (vulkan)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this tool. Everything present here — build commit, backend, KV quant, batch sizes, flash attention — is machine-reported.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk f16 -ctv f16 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0049 — Qwen3.6-35B-A3B-UD-Q4_K_XL.gguf @ UD-Q4_K_XL
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d6d (vulkan)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this tool. Everything present here — build commit, backend, KV quant, batch sizes, flash attention — is machine-reported.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk q8_0 -ctv q8_0 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0050 — Qwen3.6-35B-A3B-UD-Q4_K_XL.gguf @ UD-Q4_K_XL
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@ed89854 (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this tool. Everything present here — build commit, backend, KV quant, batch sizes, flash attention — is machine-reported.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk bf16 -ctv bf16 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0051 — Qwen3.6-35B-A3B-UD-Q4_K_XL.gguf @ UD-Q4_K_XL
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@ed89854 (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this tool. Everything present here — build commit, backend, KV quant, batch sizes, flash attention — is machine-reported.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk f16 -ctv f16 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0052 — Qwen3.6-35B-A3B-UD-Q4_K_XL.gguf @ UD-Q4_K_XL
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this tool. Everything present here — build commit, backend, KV quant, batch sizes, flash attention — is machine-reported.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk bf16 -ctv bf16 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0053 — Qwen3.6-35B-A3B-UD-Q4_K_XL.gguf @ UD-Q4_K_XL
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this tool. Everything present here — build commit, backend, KV quant, batch sizes, flash attention — is machine-reported.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk f16 -ctv f16 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0054 — Qwen3.6-35B-A3B-UD-Q4_K_XL.gguf @ UD-Q4_K_XL
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@a94d563 (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this tool. Everything present here — build commit, backend, KV quant, batch sizes, flash attention — is machine-reported.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk bf16 -ctv bf16 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0055 — Qwen3.6-35B-A3B-UD-Q4_K_XL.gguf @ UD-Q4_K_XL
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@a94d563 (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this tool. Everything present here — build commit, backend, KV quant, batch sizes, flash attention — is machine-reported.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk f16 -ctv f16 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0056 — Qwen3.6-35B-A3B-UD-Q4_K_XL.gguf @ UD-Q4_K_XL
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d6d (vulkan)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this tool. Everything present here — build commit, backend, KV quant, batch sizes, flash attention — is machine-reported.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk f16 -ctv f16 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0057 — NVIDIA-Nemotron-3-Super-120B-A12B-UD-Q4_K_M-00001-of-00003.gguf @ UD-Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; perf-matrix-gtt120 re-test: the same nemotron3-super ROCm f16 cell that OOM'd in the main perf-matrix run (2026-08-13, default GTT ceiling), re-run 2026-08-14 with the kernel boot param amdgpu.gttsize=122880 (122 GiB GTT, vs the fleet default ~84 GiB — see strix-halo-host-config-reference) to test whether more GTT headroom rescues the cell. It did not: see run-0234. No memory block recorded — the process never reached a state where weights+KV footprint could be measured; it thrashed for 55 minutes before the kernel refused further SVM mapping. Everything present here is machine-asserted at launch (protocol_check_build), not measured mid-run.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk f16 -ctv f16 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0058 — NVIDIA-Nemotron-3-Super-120B-A12B-UD-Q4_K_M-00001-of-00003.gguf @ UD-Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; cfg-0057 with quantised KV — the second half of the perf-matrix-gtt120 re-test (2026-08-14, amdgpu.gttsize=122880). Halving the KV footprint via q8_0 did not avoid the OOM either: see run-0235. Same recording gaps as cfg-0057 — no memory block, machine-asserted build only.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk q8_0 -ctv q8_0 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0059 — NVIDIA-Nemotron-3-Super-120B-A12B-UD-Q4_K_M-00001-of-00003.gguf @ UD-Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Same arm as cfg-0057 (nemotron3-super-rocm-kvf16-faon-d32768), re-run 2026-08-14 with the one flag cfg-0057 was missing: --load-mode none. cfg-0057 and its predecessor perf-matrix cell both OOM'd (run-0234; status.tsv on aihydra for the original 2026-08-13 cell) because the perf-matrix/perf-matrix-gtt120 llama-bench queue scripts never carried --load-mode none, so ROCm was benchmarked under mmap on a model within ~40 GiB of the box's full 122 GiB — see clm-0053. On the identical binary with mmap off, this cell runs cleanly: run-0236, run-0237. No memory block recorded — ingested from llama-bench, which measures throughput and not footprint; everything present here (build commit, backend, KV quant, batch sizes, flash attention, load-mode) is machine-reported at launch.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk f16 -ctv f16 -t 16 --load-mode none - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0060 — Qwen3.6-35B-A3B-UD-Q4_K_XL.gguf @ UD-Q4_K_XL
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; cfg-0032 (the backend-decider ROCm f16 d32768 arm behind clm-0050's Qwen3.6-35B numbers, benchmarked under llama-bench's mmap default) with --load-mode none added and nothing else changed — the matched no-mmap companion built 2026-08-14 to test whether the mmap default is a throughput confound for a model that comfortably fits in memory (clm-0053's Qwen3.6-35B control pair; contrast Nemotron-3-Super, which does not fit and could not run under mmap at all — cfg-0057/run-0234). Same recording gaps as cfg-0032: ingested from llama-bench, no memory block, machine- reported build fingerprint only.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk f16 -ctv f16 -t 16 --load-mode none - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0061 — NVIDIA-Nemotron-3-Super-120B-A12B-UD-Q4_K_M-00001-of-00003.gguf @ UD-Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Same llama-bench invocation flags as cfg-0057 (mmap default, no --load-mode) — the difference this config exists to isolate is a KERNEL BOOT parameter, not a benchmark flag: this arm was measured on a boot with amdgpu.no_system_mem_limit=1 added to the cmdline (queue.log, memgate-test STAGE=postboot: "no_system_mem_limit: Y"), specifically to test whether that flag changes the mmap-under-memory-pressure failure mode. It did not prevent the failure, but it changed its SHAPE: see run-0242. Ingested from nemotron-postboot-mmap-rep1.stderr/queue.log/journalctl (aihydra ~/bench-results/memgate-test/) rather than a completed llama-bench JSON, because the process never produced one. No memory block recorded — same gaps as cfg-0057.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk f16 -ctv f16 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0062 — Qwen3.8-27B-Q8_0.gguf @ Q8_0
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this tool. Everything present here — build commit, backend, KV quant, batch sizes, flash attention — is machine-reported at launch. --load-mode none comes from the queue invocation itself (bench/queue-qwen38-screen.sh passes -lm none on every rep), added per clm-0053's rule that benchmark flags must match the serving standard on unified memory.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk f16 -ctv f16 -t 16 --load-mode none - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0063 — Qwen3.8-27B-UD-Q4_K_XL.gguf @ UD-Q4_K_XL
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this tool. Everything present here — build commit, backend, KV quant, batch sizes, flash attention — is machine-reported at launch. --load-mode none comes from the queue invocation itself (bench/queue-qwen38-screen.sh passes -lm none on every rep), added per clm-0053's rule that benchmark flags must match the serving standard on unified memory.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk f16 -ctv f16 -t 16 --load-mode none - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0064 — Qwen3.8-27B-Q8_0.gguf @ Q8_0
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d6d (vulkan)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this tool. Everything present here — build commit, backend, KV quant, batch sizes, flash attention — is machine-reported at launch. --load-mode none comes from the queue invocation itself (bench/queue-qwen38-screen.sh passes -lm none on every rep), added per clm-0053's rule that benchmark flags must match the serving standard on unified memory.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk f16 -ctv f16 -t 16 --load-mode none - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0065 — Qwen3.8-27B-UD-Q4_K_XL.gguf @ UD-Q4_K_XL
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d6d (vulkan)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this tool. Everything present here — build commit, backend, KV quant, batch sizes, flash attention — is machine-reported at launch. --load-mode none comes from the queue invocation itself (bench/queue-qwen38-screen.sh passes -lm none on every rep), added per clm-0053's rule that benchmark flags must match the serving standard on unified memory.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk f16 -ctv f16 -t 16 --load-mode none - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0066 — Qwen3.8-27B-Q8_0.gguf @ Q8_0
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- memory block not recorded at test time
- flags:
-ngl 999 -fa on -c 32768 -dev ROCm0 --spec-type draft-mtp --spec-draft-n-max 3 --jinja --reasoning-format deepseek --chat-template-kwargs {"reasoning_effort":"medium"} -n 4096 --parallel 1 --load-mode none --slots - chat template: not recorded
- capability: measured
cfg-0067 — Qwen3.8-27B-Q8_0.gguf @ Q8_0
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- memory block not recorded at test time
- flags:
-ngl 999 -fa on -c 32768 -dev ROCm0 --load-mode none --parallel 1 --spec-type draft-mtp --spec-draft-n-max 3 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0068 — Qwen3.8-27B-Q8_0.gguf @ Q8_0
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (vulkan)
- derived runtime.tree: upstream
- memory block not recorded at test time
- flags:
-ngl 999 -fa on -c 32768 -dev Vulkan0 --load-mode none --parallel 1 --spec-type draft-mtp --spec-draft-n-max 3 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0069 — Qwen3.8-27B-UD-Q4_K_XL.gguf @ UD-Q4_K_XL
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (vulkan)
- derived runtime.tree: upstream
- memory block not recorded at test time
- flags:
-ngl 999 -fa on -c 32768 -dev Vulkan0 --load-mode none --parallel 1 --spec-type draft-mtp --spec-draft-n-max 3 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0070 — Qwen3.6-35B-A3B-UD-Q4_K_XL.gguf @ UD-Q4_K_XL
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@ce7689f (rocm)
- derived runtime.tree: upstream
- memory block not recorded at test time
- flags:
-ngl 999 -fa on -c 32768 --parallel 1 --load-mode none --jinja --reasoning-format deepseek --slots -ctk f16 -ctv f16 -n 4096 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0071 — Qwen3.6-35B-A3B-UD-Q4_K_XL.gguf @ UD-Q4_K_XL
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@ce7689f (rocm)
- derived runtime.tree: upstream
- memory block not recorded at test time
- flags:
-ngl 999 -fa on -c 32768 --parallel 1 --load-mode none --jinja --reasoning-format deepseek --slots -ctk q8_0 -ctv q8_0 -n 4096 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0072 — NVIDIA-Nemotron-3-Super-120B-A12B-UD-Q4_K_M-00001-of-00003.gguf @ UD-Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- memory block not recorded at test time
- flags:
-ngl 999 -fa on -c 32768 -dev ROCm0 --load-mode none --jinja --reasoning-format deepseek -n 4096 --slots - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0073 — Qwen3.8-27B-UD-Q4_K_XL.gguf @ UD-Q4_K_XL
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- memory block not recorded at test time
- flags:
-ngl 999 -fa on -c 32768 -dev ROCm0 --load-mode none -ctk q8_0 -ctv q8_0 --parallel 1 --slots --spec-type draft-mtp --spec-draft-n-max 3 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0074 — Qwen3.8-27B-UD-Q4_K_XL.gguf @ UD-Q4_K_XL
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- memory block not recorded at test time
- flags:
-ngl 999 -fa on -c 32768 -dev ROCm0 --load-mode none -ctk q8_0 -ctv q8_0 --parallel 1 --slots --spec-type draft-mtp --spec-draft-n-max 7 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0075 — Qwen3.8-27B-Q8_0.gguf @ Q8_0
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (vulkan)
- derived runtime.tree: upstream
- memory block not recorded at test time
- flags:
-ngl 999 -fa on -c 32768 -dev Vulkan0 --load-mode none --parallel 1 --slots --spec-type draft-mtp --spec-draft-n-max 5 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0076 — Qwen3.8-27B-Q8_0.gguf @ Q8_0
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- memory block not recorded at test time
- flags:
-ngl 999 -fa on -c 32768 -dev ROCm0 --load-mode none --parallel 1 --slots --spec-type draft-mtp --spec-draft-n-max 5 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0077 — Qwen3.8-27B-Q8_0.gguf @ Q8_0
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (vulkan)
- derived runtime.tree: upstream
- memory block not recorded at test time
- flags:
-ngl 999 -fa on -c 32768 -dev Vulkan0 --load-mode none -ctk q8_0 -ctv q8_0 --parallel 1 --slots --spec-type draft-mtp --spec-draft-n-max 5 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0078 — Qwen3.8-27B-UD-Q4_K_XL.gguf @ UD-Q4_K_XL
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- memory block not recorded at test time
- flags:
-ngl 999 -fa on -c 32768 -dev ROCm0 --load-mode none -ctk q8_0 -ctv q8_0 --parallel 1 --slots --spec-type draft-mtp --spec-draft-n-max 5 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0079 — Qwen3.8-27B-Q8_0.gguf @ Q8_0
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (vulkan)
- derived runtime.tree: upstream
- memory block not recorded at test time
- flags:
-ngl 999 -fa on -c 32768 -dev Vulkan0 --load-mode none --parallel 1 --slots --spec-type draft-mtp --spec-draft-n-max 3 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0080 — Qwen3.8-27B-Q8_0.gguf @ Q8_0
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (vulkan)
- derived runtime.tree: upstream
- memory block not recorded at test time
- flags:
-ngl 999 -fa on -c 32768 -dev Vulkan0 --load-mode none --parallel 1 --slots --spec-type draft-mtp --spec-draft-n-max 5 --spec-draft-p-min 0.9 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0081 — Qwen3.8-27B-UD-Q4_K_XL.gguf @ UD-Q4_K_XL
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- memory block not recorded at test time
- flags:
-ngl 999 -fa on -c 32768 -dev ROCm0 --load-mode none -ctk q8_0 -ctv q8_0 --parallel 1 --slots --spec-type draft-mtp --spec-draft-n-max 7 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0082 — DeepSeek-V4-Flash-0731-UD-IQ3_XXS-00001-of-00004.gguf @ UD-IQ3_XXS
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk f16 -ctv f16 -t 16 --load-mode none - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0083 — DeepSeek-V4-Flash-0731-UD-IQ3_XXS-00001-of-00004.gguf @ UD-IQ3_XXS
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d6d (vulkan)
- derived runtime.tree: upstream
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk f16 -ctv f16 -t 16 --load-mode none - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0084 — DeepSeek-V4-Flash-0731-UD-IQ3_XXS-00001-of-00004.gguf @ UD-IQ3_XXS
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- memory block not recorded at test time
- flags:
-ngl 999 -fa on -c 32768 -dev ROCm0 -ctk f16 -ctv f16 --jinja --reasoning-format deepseek -n 4096 --parallel 1 --load-mode none --slots - chat template: not recorded
- capability: measured
cfg-0085 — Qwen3.8-27B-UD-Q4_K_XL.gguf @ UD-Q4_K_XL
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (vulkan)
- derived runtime.tree: upstream
- memory block not recorded at test time
- flags:
-ngl 999 -fa on -c 32768 -dev Vulkan0 --load-mode none --parallel 1 --slots --spec-type draft-mtp --spec-draft-n-max 5 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0086 — Qwen3.8-27B-Q8_0.gguf @ Q8_0
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@ece963f (vulkan)
- derived runtime.tree: upstream
- memory block not recorded at test time
- flags:
-ngl 999 -fa on -c 32768 -dev Vulkan0 --load-mode none --parallel 1 --slots --spec-type draft-mtp --spec-draft-n-max 5 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0087 — NVIDIA-Nemotron-3-Super-120B-A12B-UD-Q4_K_M-00001-of-00003.gguf @ UD-Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk f16 -ctv f16 -t 16 --load-mode none - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0088 — NVIDIA-Nemotron-3-Super-120B-A12B-UD-Q4_K_M-00001-of-00003.gguf @ UD-Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d6d (vulkan)
- derived runtime.tree: upstream
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk f16 -ctv f16 -t 16 --load-mode none - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0089 — NVIDIA-Nemotron-3-Super-120B-A12B-UD-Q4_K_M-00001-of-00003.gguf @ UD-Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d6d (vulkan)
- derived runtime.tree: upstream
- memory block not recorded at test time
- flags:
-ngl 999 -fa on -c 32768 -dev Vulkan0 --load-mode none --parallel 1 --jinja --reasoning-format deepseek -n 4096 --slots - chat template: not recorded
- capability: measured
cfg-0090 — laguna-s-2.1-Q4_K_M.gguf @ Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -dev ROCm0 -ctk f16 -ctv f16 -lm none - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0091 — laguna-s-2.1-Q4_K_M.gguf @ Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d6d (vulkan)
- derived runtime.tree: upstream
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -dev Vulkan0 -ctk f16 -ctv f16 -lm none - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0092 — laguna-s-2.1-Q4_K_M.gguf @ Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- memory block not recorded at test time
- flags:
-ngl 999 -fa on -c 32768 -dev ROCm0 --load-mode none --parallel 1 --jinja --reasoning-format deepseek -n 4096 --slots - chat template: not recorded
- capability: measured
cfg-0093 — Ornith-1.0-35B-Q8_0.gguf @ Q8_0
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- memory block not recorded at test time
- flags:
-ngl 999 -fa on -c 32768 -dev ROCm0 --jinja --reasoning-format deepseek -n 4096 -np 1 --load-mode none --slots - chat template: not recorded
- capability: measured
cfg-0094 — Ornith-1.0-35B-Q8_0.gguf @ Q8_0
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this tool. Everything present here — build commit, backend, KV quant, batch sizes, flash attention — is machine-reported. llama-bench's own -ctk/-ctv fields report f16 by default (no override passed on the command line; matches this candidate's 2026-08-16 screen, cfg-0093). Depths 0/32768 ran N=3 fresh-process reps (fresh server per rep, page cache not deliberately dropped between reps -- see queue.log); 131072 ran N=1 (deepseek-v4-flash/laguna-s21 convention: the number does not move at this scatter level on this box, and time is precious on a ~35-37GB resident model). Scatter across the 3 reps was clean at every depth (max CV 0.9%, d0 prefill) -- no defective-path signature. No MTP/speculation: this job's own from-scratch GGUF header parser (no numpy/gguf-py dependency available on-box) read all 733 tensor names directly and found zero nextn/mtp/eagle/medusa/draft tensors -- plain decode throughout. Output sanity gate passed before recording any throughput (chat-endpoint coherence check, 'the quick brown fox' echoed correctly).
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk f16 -ctv f16 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0095 — Ornith-1.0-35B-Q8_0.gguf @ Q8_0
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d6d (vulkan)
- derived runtime.tree: upstream
- memory block not recorded at test time
- flags:
-ngl 999 -fa on -c 32768 -dev Vulkan0 --load-mode none --parallel 1 --jinja --reasoning-format deepseek -n 4096 --slots - chat template: not recorded
- capability: measured
cfg-0096 — Ornith-1.0-35B-Q8_0.gguf @ Q8_0
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d6d (vulkan)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this tool. Everything present here — build commit, backend, KV quant, batch sizes, flash attention — is machine-reported. Stock 3653e6d6d Vulkan binary (~/src/llama.cpp-vk3653e6d/build/bin; house convention: reports build_commit 3653e6d6d, prefix-matches 3653e6d in protocol.json known_builds, same as cfg-0083/deepseek-v4-flash and laguna-s-21's cfg-0091). Depths 0/32768 ran N=3 fresh-process reps; 131072 ran N=1 (house convention). PARTIAL-SERIES caveat does NOT apply here, unlike laguna-s-21's cfg-0091: this candidate's Vulkan arm completed ALL THREE depths cleanly, including d131072 (238.5 pp / 32.4 tg t/s) -- no device loss on either allowed attempt, and only needed the first. Scatter across the 3 reps was clean at every depth (max CV 0.2%, d32768 prefill) -- tighter than the ROCm arm's own scatter at the same depths. No MTP/speculation: same from-scratch GGUF parse as cfg-0094, zero nextn/mtp/eagle/medusa/draft tensors among all 733. Output sanity gate passed via the chat endpoint before recording any throughput.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk f16 -ctv f16 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0097 — Qwen3.6-27B-Q4_K_M.gguf @ Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this tool. Everything present here — build commit, backend, KV quant, batch sizes, flash attention — is machine-reported. Depths 0/32768 ran N=3 fresh-process reps (fresh server per rep, page cache not deliberately dropped between reps); 131072 ran N=1 (house convention: the number does not move at this scatter level on this box, and time is precious on a resident model). Scatter across the 3 reps was clean (max CV 1.6%, d32768 prefill; decode CV <0.02% everywhere) — no defective-path signature. No MTP/speculation possible in this GGUF: this job's own from-scratch GGUF header parser (851 tensors) found zero nextn/mtp/ eagle/medusa/draft tensors, and a live --spec-type draft-mtp load attempt refuses to start on this exact binary ("model doesn't contain MTP layers") — plain decode throughout. Output sanity gate passed before recording any throughput (chat-endpoint coherence check, "the quick brown fox" echoed correctly).
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk f16 -ctv f16 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0098 — Qwen3.6-27B-Q4_K_M.gguf @ Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d6d (vulkan)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Stock 3653e6d6d Vulkan binary (~/src/llama.cpp-vk3653e6d/build/bin; house convention: reports build_commit 3653e6d6d, prefix-matches 3653e6d in protocol.json known_builds). Depths 0/32768 ran N=3 fresh-process reps and completed cleanly (max CV 3.0%, d32768 prefill — at the house's flag threshold, worth naming though decode at the same depth was clean at 0.03%, no defective-path pattern beyond ordinary prefill scatter). PARTIAL SERIES: d131072 is NOT in this config's throughput record — Vulkan device-lost on BOTH allowed attempts (2-attempt cap; vk::DeviceLostError, GPU wedged, recovered via kernel amdgpu ring reset both times), matching the laguna-s-21/deepseek-v4-flash device-loss precedent at comparable depth (ornith-35b broke that streak once; this candidate does not). No MTP/speculation: same from-scratch GGUF parse as cfg-0097, zero nextn/mtp/eagle/medusa/draft tensors; a live --spec-type draft-mtp load attempt on this exact Vulkan binary also refuses to start with the identical "model doesn't contain MTP layers" error. Output sanity gate passed via the chat endpoint before recording any throughput.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk f16 -ctv f16 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0099 — Qwen3.6-27B-Q4_K_M.gguf @ Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- memory block not recorded at test time
- flags:
-ngl 999 -fa on -c 32768 -ctk f16 -ctv f16 --jinja --reasoning-format deepseek -n 4096 --parallel 1 --load-mode none - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0100 — Qwen3.6-27B-Q4_K_M.gguf @ Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this tool. Everything present here — build commit, backend, KV quant, batch sizes, flash attention — is machine-reported.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk f16 -ctv f16 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0101 — always42-universal.gguf @ F16
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- derived ctx_total: 8,192(8,192 × 1 slot(s))
- derived total_gb: 1.15 · headroom_gb: 118.85 · pool 120
- derived pool_utilisation (static estimate): 1.0%
- derived risk_utilisation: 1.0% — basis observed
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk f16 -ctv f16 -t 16 - chat template: not recorded
- capability: measured
cfg-0102 — always42-universal.gguf @ F16
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d6d (vulkan)
- derived runtime.tree: upstream
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk f16 -ctv f16 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0103 — Qwen3.8-27B-Q8_0.gguf @ Q8_0
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Depth-ladder backfill session (2026-08-17), same fingerprint as cfg-0062 (identical llama-bench invocation: build 3653e6d, ROCm, Q8_0, f16 KV, -fa 1, --load-mode none — only depth differs, which is not a fingerprint field). A fresh config id is minted per house convention (each backfill invocation gets its own config record rather than silently appending to a prior session's), matching the qwen36-27b-mtp d204800 backfill precedent (cfg-0100 vs its own cfg-0097). Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Everything present here — build commit, backend, KV quant, batch sizes, flash attention, load-mode — is machine-reported at launch.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk f16 -ctv f16 -t 16 --load-mode none - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0104 — Ornith-1.0-35B-Q8_0.gguf @ Q8_0
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Depth-ladder backfill session (2026-08-17), same fingerprint as cfg-0094 (identical llama-bench invocation: build 3653e6d, ROCm, Q8_0, f16 KV, -fa 1 — the `flags` text matches cfg-0094 verbatim so the site's chart series can merge them; --load-mode none was passed on the actual command line, same as cfg-0094's own invocation, it is simply not part of either config's recorded flags string — only depth differs between the two configs' actual runs, not a fingerprint field). A fresh config id is minted per house convention (each backfill invocation gets its own config record). Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Everything present here — build commit, backend, KV quant, batch sizes, flash attention, load-mode — is machine-reported at launch.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk f16 -ctv f16 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0105 — Ornith-1.0-35B-Q8_0.gguf @ Q8_0
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d6d (vulkan)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Depth-ladder backfill session (2026-08-17), same fingerprint as cfg-0096 (identical llama-bench invocation: build 3653e6d6d, Vulkan, Q8_0, f16 KV, -fa 1 — the `flags` text matches cfg-0096 verbatim so the site's chart series can merge them; --load-mode none was passed on the actual command line, same as cfg-0096's own invocation, it is simply not part of either config's recorded flags string — only depth differs between the two configs' actual runs, not a fingerprint field). Stock 3653e6d6d Vulkan binary (~/src/llama.cpp-vk3653e6d/build/bin, house convention). A fresh config id is minted per house convention (each backfill invocation gets its own config record). Ingested from llama-bench, which measures throughput and not footprint. Everything present here — build commit, backend, KV quant, batch sizes, flash attention, load-mode — is machine-reported at launch.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk f16 -ctv f16 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0106 — Ornith-1.0-35B-Q8_0.gguf @ Q8_0
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d6d (vulkan)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Part 2 deep-cell session (2026-08-17), same fingerprint as cfg-0105 (identical llama-bench invocation: build 3653e6d6d, Vulkan, Q8_0, f16 KV, -fa 1, --load-mode none -- only depth differs, which is not a fingerprint field). This is the model-max cell (d262144) that extends cfg-0105's d204800 Vulkan-survival streak to ornith-35b's declared context ceiling. A fresh config id is minted per house convention. Ingested from llama-bench, which measures throughput and not footprint.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk f16 -ctv f16 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0107 — NVIDIA-Nemotron-3-Super-120B-A12B-UD-Q4_K_M-00001-of-00003.gguf @ UD-Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Part 2 deep-cell session (2026-08-17), same fingerprint as cfg-0087 (identical llama-bench invocation: build 3653e6d, ROCm, UD-Q4_K_M, f16 KV, -fa 1, --load-mode none -- only depth differs). Extends this candidate's ROCm throughput matrix past its prior deepest tested cell to d131072 and d204800. A fresh config id is minted per house convention. Ingested from llama-bench, which measures throughput and not footprint.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk f16 -ctv f16 -t 16 --load-mode none - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0108 — NVIDIA-Nemotron-3-Super-120B-A12B-UD-Q4_K_M-00001-of-00003.gguf @ UD-Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d6d (vulkan)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Part 2 deep-cell session (2026-08-17), same fingerprint as cfg-0088 (identical llama-bench invocation: build 3653e6d6d, Vulkan, UD-Q4_K_M, f16 KV, -fa 1, --load-mode none -- only depth differs). d131072 attempt: SURVIVED cleanly on the first of a 2-attempt cap, no vk::DeviceLostError, no GPU wedge -- notable given this box's #25664 device-loss dataset at comparable depth on other candidates (laguna-s-21, deepseek-v4-flash). A fresh config id is minted per house convention. Ingested from llama-bench, which measures throughput and not footprint.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk f16 -ctv f16 -t 16 --load-mode none - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0109 — laguna-s-2.1-Q4_K_M.gguf @ Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Part 2 deep-cell session (2026-08-17), same fingerprint as cfg-0090 (identical llama-bench invocation: build 3653e6d, ROCm, Q4_K_M, f16 KV, --load-mode none -- only depth differs). ROCm is this candidate's required backend at any useful depth (Vulkan lost the GPU device on both allowed attempts at d131072, clm-0065); this cell extends the ROCm-only series to d204800. A fresh config id is minted per house convention. Ingested from llama-bench, which measures throughput and not footprint.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -dev ROCm0 -ctk f16 -ctv f16 -lm none - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0110 — Qwen3-Coder-Next-Q8_0-00001-of-00004.gguf @ Q8_0
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- memory block not recorded at test time
- flags:
-ngl 999 -fa on -c 32768 -dev ROCm0 --jinja --reasoning-format deepseek -n 4096 -np 1 --load-mode none --slots - chat template: not recorded
- capability: measured
cfg-0111 — Qwen3-Coder-Next-Q8_0-00001-of-00004.gguf @ Q8_0
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d6d (vulkan)
- derived runtime.tree: upstream
- memory block not recorded at test time
- flags:
-ngl 999 -fa on -c 32768 -dev Vulkan0 --jinja --reasoning-format deepseek -n 4096 --parallel 1 --load-mode none --slots - chat template: not recorded
- capability: measured
cfg-0112 — Qwen3-Coder-Next-Q8_0-00001-of-00004.gguf @ Q8_0
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this tool. Everything present here — build commit, backend, KV quant, batch sizes, flash attention — is machine-reported.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk f16 -ctv f16 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0113 — Qwen3-Coder-Next-Q8_0-00001-of-00004.gguf @ Q8_0
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d6d (vulkan)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this tool. Everything present here — build commit, backend, KV quant, batch sizes, flash attention — is machine-reported.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk f16 -ctv f16 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0114 — Ling-3.0-flash-Q4_K_M.gguf @ Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@7077abb (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench JSON, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this record. The rows behind this config cover d0 and d32768, each as a single llama-bench invocation containing prefill and decode rows. Capability guard is recorded separately as run-0434 against the matching serving config cfg-0115.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk f16 -ctv f16 -t 16 -lm none - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0115 — Ling-3.0-flash-Q4_K_M.gguf @ Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@7077abb (rocm)
- derived runtime.tree: upstream
- memory block not recorded at test time
- flags:
-ngl 999 -fa on -c 32768 --parallel 1 --load-mode none --host 127.0.0.1 --port 8090 --jinja -rea off - chat template: not recorded
- capability: measured
cfg-0116 — DeepSeek-V4-Flash-0731-UD-IQ3_XXS-00001-of-00004.gguf @ UD-IQ3_XXS
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@baf6360be (vulkan)
- derived runtime.tree: fork — carrying con-0002
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -dev Vulkan0 --load-mode none -ctk f16 -ctv f16 -p 1024 -n 256 -b 2048 -ub 512 -r 1 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0117 — DeepSeek-V4-Flash-0731-UD-IQ3_XXS-00001-of-00004.gguf @ UD-IQ3_XXS
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@baf6360be (vulkan)
- derived runtime.tree: fork — carrying con-0002
- memory block not recorded at test time
- flags:
-ngl 999 -fa on -c 32768 -ctk f16 -ctv f16 --load-mode none --host 127.0.0.1 --port 8101 --jinja --reasoning-format deepseek -n 4096 --slots - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0118 — gemma-4-26B-A4B-it-UD-Q4_K_M.gguf @ Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- memory block not recorded at test time
- flags:
-ngl 999 -fa on -c 2048 -ctk f16 -ctv f16 --parallel 1 --load-mode none --host 127.0.0.1 --port 8102 --jinja -n 4096 - chat template: not recorded
- capability: measured
cfg-0119 — Ling-3.0-flash-Q4_K_M.gguf @ Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@7077abb (rocm)
- derived runtime.tree: upstream
- memory block not recorded at test time
- flags:
-ngl 999 -dev ROCm0 -fa on -ctk f16 -ctv f16 -c 32768 --load-mode none --parallel 1 --jinja --spec-type draft-mtp --spec-draft-n-max 1 -ctkd f16 -ctvd f16 --host 127.0.0.1 --port 8090 - chat template: not recorded
- capability: measured
cfg-0120 — Ling-3.0-flash-Q4_K_M.gguf @ Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@7077abb (rocm)
- derived runtime.tree: upstream
- memory block not recorded at test time
- flags:
-ngl 999 -dev ROCm0 -fa on -ctk f16 -ctv f16 -c 32768 --load-mode none --parallel 1 --jinja --spec-type draft-mtp --spec-draft-n-max 2 -ctkd f16 -ctvd f16 --host 127.0.0.1 --port 8090 - chat template: not recorded
- capability: measured
cfg-0121 — Ling-3.0-flash-Q4_K_M.gguf @ Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@7077abb (rocm)
- derived runtime.tree: upstream
- memory block not recorded at test time
- flags:
-ngl 999 -dev ROCm0 -fa on -ctk f16 -ctv f16 -c 32768 --load-mode none --parallel 1 --jinja --spec-type draft-mtp --spec-draft-n-max 3 -ctkd f16 -ctvd f16 --host 127.0.0.1 --port 8090 - chat template: not recorded
- capability: measured
cfg-0122 — NVIDIA-Nemotron-3.5-Lightning-30B-A3B-Q4_K_M.gguf @ Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@7077abb (rocm)
- derived runtime.tree: upstream
- memory block not recorded at test time
- flags:
-ngl 999 -dev ROCm0 -fa on -ctk f16 -ctv f16 -c 32768 --load-mode none --parallel 1 --host 127.0.0.1 --port 8090 --spec-type none - chat template: not recorded
- capability: measured
cfg-0123 — NVIDIA-Nemotron-3.5-Lightning-30B-A3B-Q4_K_M.gguf @ Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@7077abb (rocm)
- derived runtime.tree: upstream
- memory block not recorded at test time
- flags:
-ngl 999 -dev ROCm0 -fa on -ctk f16 -ctv f16 -c 32768 --load-mode none --parallel 1 --host 127.0.0.1 --port 8090 --spec-type draft-mtp --spec-draft-n-max 1 -ctkd f16 -ctvd f16 - chat template: not recorded
- capability: measured
cfg-0124 — NVIDIA-Nemotron-3.5-Lightning-30B-A3B-Q4_K_M.gguf @ Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@7077abb (rocm)
- derived runtime.tree: upstream
- memory block not recorded at test time
- flags:
-ngl 999 -dev ROCm0 -fa on -ctk f16 -ctv f16 -c 32768 --load-mode none --parallel 1 --host 127.0.0.1 --port 8090 --spec-type draft-mtp --spec-draft-n-max 2 -ctkd f16 -ctvd f16 - chat template: not recorded
- capability: measured
cfg-0125 — NVIDIA-Nemotron-3.5-Lightning-30B-A3B-Q4_K_M.gguf @ Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@7077abb (rocm)
- derived runtime.tree: upstream
- memory block not recorded at test time
- flags:
-ngl 999 -dev ROCm0 -fa on -ctk f16 -ctv f16 -c 32768 --load-mode none --parallel 1 --host 127.0.0.1 --port 8090 --spec-type draft-mtp --spec-draft-n-max 3 -ctkd f16 -ctvd f16 - chat template: not recorded
- capability: measured
cfg-0126 — LFM2-24B-A2B-Q4_K_M.gguf @ Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@ce7689f (rocm)
- derived runtime.tree: upstream
- memory block not recorded at test time
- flags:
-ngl 999 -dev ROCm0 --load-mode none -fa on --parallel 1 -ctk f16 -ctv f16 -c 6144 --no-mmap --jinja -n 4096 --host 127.0.0.1 --port 8103 - chat template: not recorded
- capability: measured
cfg-0127 — Qwen3.6-35B-A3B-UD-Q4_K_XL.gguf @ UD-Q4_K_XL
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@cd644c39545aac3dca63261f99a9bfc35956cb25 (vulkan)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from admitted llama-bench r7 artifacts in /home/aihydra/bench-results/pr25494-vulkan-q8kv-stock-r7. Source/build provenance admitted by hbreviewer t_ce7d961e: HEAD cd644c39545aac3dca63261f99a9bfc35956cb25, git describe b10362-158-gcd644c395, tree clean, binary /home/aihydra/src/llama.cpp-pr25494-stock-cd644c/build/bin/llama-bench stat size 17904. Model file stat size 22360456160 at /home/aihydra/models/qwen36-35b/Qwen3.6-35B-A3B-UD-Q4_K_XL.gguf. llama-bench reports model_size 22349466112, cpu AMD RYZEN AI MAX+ 395 w/ Radeon 8060S, gpu Radeon 8060S Graphics (RADV STRIX_HALO). Throughput only: no capability guard was run for this current-stock Vulkan r7 boundary.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 --load-mode none -dev Vulkan0 -p 1024 -n 256 -o json -ctk f16 -ctv f16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0128 — Qwen3.6-35B-A3B-UD-Q4_K_XL.gguf @ UD-Q4_K_XL
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@cd644c39545aac3dca63261f99a9bfc35956cb25 (vulkan)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from admitted llama-bench r7 artifacts in /home/aihydra/bench-results/pr25494-vulkan-q8kv-stock-r7. Source/build provenance admitted by hbreviewer t_ce7d961e: HEAD cd644c39545aac3dca63261f99a9bfc35956cb25, git describe b10362-158-gcd644c395, tree clean, binary /home/aihydra/src/llama.cpp-pr25494-stock-cd644c/build/bin/llama-bench stat size 17904. Model file stat size 22360456160 at /home/aihydra/models/qwen36-35b/Qwen3.6-35B-A3B-UD-Q4_K_XL.gguf. llama-bench reports model_size 22349466112, cpu AMD RYZEN AI MAX+ 395 w/ Radeon 8060S, gpu Radeon 8060S Graphics (RADV STRIX_HALO). Throughput only: no capability guard was run for this current-stock Vulkan r7 boundary.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 --load-mode none -dev Vulkan0 -p 1024 -n 256 -o json -ctk q8_0 -ctv q8_0 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0129 — Ornith-1.0-35B-UD-Q4_K_XL.gguf @ UD-Q4_K_XL
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d6d (vulkan)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Bounded HO-005 r3 preflight/smoke ingestion only, admitted by hbreviewer t_2cabe4fa from /home/aihydra/bench-results/ho005-ornith-q4kxl-vulkan-preflight-r3, then a successor admitted full-26 tau2 capability run added run-0498 and run-0499 on the same artifact/runtime/backend/f16-KV serving fingerprint. Artifact identity: /home/aihydra/models/ornith-35b/Ornith-1.0-35B-UD-Q4_K_XL.gguf, 22,324,804,000 bytes, sha256 67081ae4a1a291bd6c72834094ea056332cb3cb5fa15e88536ec7f233a475b71, sidecar/local sha match, source unsloth/Ornith-1.0-35B-GGUF per candidate record. Build/backend fingerprint: /home/aihydra/src/llama.cpp-vk3653e6d HEAD 3653e6d6d547ec763317d9ecd0ace334a7e21359, build_commit 3653e6d6d, Vulkan0 Radeon 8060S, f16 KV, --load-mode none, tau2 head 668d3bcd135c02aa3438f987ef45735b7c163ee3. FIT/header/load proof was healthy at c32768 (2026-08-20T00:17:59Z..00:18:15Z, rc=0, gtt_used_bytes 22444122112). The later full-26 arm used the exact serving flags recorded above, passed a fresh 4/4 guard, and measured 22/26 with mean_reward 0.8461538461538461 at Q4_K_XL. capability_basis is measured for this Q4_K_XL boundary only; it is not inherited from Q8_0 and does not establish Q4/Q8 equivalence. HO-005 reviewer-admitted production-depth performance matrix r2 later reused this same UD-Q4_K_XL / Vulkan0 / f16/f16 KV / build_commit 3653e6d6d throughput boundary for llama-bench pp1024/tg256 cells at depths 0/32768/131072/204800/262144 (run-0507 through run-0516) after a fresh schema-visible guard (run-0500). These are performance records only, not a new tau2 capability claim, not Q4/Q8 equivalence, and not a broad production recommendation; energy is structured-unjoined on the new records because the hbrunner ingestion profile lacked Home Assistant history credentials.
- memory block not recorded at test time
- flags:
Serving/capability boundary: -ngl 999 -fa 1 -b 2048 -ub 512 -c 32768 -dev Vulkan0 -ctk f16 -ctv f16 -t 16 --load-mode none --parallel 1 --jinja --reasoning-format deepseek -n 4096 --slots. Bounded throughput rows using this same artifact/runtime/backend/f16-KV boundary add their llama-bench-specific -p/-n/-d/-o json flags on the run records. - chat template: not recorded
- capability: measured
cfg-0130 — Ornith-1.0-35B-UD-Q4_K_XL.gguf @ UD-Q4_K_XL
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; HO-005 reviewer-admitted production-depth performance matrix from /home/aihydra/bench-results/ho005-ornith-q4q8-tuning-matrix-r2. Artifact identity is the unsloth UD-Q4_K_XL file already recorded on ornith-35b (22,324,804,000 bytes, sha256 67081ae4a1a291bd6c72834094ea056332cb3cb5fa15e88536ec7f233a475b71). Bench JSON reports build_commit 3653e6d/build_number 1 on ROCm, while runtime source-path inspection during review found /home/aihydra/src/llama.cpp clean at HEAD ce7689f875f724ece08b1a9fbb2a83971120b4d2; preserve both rather than rewriting the source-tree identity. This config records guarded performance only, not tau2 capability and not a production recommendation.
- memory block not recorded at test time
- flags:
Throughput/guard boundary: -ngl 999 -fa 1 -b 2048 -ub 512 --load-mode none -ctk f16 -ctv f16 -t 16; llama-bench rows add pp1024/tg256 depth flags. Guard used c32768/depth8000 with raw llama-server, --parallel 1 / -np 1 semantics, no speculation, no MTP, no drafter, no proxy. - chat template: not recorded
- capability: measured
cfg-0131 — Ornith-1.0-35B-UD-Q4_K_XL.gguf @ UD-Q4_K_XL
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d6d (vulkan)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; HO-005 reviewer-admitted q8_0/q8_0 KV production-depth performance probe from /home/aihydra/bench-results/ho005-ornith-q4q8-tuning-matrix-r2, using the same UD-Q4_K_XL weight artifact as cfg-0129 and Vulkan build_commit 3653e6d6d. This is a KV-policy performance probe only. It does not establish lossy-KV quality, Q4/Q8 capability equivalence, q4_0 KV behavior, or a broad production recommendation.
- memory block not recorded at test time
- flags:
Throughput/guard boundary: -ngl 999 -fa 1 -b 2048 -ub 512 --load-mode none -dev Vulkan0 -ctk q8_0 -ctv q8_0 -t 16; llama-bench rows add pp1024/tg256 depth flags. Guard used c32768/depth8000 with raw llama-server, --parallel 1 / -np 1 semantics, no speculation, no MTP, no drafter, no proxy. - chat template: not recorded
- capability: measured
cfg-0132 — Qwen3.6-35B-A3B-MTP-UD-Q4_K_M.gguf @ UD-Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@7077abbe14c510cb829c93a1328c2815b5805ebd (ROCm0/gfx1151)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Reviewer-admitted HO-009 plain baseline for the genuine Qwen3.6-35B-A3B-MTP UD-Q4_K_M artifact at /home/aihydra/bench-results/ho009-qwen36-35b-a3b-mtp-refinement-r2. This is an activation/safety/performance boundary only, not a capability or production claim.
- memory block not recorded at test time
- flags:
HO-009 plain/control fingerprint: /home/aihydra/src/llama.cpp-master-bailingmoe3, server build 50 commit 7077abb, ROCm0 Radeon 8060S, --load-mode none, -ngl 999, flash attention on, -b 2048, -ub 512, f16/f16 KV, --parallel 1, no proxy/llama-swap/ngram/external drafter/speculation. - chat template: not recorded
- capability: measured
cfg-0133 — Qwen3.6-35B-A3B-MTP-UD-Q4_K_M.gguf @ UD-Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@7077abbe14c510cb829c93a1328c2815b5805ebd (ROCm0/gfx1151)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Stage A activation/safety arm only. Positive draft counters were admitted, but n1 was not promoted as the best clean arm.
- memory block not recorded at test time
- flags:
HO-009 MTP n_max=1 Stage A fingerprint: same as cfg-0132 plus native model MTP/NextN path with --spec-draft-n-max 1. No ngram/external drafter, no KV quant, no proxy/llama-swap, no n_max>=4. - chat template: not recorded
- capability: measured
cfg-0134 — Qwen3.6-35B-A3B-MTP-UD-Q4_K_M.gguf @ UD-Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@7077abbe14c510cb829c93a1328c2815b5805ebd (ROCm0/gfx1151)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Reviewer-admitted best clean MTP arm for HO-009 through d204800 only. n2 d262144 server-perf is preserved as a dropped HTTP 400 context-fit/refusal cell and must not be substituted.
- memory block not recorded at test time
- flags:
HO-009 promoted MTP n_max=2 fingerprint: same model/runtime/backend/KV/batch as cfg-0132 plus native model MTP/NextN path with --spec-draft-n-max 2. Depth performance uses raw llama-server /completion timing because this runtime's llama-bench exposes no MTP/spec flags; no ngram/external drafter, no KV quant, no proxy/llama-swap, no n_max>=4. - chat template: not recorded
- capability: measured
cfg-0135 — Qwen3.6-35B-A3B-MTP-UD-Q4_K_M.gguf @ UD-Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@7077abbe14c510cb829c93a1328c2815b5805ebd (ROCm0/gfx1151)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Stage A activation/safety arm only. Positive draft counters were admitted, but n3 was not promoted because n2 was faster/cleaner; no n_max>=4 safety claim is made.
- memory block not recorded at test time
- flags:
HO-009 MTP n_max=3 Stage A fingerprint: same as cfg-0132 plus native model MTP/NextN path with --spec-draft-n-max 3. No ngram/external drafter, no KV quant, no proxy/llama-swap, no n_max>=4. - chat template: not recorded
- capability: measured
cfg-0136 — Qwen3.8-27B-Q8_0.gguf @ Q8_0
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@0b0f35d0ed745bd6e40eb248e320cacbb1a546ab (Vulkan0/Strix-Halo)
- derived runtime.tree: fork — carrying con-0002
- backfilled reconstructed from the archive; Reviewer-admitted HO-004 negative Stage A record from /home/aihydra/bench-results/ho004-qwen38-dflash2-matched-r1. The arm loaded/responded but failed the Stage A sentinel gate on code/toolish prompts; guard and Stage B performance were not run, and no throughput or production claim is admitted.
- memory block not recorded at test time
- flags:
HO-004 plain/control Stage A fingerprint: /home/aihydra/src/strix-halo-v065-run-20260820195059, asset e9e8d0067604b02ff3d69946ede758a4708d6b75791aa55f8340c0ae43c5a234, manifest 034b652092d91f54a55f6fb7bc0101a2669da3864fd5d732a562ca8fc2826f11, Mesa d18d598e2, Qwen3.8-27B Q8_0 target, --load-mode none, -ngl 999, flash attention on, -b 2048, -ub 512, f16/f16 KV, --parallel 1, temp 0, seed 4242, no speculation, no MTP, no DFlash2 drafter, no ngram/proxy/llama-swap/ROCm. - chat template: not recorded
- capability: measured
cfg-0137 — Qwen3.8-27B-Q8_0.gguf @ Q8_0
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@0b0f35d0ed745bd6e40eb248e320cacbb1a546ab (Vulkan0/Strix-Halo)
- derived runtime.tree: fork — carrying con-0002
- backfilled reconstructed from the archive; Stage A activation/safety arm only. Positive draft acceptance counters were reviewer-admitted, but the arm failed the Stage A sentinel gate on code/toolish prompts; no guard or Stage B performance cell is admitted and no n_max>=4 safety claim is made.
- memory block not recorded at test time
- flags:
HO-004 stock MTP n_max=3 Stage A fingerprint: same target/runtime/backend/KV/batch/load-mode as cfg-0136 plus native draft-mtp path with --spec-type draft-mtp --spec-draft-n-max 3. No DFlash2 drafter, no ngram/proxy/llama-swap/ROCm, no stock MTP n_max>=4, no alternate target quant, no KV quant substitution. - chat template: not recorded
- capability: measured
cfg-0138 — Qwen3.8-27B-Q8_0.gguf + Qwen3.8-27B-DFlash2-Q8_0.gguf @ Q8_0 + DFlash2-Q8_0
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@0b0f35d0ed745bd6e40eb248e320cacbb1a546ab (Vulkan0/Strix-Halo)
- derived runtime.tree: fork — carrying con-0002, con-0003
- backfilled reconstructed from the archive; Stage A activation/safety arm only. Positive DFlash2 draft acceptance counters were reviewer-admitted, but the arm failed the Stage A sentinel gate on code/toolish prompts; no guard, Stage B performance, energy-efficiency, or production recommendation claim is admitted.
- memory block not recorded at test time
- flags:
HO-004 DFlash2 Q8_0 width-4 Stage A fingerprint: same target/runtime/backend/KV/batch/load-mode as cfg-0136 plus external drafter /home/aihydra/models/qwen38-27b-dflash2/Qwen3.8-27B-DFlash2-Q8_0.gguf with -md DFlash2-Q8_0 -ngld 999 --spec-type draft-dflash --spec-draft-n-max 4. No stock MTP, no ngram/proxy/llama-swap/ROCm, no alternate target quant, no KV quant substitution. - chat template: not recorded
- capability: measured
cfg-0139 — GLM-4.7-Flash-Q4_K_M.gguf @ Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (ROCm0/Strix-Halo)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Reviewer-admitted HG-004 Phase-1 boundary diagnostic only from /home/aihydra/bench-results/hg004-glm47-phase1-boundary-r1/hg004-glm47-phase1-boundary-r1-20260820T231235Z and /Users/wardenmac/.openclaw/workspace/halobench/reconcile/HG-004-phase1-boundary-review.md. The record preserves one raw llama-server HTTP 200 response capture whose normalized assistant content was empty with no tool_calls, non-empty reasoning_content, and finish_reason length. Tau2 Phase-2 was not launched, so no score, task pass count, throughput, capability verdict, or production recommendation is represented.
- memory block not recorded at test time
- flags:
HG-004 Phase-1 boundary diagnostic fingerprint: raw /home/aihydra/src/llama.cpp/build/bin/llama-server on 127.0.0.1:8104, llama.cpp checkout ce7689f875f724ece08b1a9fbb2a83971120b4d2 with clean tree, server binary sha256 17f55ef427a56ebd98aabf0d3c48e43a93e384e1db5075c5c7bed5ed6a0aad01, version 1 (3653e6d), GLM-4.7-Flash Q4_K_M model sha256 29837ed2c0fc5f51981adf8ac8083fcf80743c598381f13e9f06cbad0498b174 bytes 18312339808, --load-mode none, --no-mmap, -ngl 999, -dev ROCm0, -fa on, -ctk f16, -ctv f16, -c 32768, -n 4096, --parallel 1, --jinja, --reasoning-format auto, pinned Tau2/native-template task-1 path with frozen and embedded template sha256 d63ad536c3c81880043e22ec7fd08db42b4d8fb7c89c7138bc562bfa25281375, task-1 sha256 365091613cf731aaadac059d7a39b32b95a070bf358b4bd9f04ad836eec1a614, tools-14 sha256 bdba5499d11340d02e92688d62f942f2f702667368d04e52772520aa8fa7ea22, simulator openrouter/anthropic/claude-haiku-4.5 attempts=1. No tau2 Phase-2, benchmark, throughput, or capability result is admitted from this diagnostic. - chat template: d63ad536c3c81880043e22ec7fd08db42b4d8fb7c89c7138bc562bfa25281375
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0140 — Ornith-1.5-35B-Q4_K_M.gguf @ Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; HO-011 reviewer-admitted from /home/aihydra/bench-results/ho011-ornith15-q4km-preflight-r6 (screening) and /home/aihydra/bench-results/ho011-ornith15-q4km-full26-r1 (full-26 tau2). The Q4_K_M quant class is measured on its own boundary; it does not inherit Ornith-1.0 Q8_0 or UD-Q4_K_XL capability, does not establish Ornith-1.0/1.5 equivalence, and does not create an MTP/speculation claim. Energy windows are recorded but not joined — this hbreviewer profile lacks Home Assistant credentials.
- memory block not recorded at test time
- flags:
Serving/capability boundary: -ngl 999 -fa on -b 2048 -ub 512 -c 32768 -dev ROCm0 -ctk f16 -ctv f16 -t 16 --load-mode none --parallel 1 --jinja -n 4096. Guard at depth 8000, tau2 airline with openrouter/anthropic/claude-haiku-4.5 simulator, temp 0, seed 42, num_trials 1, max_steps 200, max_errors 50. Bounded throughput cells at d262144 pp1024/tg256 N=5 fresh-process reps under this same artifact/runtime/backend/f16-KV boundary. Model artifact: /home/aihydra/models/ornith-15-35b/Ornith-1.5-35B-Q4_K_M.gguf, 21,713,462,848 bytes, sha256 ca6ea26329c88b78ffd90a85163be2e746c2fafd1024f56db47e499f117f9a7f, source ornith-ai/Ornith-1.5-35B-A3B-GGUF. Tau2 head: 668d3bcd135c02aa3438f987ef45735b7c163ee3. - chat template: unknown
- capability: measured
cfg-0141 — Qwen3.6-35B-A3B-MTP-UD-Q4_K_M.gguf @ UD-Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@7077abbe14c510cb829c93a1328c2815b5805ebd (ROCm0/gfx1151)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; HU-004 reviewer-admitted long-prompt guard control arm only. Guard-cleared evidence (all long-prompt fixtures non-empty content / non-empty required tool arguments, no silent empty success) is a guard-pass result, not a capability, throughput, or production result. This config does not establish agentic/tau2 capability or served-path production performance.
- memory block not recorded at test time
- flags:
HU-004 long-prompt guard plain/control arm on raw llama-server: model /home/aihydra/models/qwen3.6-35b-a3b-mtp/Qwen3.6-35B-A3B-UD-Q4_K_M.gguf, -ngl 999 -dev ROCm0 -fa on -ctk f16 -ctv f16 -c 204800 --load-mode none -b 2048 -ub 512 --parallel 1 --host 127.0.0.1 --port 8114 --spec-type none. Server binary build-rocm/bin/llama-server, build 50 commit 7077abb — the same binary used by HO-009 cfg-0132/cfg-0134. No proxy/llama-swap/ngram, no MTP/NextN on this plain control arm. - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0142 — Qwen3.6-35B-A3B-MTP-UD-Q4_K_M.gguf @ UD-Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@7077abbe14c510cb829c93a1328c2815b5805ebd (ROCm0/gfx1151)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; HU-004 reviewer-admitted long-prompt guard target arm only. Guard-cleared evidence is a guard-pass result, not a capability, throughput, or production result. No n2 d262144 substitution and no n_max>=4 safety claim are made.
- memory block not recorded at test time
- flags:
HU-004 long-prompt guard n2 target arm on raw llama-server: model /home/aihydra/models/qwen3.6-35b-a3b-mtp/Qwen3.6-35B-A3B-UD-Q4_K_M.gguf, -ngl 999 -dev ROCm0 -fa on -ctk f16 -ctv f16 -c 204800 --load-mode none -b 2048 -ub 512 --parallel 1 --host 127.0.0.1 --port 8114 --spec-type draft-mtp --spec-draft-n-max 2 -ctkd f16 -ctvd f16. Server binary build-rocm/bin/llama-server, build 50 commit 7077abb — the same binary/path used by HO-009 cfg-0134 with native model MTP/NextN. No ngram/external drafter, no proxy/llama-swap, no n_max>=4. - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0143 — Akicou/inclusionAI_LLaDA2.2-flash-GGUF/LLaDA2.2-flash-Q4_K_S.gguf @ Q4_K_S
- host aihydra · engine igpu · runtime Akicou/diffuse-cpp@b799157 (rocm)
- derived runtime.tree: fork — carrying con-0005
- memory block not recorded at test time
- flags:
-ngl 99 -t 8 -s 32 --cache --cache-ctx 16384 --host 127.0.0.1 --port 8080 - chat template: not recorded
- capability: measured
cfg-0144 — Qwen3.6-35B-A3B-MTP-UD-Q4_K_M.gguf @ UD-Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@7077abbe14c510cb829c93a1328c2815b5805ebd (ROCm0/gfx1151)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; The -rea off delta is the deliberate DIAG-recommended fix, not a deviation. This cfg is scoped to the plain-vs-n2 capability comparison on build 7077abb only; no production/quality/throughput equivalence claim.
- memory block not recorded at test time
- flags:
HO-009-AB PLAIN (reasoning-OFF) capability fingerprint: same model/runtime/ backend as cfg-0132 (f16/f16 KV, -ngl 999 -fa on -c 32768 --load-mode none --parallel 1 --jinja -b 2048 -ub 512 -t 16, no spec/MTP flags) PLUS the tau2-protocol agent settings and -rea off. Reasoning disabled on the tau2 path to remove the llm_utils.generate() reasoning_content-drop confound (DIAG verdict, HO-009-DIAG t_fa64ef03). temp 0, seed 42, max_tokens 4096. - chat template: not recorded
- capability: measured
cfg-0145 — Qwen3.6-35B-A3B-MTP-UD-Q4_K_M.gguf @ UD-Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@7077abbe14c510cb829c93a1328c2815b5805ebd (ROCm0/gfx1151)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; The -rea off delta is the deliberate DIAG-recommended fix, not a deviation. Scoped to the plain-vs-n2 capability comparison on build 7077abb only. No n2 quality-equivalence, production recommendation, or throughput claim beyond the plain-vs-n2 capability comparison on this card.
- memory block not recorded at test time
- flags:
HO-009-AB n2/MTP (reasoning-OFF) capability fingerprint: same model/runtime/ backend as cfg-0134 plus -rea off on the tau2 path. MTP engaged via --spec-type draft-mtp --spec-draft-n-max 2 -ctkd f16 -ctvd f16 (draft KV stays f16 = no KV quant, matching clm-0097/cfg-0134; --spec-draft-n-max 2 alone is a silent no-op on build 7077abb, per hborchestrator fingerprint resolution). f16/f16 target KV, -ngl 999 -fa on -c 32768 --load-mode none --parallel 1 --jinja -b 2048 -ub 512 -t 16. temp 0, seed 42, max_tokens 4096. n_max=2 ONLY; no n_max>=4, no KV quant, no alt backend. - chat template: not recorded
- capability: measured
cfg-0146 — Qwen3.8-27B-Q8_0.gguf @ Q8_0
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@2586f6eddae19bef3dd21a1a0b109cca7bcf1c32 (Vulkan0/Strix-Halo)
- derived runtime.tree: fork — carrying con-0002, con-0009
- backfilled reconstructed from the archive; Recorded from the runner command line and MANIFEST.txt at review time; the full build-commit/Mesa/source lineage is in the artifact MANIFEST (mesa d18d598e2). Sentinal outcome measured: plain/control PASS Stage-A (empty-assistant 0, empty-tool-call 0, single-token-EOS 0, code-valid 3/3, tool-correct 3/3) plus the 4/4 house guard, so the null-update diagnostic resolved. This config is ONLY a Stage-A sentinel/gate record: it carries no throughput, no Stage-B production-depth performance, and no capability score.
- memory block not recorded at test time
- flags:
HO-004-DIAG v0.6.10 plain/control Stage-A sentinel fingerprint (re-tests the v0.5.x base-platform empty-content defect after the v0.6.10 MTP-rollback and DSpark re-land): /home/aihydra/src/strix-halo-v0610/vulkan/llama-server, MANIFEST.txt "Built 2026-08-22T10:47:53Z source 2586f6eddae19bef3dd21a1a0b109cca7bcf1c32", target Qwen3.8-27B Q8_0 (sha256 57484e...), --load-mode none, -ngl 999, -fa on, -b 2048, -ub 512, f16/f16 KV, --parallel 1, --jinja, -c 32768, --temp 0, --seed 4242, no speculation, no MTP, no DFlash2 drafter, no ngram/proxy/llama-swap/ROCm. Same portable Vulkan family and KV/batch/load-mode as cfg-0136 but on the v0.6.10 build. - chat template: not recorded
- capability: measured
cfg-0147 — Qwen3.8-27B-Q8_0.gguf @ Q8_0
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@2586f6eddae19bef3dd21a1a0b109cca7bcf1c32 (Vulkan0/Strix-Halo)
- derived runtime.tree: fork — carrying con-0002, con-0009
- backfilled reconstructed from the archive; Recorded from the runner command line at review time. Sentinal outcome: stock-MTP n_max=3 PASS Stage-A (empty-assistant 0, empty-tool-call 0, single-token-EOS 0, code-valid 3/3, tool-correct 3/3) plus the 4/4 house guard, with log-backed draft activation (draft acceptance 0.8218/0.9531). This config is ONLY a Stage-A sentinel/gate record: no throughput, no Stage-B production-depth performance, and no MTP safety claim beyond n_max=3.
- memory block not recorded at test time
- flags:
HO-004-DIAG v0.6.10 stock-MTP Stage-A sentinel fingerprint: identical target/runtime/backend/KV/batch/load-mode/-c 32768/--temp 0/--seed 4242 as cfg-0146 plus the native draft-mtp path with --spec-type draft-mtp --spec-draft-n-max 3 (HARD ceiling 3; n_max>=4 corrupts generations on this family, clm-0055). No DFlash2 drafter, no ngram/proxy/llama-swap/ROCm, no n_max>=4, no KV quant, no alternate target quant. - chat template: not recorded
- capability: measured
cfg-0148 — Qwen3.8-27B-Q8_0.gguf + Qwen3.8-27B-DFlash2-Q8_0.gguf @ Q8_0 + DFlash2-Q8_0
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@2586f6eddae19bef3dd21a1a0b109cca7bcf1c32 (Vulkan0/Strix-Halo)
- derived runtime.tree: fork — carrying con-0002, con-0009, con-0003
- backfilled reconstructed from the archive; Recorded from the runner command line at review time. Sentinal outcome: DFlash2 Q8_0 n_max=3 PASS Stage-A (empty-assistant 0, empty-tool-call 0, single-token-EOS 0, code-valid 3/3, tool-correct 3/3) plus the 4/4 house guard, with log-backed DFlash2 draft activation (draft acceptance 0.8505). This config is ONLY a Stage-A sentinel/gate record: no throughput, no Stage-B production-depth performance, and no DFlash2 quality-equivalence or throughput-improvement claim (G2 varied-prompt invariance is NEGATIVE).
- memory block not recorded at test time
- flags:
HO-004-DIAG v0.6.10 DFlash2 Stage-A sentinel fingerprint: identical target/runtime/backend/KV/batch/load-mode/-c 32768/--temp 0/--seed 4242 as cfg-0146 plus the external DFlash2 drafter /home/aihydra/models/qwen38-27b-dflash2/Qwen3.8-27B-DFlash2-Q8_0.gguf with -md DFlash2-Q8_0 -ngld 999 --spec-type draft-dflash --spec-draft-n-max 3 (card forbids n_max>=4; prior matched-r1 used 4). No stock MTP, no ngram/proxy/llama-swap/ROCm, no KV quant, no alternate target quant. - chat template: not recorded
- capability: measured
cfg-0149 — Qwen3.6-35B-A3B-MTP-UD-Q4_K_M.gguf @ UD-Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@2586f6eddae19bef3dd21a1a0b109cca7bcf1c32 (Vulkan0/Strix-Halo (strix-halo-v0610 fork via RADV, NOT ggml-org ROCm0))
- derived runtime.tree: fork — carrying con-0002, con-0009
- backfilled reconstructed from the archive; Recorded from the raw tau2 artifacts and the v0.6.10 fork MANIFEST (payload source 2586f6edd, Vulkan/RADV). The -rea off delta is the deliberate DIAG-recommended fix already applied in cfg-0144/0145. Scoped to the v0.6.10-fork-on-Vulkan vs 7077abb-upstream-on-ROCm0 capability recovery question only; no DFlash2/stock-MTP/n_max>=4/reasoning-on/alt-model/ production-throughput claim.
- memory block not recorded at test time
- flags:
HO-009-v0610 PLAIN (reasoning-OFF) capability fingerprint on the v0.6.10 fork/Vulkan runtime, mirroring cfg-0144 flags exactly (f16/f16 KV, -ngl 999 -fa on -c 32768 --load-mode none --parallel 1 --jinja -b 2048 -ub 512 -t 16, no spec/MTP flags) PLUS the tau2-protocol agent settings and -rea off on the tau2 path. temp 0, seed 42, max_tokens 4096. Same model/quant/flags as cfg-0144 but built on the v0.6.10 fork under Vulkan/RADV instead of upstream 7077abb ROCm0/gfx1151 — repo + backend + build all changed together, so this is a stack effect, not a build-only change. - chat template: not recorded
- capability: measured
cfg-0150 — Qwen3.6-35B-A3B-MTP-UD-Q4_K_M.gguf @ UD-Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@2586f6eddae19bef3dd21a1a0b109cca7bcf1c32 (Vulkan0/Strix-Halo (strix-halo-v0610 fork via RADV, NOT ggml-org ROCm0))
- derived runtime.tree: fork — carrying con-0002, con-0009
- backfilled reconstructed from the archive; Recorded from the raw tau2 artifacts and the v0.6.10 fork MANIFEST (payload source 2586f6edd, Vulkan/RADV). The -rea off delta is the deliberate DIAG-recommended fix from cfg-0145. Scoped to the v0.6.10-fork-on-Vulkan vs 7077abb-upstream-on-ROCm0 capability recovery question only; no DFlash2/stock-MTP/n_max>=4/reasoning-on/alt-model/production-throughput claim.
- memory block not recorded at test time
- flags:
HO-009-v0610 n2/MTP (reasoning-OFF) capability fingerprint on the v0.6.10 fork/Vulkan runtime, mirroring cfg-0145 (same model/quant/flags as cfg-0149 PLUS --spec-type draft-mtp --spec-draft-n-max 2 -ctkd f16 -ctvd f16; draft KV stays f16 = no KV quant). tau2-protocol agent settings, -rea off on the tau2 path, f16/f16 target KV, -ngl 999 -fa on -c 32768 --load-mode none --parallel 1 --jinja -b 2048 -ub 512 -t 16, temp 0 seed 42 max_tokens 4096. n_max=2 ONLY; no n_max>=4, no KV quant, no alt drafter. Built on the v0.6.10 fork under Vulkan/RADV vs upstream 7077abb ROCm0 — repo + backend + build all changed together (stack effect). - chat template: not recorded
- capability: measured
cfg-0151 — Qwen3.5-122B-A10B-UD-Q4_K_M-00001-of-00003.gguf @ UD-Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@ce7689f (ROCm0/gfx1151)
- derived runtime.tree: upstream
- memory block not recorded at test time
- flags:
HO-012 production-optimisation matrix Cell-1 baseline (b2048/ub512, q8_0/q8_0 KV) + Cell-2 depth probe fingerprint. llama-server build ce7689f (llama.cpp-kvfix tree), -ngl 999 -fa 1 -c 32768 --parallel 1 --load-mode none --jinja -rea off -ctk q8_0 -ctv q8_0 -b 2048 -ub 512 -t 16 -n 4096, ROCm0 Radeon 8060S gfx1151 122880 MiB. Cell-2 depth via llama-bench -p 16384 -n 256 -d 262144 -r 1 on the same b2048/ub512 q8_0/q8_0 signature. From per-execution command files (server-c1-b2048-ub512.command.txt and cell2-depth/c2-d262144.command.txt). - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0152 — Qwen3.5-122B-A10B-UD-Q4_K_M-00001-of-00003.gguf @ UD-Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@ce7689f (ROCm0/gfx1151)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; From reviewer-admitted HO-012 evidence. batch/ubatch is a perf lever; this is a throughput/matrix arm only, no capability measurement or claim. q8_0/q8_0 KV, b1024/ub512, build ce7689f, ROCm0. Intra-suite comparison only.
- memory block not recorded at test time
- flags:
HO-012 Cell-1 batch/ubatch sweep arm b1024/ub512: llama-server build ce7689f (llama.cpp-kvfix), -ngl 999 -fa 1 -c 32768 --parallel 1 --load-mode none --jinja -rea off -ctk q8_0 -ctv q8_0 -b 1024 -ub 512 -t 16 -n 4096, ROCm0 Radeon 8060S gfx1151 122880 MiB. Identical to cfg-0151 EXCEPT -b 1024. From per-execution server command file. - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0153 — Qwen3.5-122B-A10B-UD-Q4_K_M-00001-of-00003.gguf @ UD-Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@ce7689f (ROCm0/gfx1151)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; From reviewer-admitted HO-012 evidence. batch/ubatch is a perf lever; this is a throughput-only production-optimisation config, no capability measurement or claim. q8_0/q8_0 KV, b512/ub512, build ce7689f, ROCm0. Intra-suite comparison only — no cross-model or equivalency claim.
- memory block not recorded at test time
- flags:
HO-012 Cell-1 batch/ubatch sweep arm b512/ub512: llama-server on build ce7689f (llama.cpp-kvfix), -ngl 999 -fa 1 -c 32768 --parallel 1 --load-mode none --jinja -rea off -ctk q8_0 -ctv q8_0 -b 512 -ub 512 -t 16 -n 4096, ROCm0 Radeon 8060S gfx1151 122880 MiB. Identical to cfg-0151 EXCEPT -b 512. From server-c1-b512-ub512.command.txt. - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0154 — Qwen3.5-122B-A10B-UD-Q4_K_M-00001-of-00003.gguf @ UD-Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@ce7689f (ROCm0/gfx1151)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; From reviewer-admitted HO-012 evidence. batch/ubatch is a perf lever; throughput/matrix arm only, no capability claim. q8_0/q8_0 KV. Intra-suite comparison only.
- memory block not recorded at test time
- flags:
HO-012 Cell-1 batch/ubatch sweep arm b512/ub256: llama-server build ce7689f (llama.cpp-kvfix), -ngl 999 -fa 1 -c 32768 --parallel 1 --load-mode none --jinja -rea off -ctk q8_0 -ctv q8_0 -b 512 -ub 256 -t 16 -n 4096, ROCm0 Radeon 8060S gfx1151 122880 MiB. Identical to cfg-0151 EXCEPT -b 512 -ub 256. From server-c1-b512-ub256.command.txt. - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0155 — Qwen3.5-122B-A10B-UD-Q4_K_M-00001-of-00003.gguf @ UD-Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@ce7689f (ROCm0/gfx1151)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; From reviewer-admitted HO-012 evidence. batch/ubatch is a perf lever; matrix arm only, no capability claim. q8_0/q8_0 KV. Intra-suite comparison only.
- memory block not recorded at test time
- flags:
HO-012 Cell-1 batch/ubatch sweep arm b512/ub128: llama-server build ce7689f (llama.cpp-kvfix), -ngl 999 -fa 1 -c 32768 --parallel 1 --load-mode none --jinja -rea off -ctk q8_0 -ctv q8_0 -b 512 -ub 128 -t 16 -n 4096, ROCm0 Radeon 8060S gfx1151 122880 MiB. Identical to cfg-0151 EXCEPT -b 512 -ub 128. From server-c1-b512-ub128.command.txt. - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0156 — Qwen3.5-122B-A10B-UD-Q4_K_M-00001-of-00003.gguf @ UD-Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@ce7689f (ROCm0/gfx1151)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; From reviewer-admitted HO-012 evidence. batch/ubatch is a perf lever; matrix arm only, no capability claim. q8_0/q8_0 KV. Intra-suite comparison only.
- memory block not recorded at test time
- flags:
HO-012 Cell-1 batch/ubatch sweep arm b256/ub256: llama-server build ce7689f (llama.cpp-kvfix), -ngl 999 -fa 1 -c 32768 --parallel 1 --load-mode none --jinja -rea off -ctk q8_0 -ctv q8_0 -b 256 -ub 256 -t 16 -n 4096, ROCm0 Radeon 8060S gfx1151 122880 MiB. Identical to cfg-0151 EXCEPT -b 256 -ub 256. From server-c1-b256-ub256.command.txt. - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0157 — Qwen3.5-122B-A10B-UD-Q4_K_M-00001-of-00003.gguf @ UD-Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@ce7689f (ROCm0/gfx1151)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; From reviewer-admitted HO-012 evidence. kv_quant is a LOSSY capability lever, so capability is NOT inherited — this arm is a throughput/delta cell only, guarded 4/4, no capability claim. f16/f16 KV. b2048/ub512. No cross-model or equivalence claim.
- memory block not recorded at test time
- flags:
HO-012 Cell-3 KV f16/f16 arm: llama-server build ce7689f (llama.cpp-kvfix), -ngl 999 -fa 1 -c 32768 --parallel 1 --load-mode none --jinja -rea off -ctk f16 -ctv f16 -b 2048 -ub 512 -t 16 -n 4096, ROCm0 Radeon 8060S gfx1151 122880 MiB. Identical to cfg-0151 EXCEPT -ctk f16 -ctv f16 (KV quant lever). From server-c3-f16.command.txt. - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0158 — Qwen3.5-122B-A10B-UD-Q4_K_M-00001-of-00003.gguf @ UD-Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@ce7689f (ROCm0/gfx1151)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; From reviewer-admitted HO-012 evidence. kv_quant is a LOSSY capability lever, so capability is NOT inherited; this arm is a throughput/delta cell only, guarded 4/4, no capability claim. q4_0/q4_0 KV, b2048/ub512. Intra-suite only; the q4_0 decode regression vs f16 is recorded truthfully (see clm-0111), never as equivalence.
- memory block not recorded at test time
- flags:
HO-012 Cell-3 KV q4_0 arm:llama-server build ce7689f (llama.cpp-kvfix), -ngl 999 -fa 1 -c 32768 --parallel 1 --load-mode none --jinja -rea off -ctk q4_0 -ctv q4_0 -b 2048 -ub 512 -t 16 -n 4096, ROCm0 Radeon 8060S gfx1151 122880 MiB. Identical to cfg-0151 EXCEPT -ctk q4_0 -ctv q4_0 (KV quant lever). From server-c3-q4_0.command.txt. - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0159 — Ornith-1.0-35B-UD-Q4_K_XL.gguf @ UD-Q4_K_XL
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@2586f6ed (Vulkan0/Strix-Halo (strix-halo-v0610 fork via RADV, NOT ggml-org ROCm0))
- derived runtime.tree: fork — carrying con-0002, con-0009
- backfilled reconstructed from the archive; Ingested from HO-005-REVISED-REVIEW (t_b3f466d1) ADMIT verdict. Control arm on the v0.6.10 fork (build 2586f6ed Vulkan, PRIMARY). 34.357 tps, CV 0.073%. Decode-only (pp0/tg256) — no prefill figure. Guard run-0591 (4/4) covers the signature. Interactive-decode tuning cell only; build swap is a binary lever so capability is not inherited; no capability, cross-model or production-equivalence claim.
- memory block not recorded at test time
- flags:
HO-005-REVISED interactive-decode batch/ubatch control arm b512/ub512 at d131072, Vulkan v0.6.10 fork, KV f16/f16, -ngl 999 -fa 1 -dev Vulkan0 --load-mode none --parallel 1 -t 16, pp0/tg256 decode-only, plain decode (no DQ/KV-quant/MTP). llama-bench -o json -r 5, fresh process per cell. - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0160 — Qwen3.6-35B-A3B-MTP-UD-Q4_K_M.gguf @ UD-Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@2586f6eddae19bef3dd21a1a0b109cca7bcf1c32 (Vulkan0/Strix-Halo (strix-halo-v0610 fork via RADV, NOT ggml-org ROCm0))
- derived runtime.tree: fork — carrying con-0002, con-0009
- backfilled reconstructed from the archive; Recorded from raw cell1-n4-off/fingerprint.txt (Vulkan/RADV build 2586f6edd, model sha 0b21525e, spec=n4 reasoning=off, base_flags + spec_flags recorded verbatim) and gates.txt G3/G6. ADMITTED full-26 capability arm (17/26, mean 0.653846) via HO-013-REVIEW. Scope strict to this single model/quant/build, MTP n2/n4 + reasoning on/off matrix only; no cross-model/KV-quant/DFlash2/n>4 claim.
- memory block not recorded at test time
- flags:
HO-013 cell1 n4-off (MTP n_max=4, reasoning OFF) capability fingerprint on the v0.6.10 fork/Vulkan runtime 2586f6edd. Mirrors cfg-0150 (f16/f16 KV, -ngl 999 -fa 1 -c 32768 --load-mode none --parallel 1 --jinja -rea off -n 4096 --slots -ctk f16 -ctv f16 -b 2048 -ub 512 -t 16, temp 0 seed 42 max_tokens 4096) but --spec-draft-n-max 4 (--spec-type draft-mtp -ctkd f16 -ctvd f16; draft KV stays f16 = no KV quant). n_max=4 ONLY; no reasoning-on, no KV quant, no alt drafter. - chat template: not recorded
- capability: measured
cfg-0161 — Qwen3.6-35B-A3B-MTP-UD-Q4_K_M.gguf @ UD-Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@2586f6eddae19bef3dd21a1a0b109cca7bcf1c32 (Vulkan0/Strix-Halo (strix-halo-v0610 fork via RADV, NOT ggml-org ROCm0))
- derived runtime.tree: fork — carrying con-0002, con-0009
- backfilled reconstructed from the archive; DISPOSITION-ONLY. cell2-n2-ON tau2 was never started: the reasoning-ON path was stopped at the plain-rea-ON sentinel G6 FAIL (empty assistant content, see cfg-0162 run + clm-0113), so this arm was not exercised. Recorded to document the intended cell arm per the HO-013 matrix; carries NO score, NO capability, NO sentinel result. Scope strict to single model/quant/build.
- memory block not recorded at test time
- flags:
HO-013 cell2-n2-rea-ON planned fingerprint: MTP n_max=2 (--spec-type draft-mtp --spec-draft-n-max 2 -ctkd f16 -ctvd f16) with reasoning ON (-rea on) on the v0.6.10 fork/Vulkan runtime 2586f6edd, same f16/f16 KV -ngl 999 -fa 1 -c 32768 --load-mode none --parallel 1 --jinja -n 4096 --slots -ctk f16 -ctv f16 -b 2048 -ub 512 -t 16 base. Planned as cell2-n2-ON. - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0162 — Qwen3.6-35B-A3B-MTP-UD-Q4_K_M.gguf @ UD-Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@2586f6eddae19bef3dd21a1a0b109cca7bcf1c32 (Vulkan0/Strix-Halo (strix-halo-v0610 fork via RADV, NOT ggml-org ROCm0))
- derived runtime.tree: fork — carrying con-0002, con-0009
- backfilled reconstructed from the archive; DISPOSITION-ONLY. This plain-rea-ON control FAILED the reasoning-content sentinel (G6 FAIL, rc=2): assistant_content_len=0 (EMPTY) despite reasoning_content_len=1846, num_tool_calls=2, finish_reason=tool_calls. The reasoning-ON path is BLOCKED on this fork/stack — no tau2 was run on any reasoning-ON arm, and the whole cell2/cell3 reasoning set was stopped/cancelled (see clm-0113). No capability/score. Scope strict to single model/quant/build.
- memory block not recorded at test time
- flags:
HO-013 cell2-plain-rea-ON sentinel fingerprint: PLAIN (NO MTP spec flags) with reasoning ON (-rea on) on the v0.6.10 fork/Vulkan runtime 2586f6edd. sentinel-rea fingerprint records base_flags: -ngl 999 -fa 1 -c 32768 --load-mode none --parallel 1 --jinja -rea on -n 4096 --slots -ctk f16 -ctv f16 -b 2048 -ub 512 -t 16, reasoning=on, spec=plain (NO spec_flags). This is the reasoning-content sentinel control arm. - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0163 — Qwen3.6-35B-A3B-MTP-UD-Q4_K_M.gguf @ UD-Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@2586f6eddae19bef3dd21a1a0b109cca7bcf1c32 (Vulkan0/Strix-Halo (strix-halo-v0610 fork via RADV, NOT ggml-org ROCm0))
- derived runtime.tree: fork — carrying con-0002, con-0009
- backfilled reconstructed from the archive; DISPOSITION-ONLY (gated-CANCELLED). cell3-n4-ON was always gated behind the reasoning-content sentinel result; because cfg-0162's plain-rea-ON control FAILED G6, this gated arm was correctly CANCELLED before any run. Recorded to document the intended cell arm per the HO-013 matrix; carries NO score, NO capability, NO sentinel result. Scope strict to single model/quant/build.
- memory block not recorded at test time
- flags:
HO-013 cell3-n4-ON planned (gated) fingerprint: MTP n_max=4 (--spec-type draft-mtp --spec-draft-n-max 4 -ctkd f16 -ctvd f16) with reasoning ON (-rea on) on the v0.6.10 fork/Vulkan runtime 2586f6edd, same f16/f16 KV -ngl 999 -fa 1 -c 32768 --load-mode none --parallel 1 --jinja -n 4096 --slots -ctk f16 -ctv f16 -b 2048 -ub 512 -t 16 base. - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0164 — Qwen3.8-27B-IU4-Kairic-Edge.gguf @ IU4
- host aihydra · engine igpu · runtime ciru-ai/ROCmFPX@e1da26bb8 (rocm)
- derived runtime.tree: fork — carrying con-0012
- backfilled reconstructed from the archive; Fingerprint reconstructed from the HK-001-FULLBENCH launch script and build doc (hk001-results-r1.md + scratch/hk001-launch.sh) — exact start_server invocation not re-verified on-box (the runner's logs are on aihydra, build commit e1da26bb8 / binary e121a4d3; model sha 360caf73; sidecar hashes adcbb90a / 82f93131 / 3b07e7b1). Capability is measured (guard exercised 3/3 + isolation N/A), everything on the launched served config; the throughput/depth cells are engine-path (llama-bench). Single config record for this build; every HK-001-FULLBENCH run below points at it.
- memory block not recorded at test time
- flags:
The Kairic Edge TheRock served config at c32768 on this build (compat mode; prompts even fast-greedy forbids tool calls) used by HK-001-FULLBENCH. Generic flag set (per the fullbench launch script): -m Qwen3.8-27B-IU4-Kairic-Edge.gguf -dev ROCm0 -ngl 999 -c 32768 -b 2048 -ub 512 -fa on -ctk f16 -ctv f16 -t 16 -tb 32 -np 1 --no-mmap --kairic-edge --metrics -n 4096 --slots with the PROMPTFORGE_* env block (PROMPTFORGE_MODE=iu4_ffn, IU4/GDN/GDN-Output sidecar pfs paths, HADAMARD+SEGMENTED switches). PLAIN spec only — NO MTP (n_max>=4 cliff on the Qwen3.8 family, clm-0055). --kairic-edge and the sidecars are server-only; llama-bench (the engine path that produced the depth-ladder cells) can NOT carry the flag or the GDN row, so the depth cells are recorded as engine-path on this same build+sidecar-env, per protocol 1d. - chat template: not recorded
- capability: measured
cfg-0165 — Qwen3.6-35B-A3B-MTP-UD-Q4_K_M.gguf @ UD-Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@2586f6eddae19bef3dd21a1a0b109cca7bcf1c32 (Vulkan0/Strix-Halo (strix-halo-v0610 fork via RADV, NOT ggml-org ROCm0))
- derived runtime.tree: fork — carrying con-0002, con-0009
- backfilled reconstructed from the archive; From HO-014 reviewer-admitted FILL evidence (t_dc235dcd ADMIT; runner command c1-b2048-ub512.command.txt, summary.tsv). batch/ubatch is a perf lever; the house guard (4/4) clears binary/hazard levers only and no new capability claim is made or inherited. Model sha/bytes re-verified by RUN gate G0 (0b21525e... 22,663,387,424 B). Intra-matrix comparison only.
- memory block not recorded at test time
- flags:
HO-014 production-optimisation Cell-1 batch/ubatch arm b2048/ub512 (also the Cell-1 baseline / guard fingerprint). llama-bench pp512/tg256 N=5 d131072, -m Qwen3.6-35B-A3B-UD-Q4_K_M.gguf -ngl 999 -fa on -b 2048 -ub 512 --load-mode none -dev Vulkan0 -ctk f16 -ctv f16 -t 16 -p 512 -n 256 -d 131072 -r 5, build 2586f6edd Vulkan/RADV, capability c32768, plain mode (-rea off, no MTP). Same model/quant/flags as cfg-0149 but under the stris-halo v0.6.10 fork/Vulkan; this record fixes the batch/ubatch sweep arm for the production-optimisation matrix, not a capability measurement. - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0166 — Qwen3.6-35B-A3B-MTP-UD-Q4_K_M.gguf @ UD-Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@2586f6eddae19bef3dd21a1a0b109cca7bcf1c32 (Vulkan0/Strix-Halo (strix-halo-v0610 fork via RADV, NOT ggml-org ROCm0))
- derived runtime.tree: fork — carrying con-0002, con-0009
- backfilled reconstructed from the archive; From HO-014 reviewer-admitted FILL + RUN evidence (t_dc235dcd ADMIT: summary.tsv best-ttft b1024/ub512 1353.1 ms; cell2-summary.tsv; RUN cell3 f16 arm at this config). batch/ubatch + depth are perf levers; the house guard (4/4, RUN + FILL both) clears binary/hazard levers only and no new capability claim is made or inherited. Intra-matrix/production-optimisation only; no cross-model/build/backend claim.
- memory block not recorded at test time
- flags:
HO-014 PRODUCTION config b1024/ub512 f16/f16 KV — the Cell-1 best-TTFT arm (Cell-1 tdepths), and the base for the Cell-2 depth probe (d131072→d200000→ d262144) and the RUN Cell-3 f16 arm. llama-bench pp512/tg256 N=5 d131072 for the Cell-1 latency arm (command c1-b1024-ub512.command.txt); pp16384/tg256 N=3 for the Cell-2 depth ladder. -m Qwen3.6-35B-A3B-UD-Q4_K_M.gguf -ngl 999 -fa on -b 1024 -ub 512 --load-mode none -dev Vulkan0 -ctk f16 -ctv f16 -t 16, build 2586f6edd Vulkan/RADV, plain mode (-rea off, no MTP). Model sha/bytes re-verified by RUN gate G0. Same model/quant/flags as cfg-0149 but the production b/ub choice for interactive serving. - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0167 — Qwen3.6-35B-A3B-MTP-UD-Q4_K_M.gguf @ UD-Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@2586f6eddae19bef3dd21a1a0b109cca7bcf1c32 (Vulkan0/Strix-Halo (strix-halo-v0610 fork via RADV, NOT ggml-org ROCm0))
- derived runtime.tree: fork — carrying con-0002, con-0009
- backfilled reconstructed from the archive; From HO-014 reviewer-admitted FILL evidence (runner command c1-b512-ub512.command.txt, summary.tsv). batch/ubatch is a perf lever; guard (4/4) clears binary/hazard levers only, no capability claim. Intra-matrix only.
- memory block not recorded at test time
- flags:
HO-014 Cell-1 batch/ubatch arm b512/ub512 (f16/f16 KV). llama-bench pp512/tg256 N=5 d131072, -m Qwen3.6-35B-A3B-UD-Q4_K_M.gguf -ngl 999 -fa on -b 512 -ub 512 --load-mode none -dev Vulkan0 -ctk f16 -ctv f16 -t 16 -p 512 -n 256 -d 131072 -r 5, build 2586f6edd Vulkan/RADV, plain mode (-rea off, no MTP). - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0168 — Qwen3.6-35B-A3B-MTP-UD-Q4_K_M.gguf @ UD-Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@2586f6eddae19bef3dd21a1a0b109cca7bcf1c32 (Vulkan0/Strix-Halo (strix-halo-v0610 fork via RADV, NOT ggml-org ROCm0))
- derived runtime.tree: fork — carrying con-0002, con-0009
- backfilled reconstructed from the archive; From HO-014 reviewer-admitted FILL evidence (runner command c1-b512-ub256.command.txt, summary.tsv). batch/ubatch is a perf lever; guard (4/4) clears binary/hazard levers only, no capability claim. Intra-matrix only.
- memory block not recorded at test time
- flags:
HO-014 Cell-1 batch/ubatch arm b512/ub256 (f16/f16 KV). llama-bench pp512/tg256 N=5 d131072, -m Qwen3.6-35B-A3B-UD-Q4_K_M.gguf -ngl 999 -fa on -b 512 -ub 256 --load-mode none -dev Vulkan0 -ctk f16 -ctv f16 -t 16 -p 512 -n 256 -d 131072 -r 5, build 2586f6edd Vulkan/RADV, plain mode (-rea off, no MTP). - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0169 — Qwen3.6-35B-A3B-MTP-UD-Q4_K_M.gguf @ UD-Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@2586f6eddae19bef3dd21a1a0b109cca7bcf1c32 (Vulkan0/Strix-Halo (strix-halo-v0610 fork via RADV, NOT ggml-org ROCm0))
- derived runtime.tree: fork — carrying con-0002, con-0009
- backfilled reconstructed from the archive; From HO-014 reviewer-admitted FILL evidence (runner command c1-b512-ub128.command.txt, summary.tsv). batch/ubatch is a perf lever; guard (4/4) clears binary/hazard levers only, no capability claim. Intra-matrix only.
- memory block not recorded at test time
- flags:
HO-014 Cell-1 batch/ubatch arm b512/ub128 (f16/f16 KV). llama-bench pp512/tg256 N=5 d131072, -m Qwen3.6-35B-A3B-UD-Q4_K_M.gguf -ngl 999 -fa on -b 512 -ub 128 --load-mode none -dev Vulkan0 -ctk f16 -ctv f16 -t 16 -p 512 -n 256 -d 131072 -r 5, build 2586f6edd Vulkan/RADV, plain mode (-rea off, no MTP). - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0170 — Qwen3.6-35B-A3B-MTP-UD-Q4_K_M.gguf @ UD-Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@2586f6eddae19bef3dd21a1a0b109cca7bcf1c32 (Vulkan0/Strix-Halo (strix-halo-v0610 fork via RADV, NOT ggml-org ROCm0))
- derived runtime.tree: fork — carrying con-0002, con-0009
- backfilled reconstructed from the archive; From HO-014 reviewer-admitted FILL evidence (runner command c1-b256-ub256.command.txt, summary.tsv). batch/ubatch is a perf lever; guard (4/4) clears binary/hazard levers only, no capability claim. Intra-matrix only.
- memory block not recorded at test time
- flags:
HO-014 Cell-1 batch/ubatch arm b256/ub256 (f16/f16 KV). llama-bench pp512/tg256 N=5 d131072, -m Qwen3.6-35B-A3B-UD-Q4_K_M.gguf -ngl 999 -fa on -b 256 -ub 256 --load-mode none -dev Vulkan0 -ctk f16 -ctv f16 -t 16 -p 512 -n 256 -d 131072 -r 5, build 2586f6edd Vulkan/RADV, plain mode (-rea off, no MTP). - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0171 — Qwen3.6-35B-A3B-MTP-UD-Q4_K_M.gguf @ UD-Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@2586f6eddae19bef3dd21a1a0b109cca7bcf1c32 (Vulkan0/Strix-Halo (strix-halo-v0610 fork via RADV, NOT ggml-org ROCm0))
- derived runtime.tree: fork — carrying con-0002, con-0009
- backfilled reconstructed from the archive; From HO-014-RUN reviewer-admitted Cell-3 evidence (cell3-verdict.md; G3 arch pre-probe PASS 11 quantisable-KV layers blk3..40 step4; bench pair + KLD). KV quant is a LOSSY capability lever, so capability is NOT inherited — this arm is a throughput/quality delta cell only. Guarded 4/4 (run-0615). No cross-model or equivalence claim; the small q4_0 delta is partly ARCHITECTURAL (hybrid: only 11/40 layers carry a quantisable full-attention KV).
- memory block not recorded at test time
- flags:
HO-014 RUN Cell-3 q4_0/q4_0 KV arm at the production config b1024/ub512 d131072, f16 baseline comparison arm is cfg-0166. llama-bench pp512/tg256 N=5 d131072, -m Qwen3.6-35B-A3B-UD-Q4_K_M.gguf -ngl 999 -fa on -b 1024 -ub 512 --load-mode none -dev Vulkan0 -ctk q4_0 -ctv q4_0 -t 16 -p 512 -n 256 -d 131072 -r 5, build 2586f6edd Vulkan/RADV, plain mode (-rea off, no MTP, no spec). Identical to cfg-0166 EXCEPT -ctk q4_0 -ctv q4_0 (KV quant lever). The Cell-3A KLD (f16 logits, replay q4_0) was measured at capability context c32768, both arms identical, after the initial ctx=131072 arm hit an instrument logits-buffer memory wall (~96 GiB RSS, earlyoom — not model OOM). - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0172 — Qwen3.8-27B-IU4-Kairic-Edge.gguf @ IU4
- host aihydra · engine igpu · runtime ciru-ai/ROCmFPX@e1da26bb8 (rocm)
- derived runtime.tree: fork — carrying con-0012
- memory block not recorded at test time
- flags:
The Kairic Edge TheRock served config at c32768 with reasoning DISABLED (-rea off) — the fix path clm-0116 prescribed for the reasoning-ON empty-assistant content-drop defect, used by HK-RERUN-REASONOFF. Server command verified on-box (tau2-reasonoff-full26-20260823T182525Z/server.command.txt): llama-server -m Qwen3.8-27B-IU4-Kairic-Edge.gguf --host 127.0.0.1 --port 8124 --jinja -dev ROCm0 -ngl 999 -c 32768 -b 2048 -ub 512 -fa on -ctk f16 -ctv f16 -t 16 -tb 32 -np 1 -ctxcp 32 --cache-ram 8192 --cache-prompt --cache-idle-slots --kairic-edge --no-mmap -rea off --metrics -n 4096 --slots, with the PROMPTFORGE_* sidecar env (iu4_ffn + GDN rows). Served model id: qwen3.8-27b-kairic-edge (served-model-id.txt). PLAIN spec only — NO MTP. The specsweep arm additionally ran the SAME server both WITH --kairic-edge (spec=on) and WITHOUT it (spec=off) at rungs {0, 32768, 131072, 204800}; gdn_output_fallback=0 across all 20 sweep cells (specsweep gdn_fallback.txt) and =48 for the tau2 full-26 server. - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0173 — Qwen3.8-27B-IU4-Kairic-Edge.gguf @ IU4
- host aihydra · engine igpu · runtime ciru-ai/ROCmFPX@e1da26bb8 (rocm)
- derived runtime.tree: fork — carrying con-0012
- memory block not recorded at test time
- flags:
HK-001-REASONING-MATRIX 5-cell served config on this build, one llama-server per cell at c32768 with PROMPTFORGE_* sidecar env (iu4_ffn + GDN rows), port 8125, -dev ROCm0, kv f16/f16 (per fingerprint.txt). Per-cell arms vary ONLY max_tokens / reasoning flag / reasoning budget: R0-B4k max_tokens=4096 + "-rea off" (chat_template_kwargs enable_thinking=false); R1-B6k max_tokens=6144, reasoning on, --reasoning-budget -1; R2-B8k max_tokens=8192, reasoning on, --reasoning-budget -1; R3-B8k-L max_tokens=8192, reasoning on, --reasoning-budget 1024 (+ enable_thinking=true); R4-B8k-M max_tokens=8192, reasoning on, --reasoning-budget 2048. tau2 agent args temperature=0.0; airline domain, seed 42, tasks 0-4, claude-haiku-4.5 simulator. gdn_output_fallback=48 (matches baseline). - chat template: not recorded
- capability: inherited from cfg-0164 across neutral lever(s):
cfg-0174 — Qwen3.8-27B-ROCMFPX-MQ-Q4.gguf + Qwen3.8-27B-DFlash2-Q8_0.gguf @ ROCmFPX-MQ-Q4 + DFlash2-Q8_0
- host aihydra · engine igpu · runtime LaurentZuijdwijk/llama.cpp@16f0799a6 (Vulkan0/RADV Mesa 26.0.3)
- derived runtime.tree: fork — carrying con-0013
- backfilled reconstructed from the archive; The archived handoff does not provide target/drafter SHA256s, a chat-template hash, a standard house-guard artifact, a plain/spec-off paired served-path floor, or a llama-bench floor. The server samples are retained as screening evidence, not a production-performance or quant-equivalence admission.
- memory block not recorded at test time
- flags:
Operator-directed autonomous tuning served-path screen: target ROCmFPX-MQ-Q4 with DFlash2 Q8_0 drafter, -dev Vulkan0 -ngl 99 -c 32768 -fa on -t 16 -np 1 --jinja --spec-type draft-dflash --spec-draft-n-max 7 -ngld 99. The separate competence smoke used native MTP n_max 4 with reasoning medium, temperature 0.1, max_tokens 8192, tasks 0-4, seed 42, and the pinned Haiku simulator. - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0175 — Qwen3.8-27B-ROCMFPX-MQ-Q4.gguf + Qwen3.8-27B-DFlash2-Q8_0.gguf @ ROCmFPX-MQ-Q4 + DFlash2-Q8_0
- host aihydra · engine igpu · runtime LaurentZuijdwijk/llama.cpp@16f0799a6 (Vulkan0/RADV Mesa 26.0.3)
- derived runtime.tree: fork — carrying con-0013
- backfilled reconstructed from the archive; The trace provides no artifact SHA256s, chat-template hash, standard house-guard result, or independent answer grading. The script's displayed depth column is the newly-prefilled-token count, not a machine-recorded cumulative-context field.
- memory block not recorded at test time
- flags:
Operator-directed autonomous cache-reuse trace: target ROCmFPX-MQ-Q4 with DFlash2 Q8_0, -dev Vulkan0 -ngl 99 -c 98304 -fa on -t 16 -np 1 --jinja --spec-type draft-dflash --spec-draft-n-max 7 -ngld 99. Each completion request used temperature 0, n_predict 80, and cache_prompt true while appending a new synthetic section. - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0176 — Qwen3.8-27B-UD-Q4_K_XL.gguf @ UD-Q4_K_XL
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@7077abb (ROCm0 / TheRock ROCm 7.14)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; The archived speed JSON has one repetition per target depth and reports the actual prompt token count, which is materially below the requested target at both retained deep samples. It has no artifact SHA256, template hash, house guard, plain/spec-off paired floor, or energy window; it cannot establish a production recommendation.
- memory block not recorded at test time
- flags:
Operator-directed autonomous served-path depth samples: -dev ROCm0 -ngl 999 -c 204800 -fa on -ctk q8_0 -ctv q8_0 -t 16 -tb 32 -np 1 --load-mode none --jinja --reasoning on --reasoning-effort medium --reasoning-format deepseek --spec-type draft-mtp --spec-draft-n-max 3. - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0177 — Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf @ UD-Q4_K_XL
- host aihydra · engine igpu · runtime unslothai/llama.cpp@250b61446efc91e3a179c8677956f2667c8fbda0 (vulkan)
- derived runtime.tree: upstream
- memory block not recorded at test time
- flags:
llama-server -m <UD-Q4_K_XL 4-shard set> -ngl 999 -fa 1 -b 2048 -ub 512 -c 32768 -dev Vulkan0 -ctk f16 -ctv f16 -t 16 --load-mode none --parallel 1 --jinja --reasoning-format deepseek --chat-template-kwargs '{"reasoning_effort":"low"}' -n 4096. The n-gram/PLE table tensor (per_layer_token_embd, 26.8 GiB iq4_nl) is placed on CPU AUTOMATICALLY by the Vulkan backend ("cannot be used with preferred buffer type Vulkan_Host, using CPU instead") — so -ot per_layer_token_embd=CPU is a NO-OP and is omitted; GTT-resident footprint = 76.7 GiB with ~43 GiB free for KV. tau2 agent args temperature=0.0, max_tokens=4096; airline domain, seed 42, tasks 0-25, claude-haiku-4.5 simulator (sim_self_play=false), max-steps 200. - chat template: not recorded
- capability: measured
cfg-0178 — Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf @ UD-Q4_K_XL
- host aihydra · engine igpu · runtime unslothai/llama.cpp@250b61446efc91e3a179c8677956f2667c8fbda0 (vulkan)
- derived runtime.tree: upstream
- memory block not recorded at test time
- flags:
llama-bench decode-depth ladder: -ngl 999 -fa 1 -ctk f16 -ctv f16 -dev Vulkan0 -lm none -b 2048 -ub 512 -p 512 -n 128 -d {0,32768,131072} -r 3 in one matrix invocation, plus d204800 -r 3 as a separate clean run. Fresh-process per cell. NOTE: d245760 failed as a cell in the 0/32768/131072/245760 matrix invocation (opaque — rc=0, truncated output, no stderr error) but d204800 (>200k production depth) ran clean standalone; 262144 is omitted (Vulkan workgroup-count GGML_ASSERT at ggml-vulkan.cpp fires > ~262140). - chat template: not recorded
- capability: inherited from cfg-0177 across neutral lever(s):
cfg-0179 — Qwen3.8-Flash-Next-Q4_0-ROCmFP4-STRIX.gguf @ Q4_0_ROCmFP4_STRIX
- host aihydra · engine igpu · runtime kingjones30/ROCmFPX@36e9acd40e10a87cd3c3ef8ec734668757dc8520 (HIP/gfx1151)
- derived runtime.tree: fork — carrying con-0017
- memory block not recorded at test time
- flags:
No model command executed. The fixed intended uninstrumented serving boundary was one slot, --n-gpu-layers 999, --flash-attn on, --fit off, --jinja, runtime-default mmap, and no MTP, draft or speculation. The separately required receipt-only capture variant attempted a Release HIP/gfx1151/native targeted llama-server build and stopped rc=2 because examples/gguf-hash/deps/sha256/sha256.c could not include rotate-bits/rotate-bits.h, despite the exact tracked header being present in the source tree. Resolved template, thinking, KV, batch, ubatch, cache and thread values were never produced because the campaign stopped before server launch. - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0180 — Qwen3.8-Flash-Next-Q4_0-ROCmFP4-STRIX.gguf @ Q4_0_ROCmFP4_STRIX
- host aihydra · engine igpu · runtime kingjones30/ROCmFPX@36e9acd40e10a87cd3c3ef8ec734668757dc8520 (HIP/gfx1151)
- derived runtime.tree: fork — carrying con-0017
- memory block not recorded at test time
- flags:
Corrected served cohort: llama-server -c 262144 -ngl 999 -fa on -fit off -ctk f16 -ctv f16 -b 2048 -ub 512 -t 16 --jinja -rea off --ctx-checkpoints 512 --checkpoint-every-n-tokens 2048 -np 1. One slot, temperature 0, no MTP/draft/speculation. Runtime-default mmap is an explicit model-intrinsic exception: this artifact's PLE path depends on lazy file-backed mapping and is not comparable to the house --load-mode none floor. - chat template: not recorded
- capability: measured
cfg-0181 — Qwen3.8-Flash-Next-Q4_0-ROCmFP4-STRIX.gguf @ Q4_0_ROCmFP4_STRIX
- host aihydra · engine igpu · runtime kingjones30/ROCmFPX@36e9acd40e10a87cd3c3ef8ec734668757dc8520 (HIP/gfx1151)
- derived runtime.tree: fork — carrying con-0017
- memory block not recorded at test time
- flags:
Source-config/default-mmap llama-bench engine-only rows: -ngl 999 -fa 1, f16 K/V, b2048/ub512, t16, pp512/tg128 at d0, d32768, d131072, d204800, and d240000; N=3 at d0/d32768 and N=1 deeper. This instrument does not represent the corrected served QSA path. Runtime-default mmap is disclosed and not equivalent to the ordinary HaloBench --load-mode none floor. - chat template: not recorded
- capability: inherited from cfg-0180 across neutral lever(s):
cfg-0182 — Qwen3.8-Flash-Next-ROCmFP4-FAST-v2-ple16.gguf @ ROCmFP4-FAST-v2-ple16
- host aihydra · engine igpu · runtime LaurentZuijdwijk/llama.cpp@5e085d123eead2e89b5c19f824fccb05727da6a2 (vulkan)
- derived runtime.tree: upstream
- memory block not recorded at test time
- flags:
llama-server -m Qwen3.8-Flash-Next-ROCmFP4-FAST-v2-ple16.gguf -ngl 99 -fa on -c 200192 -b 2048 -ub 2048 -ctk q8_0 -ctv q8_0 --parallel 1 --jinja --metrics --spec-type draft-mtp -md Qwen3.8-Flash-Next-MTP-ROCmFP4-FAST.gguf --spec-draft-adaptive; env GGML_VK_ALLOW_GRAPHICS_QUEUE=1; port 5830. The ~51B-parameter n-gram/PLE table (per_layer_token_embd) is baked per-head into the ple16 quant and stays fully GPU-resident (no -ot offload needed). Capability run (run-0653) used -c 196608 directly on llama-server; the decode-depth ladder (run-0654..run-0646) ran through the slotpin energy-metering proxy at -c 200192 (near-identical). tau2 agent args temperature=0.0, max_tokens=4096, airline seed 42, tasks 0-25, claude-haiku-4.5 simulator (sim_self_play=false), max-steps 200, single slot (concurrency 1) with adaptive MTP. - chat template: not recorded
- capability: measured
cfg-0183 — Qwen3.8-27B-UD-Q4_K_XL.gguf @ UD-Q4_K_XL
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@c530ea79c753573c98796ebfaccb1aed8f5897b5 (rocm)
- derived runtime.tree: fork — carrying con-0018
- memory block not recorded at test time
- flags:
-ngl 999 -fa on -c 49152 -b 2048 -ub 2048 -ctk f16 -ctv f16 --parallel 4 --cont-batching --slots --slot-save-path /var/lib/slotpin/kv-cache --cache-ram 8192 --reasoning-effort low --spec-type draft-mtp --spec-draft-n-max 2 --jinja --metrics --host 127.0.0.1 --port 5802 (fronted by the slotpin KV-warmth proxy on :8090, slots=4 max_inflight=4) - chat template: not recorded
- capability: measured
cfg-0184 — Qwen3.8-27B-UD-Q4_K_XL.gguf @ UD-Q4_K_XL
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@c530ea79c753573c98796ebfaccb1aed8f5897b5 (rocm)
- derived runtime.tree: fork — carrying con-0018
- memory block not recorded at test time
- flags:
-ngl 999 -fa on -c 262144 -b 2048 -ub 2048 -ctk f16 -ctv f16 --parallel 1 --slots --slot-save-path /var/lib/slotpin/kv-cache --cache-ram 8192 --reasoning-effort low --spec-type draft-mtp --spec-draft-n-max 2 --jinja --metrics --host 127.0.0.1 --port 5802 (fronted by the slotpin proxy on :8090, slots=1 max_inflight=1) - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0185 — Qwen3.8-27B-IQ4_XS-ALL-IMATRIX-Q8-OUT-MTP.gguf @ IQ4_XS-ALL-IMATRIX-Q8-OUT
- host aihydra · engine igpu · runtime pwilkin/llama.cpp@d3b5cc43d1fcfce891f2de94d5274ee40eceb21c (rocm)
- derived runtime.tree: fork — carrying con-0018, con-0019
- memory block not recorded at test time
- flags:
-dev ROCm0 -ngl 999 -fa on -fit off --load-mode none -c 49152 -b 2048 -ub 512 --parallel 4 --cont-batching --jinja --reasoning-effort low --spec-type draft-dflash --spec-draft-model Qwen3.8-27B-DFlash2-IQ4_XS.gguf --spec-draft-device ROCm0 --spec-draft-ngl 99 --spec-draft-n-min 0 --spec-draft-n-max 3 --spec-draft-p-min 0.10 --metrics --port 5803 (runtime env: ENABLE_RETAINED_PM4=1 -> DEBUG_HIP_GRAPH_PM4=1, HSA_OVERRIDE_GFX_VERSION=11.5.1, GGML_HIP_ENABLE_UNIFIED_MEMORY=1; custom ROCr+HIP runtime via LD_LIBRARY_PATH) - chat template: not recorded
- capability: measured
cfg-0186 — Qwen3.8-27B-IQ4_XS-ALL-IMATRIX-Q8-OUT-MTP.gguf @ IQ4_XS-ALL-IMATRIX-Q8-OUT
- host aihydra · engine igpu · runtime pwilkin/llama.cpp@d3b5cc43d1fcfce891f2de94d5274ee40eceb21c (rocm)
- derived runtime.tree: fork — carrying con-0018, con-0019
- memory block not recorded at test time
- flags:
-dev ROCm0 -ngl 999 -fa on -fit off --load-mode none -c 262144 -b 2048 -ub 512 --parallel 1 --jinja --spec-type draft-dflash --spec-draft-model Qwen3.8-27B-DFlash2-IQ4_XS.gguf --spec-draft-device ROCm0 --spec-draft-ngl 99 --spec-draft-n-min 0 --spec-draft-n-max 3 --spec-draft-p-min 0.10 --metrics --port 5803 (retained-PM4 runtime env as cfg-0185; fronted by slotpin :8090) - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0187 — Qwen3.8-27B-UD-Q4_K_XL.gguf @ UD-Q4_K_XL
- host aihydra · engine igpu · runtime pwilkin/llama.cpp@d3b5cc43d1fcfce891f2de94d5274ee40eceb21c (rocm)
- derived runtime.tree: fork — carrying con-0018, con-0019
- memory block not recorded at test time
- flags:
-dev ROCm0 -ngl 999 -fa on -c 49152 -b 2048 -ub 2048 -ctk f16 -ctv f16 --parallel 1 --jinja --reasoning-effort low --spec-type draft-mtp --spec-draft-n-max 2 --metrics --port 5803 (isolation probe: runtime env toggled ENABLE_RETAINED_PM4=1 [run-0672] vs =0 [run-0673]) - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
Runs (672)
Split by what is being established. Capability is what the model can do; performance is how fast it delivers that; a guard is the cheap cliff-detector that makes a performance number worth recording. Seeprotocol §1.
capability (61)
run-0003 — conceptual-c1-c8-blind@v1
- 2026-06-28 · config → cfg-0001
- depth: 0 tokens (conversational — says nothing about behaviour at 100k+) · scoring: blind-human
- time_to_correct_answer_s=1.7
run-0098 — tau2-bench-airline@v1.0.1
run-0099 — tau2-bench-airline-nothink@v1.0.1
run-0100 — tau2-bench-airline@668d3bc
run-0101 — tau2-bench-airline@668d3bc
run-0102 — tau2-bench-airline@668d3bc
run-0103 — tau2-bench-airline@668d3bc
run-0104 — tau2-bench-airline@668d3bc
run-0105 — tau2-bench-airline@668d3bc
run-0106 — tau2-bench-airline@668d3bc
run-0107 — tau2-bench-airline@668d3bc
run-0108 — tau2-bench-airline@668d3bc
run-0109 — tau2-bench-airline@668d3bc
run-0110 — tau2-bench-airline@668d3bc
run-0111 — tau2-bench-airline@668d3bc
run-0271 — tau2-bench-airline@668d3bc
run-0272 — kv-quality-patched
run-0273 — kv-quality-patched
run-0274 — nemotron3-super-q4km-screen
run-0287 — tau2-bench-airline@668d3bc
run-0316 — tau2-bench-airline@668d3bc
run-0329 — tau2-bench-airline@668d3bc
run-0331 — ornith-35b-screen
run-0345 — tau2-bench-airline@668d3bc
run-0358 — tau2-bench-airline@668d3bc
run-0360 — tau2-bench-airline@668d3bc
run-0372 — deep-thought-posttrain-fullbench
run-0432 — tau2-bench-airline@668d3bc
run-0433 — tau2-bench-airline@668d3bc
run-0439 — tau2-bench-airline@668d3bc
run-0468 — tau2-bench-airline@668d3bc
run-0469 — tau2-bench-airline@668d3bc
run-0470 — tau2-bench-airline@668d3bc
run-0471 — HG-003-lfm2-retrieval-promotion
- 2026-08-19 · config → cfg-0126
- depth: 2,048 tokens · scoring: deterministic
- runner_invalid=true
- ⚠ contaminated by: inc-0007
run-0472 — HG-003-lfm2-retrieval-promotion
run-0473 — HG-003-lfm2-retrieval-promotion
run-0495 — ho005-ornith-q4kxl-vulkan-preflight-r3
run-0499 — tau2-bench-airline@668d3bc
run-0546 — tau2-bench-airline@668d3bc
- 2026-08-21 · config → cfg-0140
- depth: 32,768 tokens · scoring: deterministic
- tasks_passed=4 · tasks_total=5 · mean_score=0.8
run-0550 — tau2-bench-airline@668d3bc
- 2026-08-21 · config → cfg-0140
- depth: 32,768 tokens · scoring: deterministic
- tasks_passed=8 · tasks_total=26 · mean_score=0.32
run-0553 — tau2-bench-airline@668d3bc
run-0555 — tau2-bench-airline@668d3bc
run-0556 — tau2-bench-airline@668d3bc
run-0569 — tau2-bench-airline@668d3bc
run-0571 — tau2-bench-airline@668d3bc
run-0576 — tau2-bench-airline@668d3bc
run-0578 — tau2-bench-airline@668d3bc
run-0599 — tau2-bench-airline@668d3bc
run-0618 — hk001-rerun-reasonoff.tau2-full26
run-0621 — hk001-reasoning-matrix.R0-B4k
run-0622 — hk001-reasoning-matrix.R1-B6k
run-0623 — hk001-reasoning-matrix.R2-B8k
run-0624 — hk001-reasoning-matrix.R3-B8k-L
run-0625 — hk001-reasoning-matrix.R4-B8k-M
run-0627 — tau2-airline-smoke@tasks-0-4
- 2026-08-24 · config → cfg-0174
- depth: 32,768 tokens · scoring: llm-judged (judge: claude-haiku-4.5)
- tasks_passed=5 · tasks_total=5 · mean_score=1
run-0632 — tau2-airline-full26
run-0638 — tau2-airline-smoke5
- 2026-08-29 · config → cfg-0180
- depth: 262,144 tokens · scoring: deterministic
- empty_arg_calls=0 · tasks_passed=5 · tasks_total=5 · mean_score=1
- ⚠ contaminated by: inc-0010
run-0639 — tau2-airline-full26
run-0653 — tau2-airline-full26
run-0658 — tau2-bench-airline@668d3bc
run-0665 — tau2-bench-airline@668d3bc
guard (76)
run-0004 — tool-loop@v2.1
- 2026-06-27 · config → cfg-0001
- tool_call_success=0.94 · turns_before_degradation=10+ · empty_arg_calls=0
run-0005 — capability-guard@v2
- 2026-08-08 · config → cfg-0006
- tool_call_success=1 · empty_arg_calls=0 · tasks_passed=4 · tasks_total=4
run-0267 — eos-cliffguard@1
run-0275 — eos-cliff-q8kv@1
run-0276 — eos-cliff-q8kv@1
run-0277 — eos-cliff-isolation@1
run-0278 — eos-cliff-isolation@1
run-0279 — eos-cliff-isolation@1
run-0280 — eos-cliff-isolation@1
run-0281 — eos-cliff-isolation@1
run-0282 — eos-cliff-isolation@1
run-0283 — eos-cliff-isolation@1
run-0284 — eos-cliff-isolation@1
run-0285 — eos-cliff-isolation@1
run-0286 — eos-cliff-isolation@1
run-0304 — eos-cliff-isolation@1
run-0305 — eos-cliff-isolation@1
run-0306 — house-capability-guard@1
run-0307 — house-capability-guard@1
run-0328 — house-capability-guard@1
run-0330 — ornith-35b-screen
run-0338 — ornith-35b-fullbench
run-0357 — qwen36-27b-mtp-fullbench
run-0359 — qwen36-27b-mtp-fullbench
run-0363 — deep-thought-posttrain-fullbench
run-0395 — house-capability-guard@1
run-0434 — house-capability-guard@1
run-0440 — house-capability-guard@1
run-0452 — house-capability-guard@1
run-0454 — house-capability-guard@1
run-0456 — house-capability-guard@1
run-0458 — house-capability-guard@1
run-0460 — house-capability-guard@1
run-0462 — house-capability-guard@1
run-0464 — house-capability-guard@1
run-0466 — house-capability-guard@1
run-0494 — ho005-ornith-q4kxl-vulkan-preflight-r3
run-0498 — guard-c32768-depth8000
run-0523 — mtp-context-proof
run-0528 — guard-c32768-depth8000
run-0529 — guard-c32768-depth8000
run-0541 — stageA-c32768-varied-prompts
run-0542 — stageA-c32768-varied-prompts
run-0543 — stageA-c32768-varied-prompts
run-0544 — hg004-phase1-boundary-diagnostic
run-0551 — guard-hu004-longprompt-8kto40k
- 2026-08-21 · config → cfg-0141
- tool_call_success=1 · empty_arg_calls=0 · tasks_passed=10 · tasks_total=10
run-0552 — guard-hu004-longprompt-8kto40k
- 2026-08-21 · config → cfg-0142
- tool_call_success=1 · empty_arg_calls=0 · tasks_passed=10 · tasks_total=10
run-0572 — stageA-c32768-varied-prompts
run-0573 — stageA-c32768-varied-prompts
run-0574 — stageA-c32768-varied-prompts
run-0591 — guard-c32768-depth8000
run-0597 — eos-cliff-c32768-varied-fullprompt
- 2026-08-23 · config → cfg-0160
- tasks_passed=8 · tasks_total=8
run-0605 — ho014a-fill-r1.guard-c32768-depth8000
- 2026-08-23 · config → cfg-0165
- tasks_passed=4 · tasks_total=4
run-0615 — ho014-run-kv-overlap.guard-c32768-depth8000
- 2026-08-23 · config → cfg-0166
- tasks_passed=4 · tasks_total=4
run-0637 — q38fn-kingjones-r2-targeted-receipt-build-gate-v1
- 2026-08-29 · config → cfg-0179
- runner_invalid=true
- ⚠ contaminated by: inc-0009
run-0660 — empty-output-monitor@pr27311
- 2026-09-05 · config → cfg-0183
- tool_call_success=1 · turns_before_degradation=none — 0 EMPTY over 25 probes across 106 min (13:42-15:23Z) · empty_arg_calls=0
- parallel 4 · prompts varied — multi-slot under speculation shares an n-gram pool across slots (clm-0016)
run-0666 — empty-output-monitor@dflash-pm4
- 2026-09-06 · config → cfg-0185
- tool_call_success=1 · turns_before_degradation=none — 0 EMPTY over 16 probes across 66 min (09:27-10:33Z) · empty_arg_calls=0
- parallel 4 · prompts varied — multi-slot under speculation shares an n-gram pool across slots (clm-0016)
performance (535)
run-0001 — tool-loop@v2.1
run-0002 — mtp-decode-sweep@v1
- 2026-07-20 · config → cfg-0002
- decode_tps=36.5 · prefill_tps=144 · ttft_ms=0 · empty_arg_calls=0
- ⚠ no guard — No guard was run. MTP is a neutral lever so capability loss is not expected by construction, but that is an argument, not a measurement.
run-0007 — spec-ab@v1
run-0008 — llama-bench@1
- 2026-08-08 · config → cfg-0008
- prefill_tps=312.037487
- ⚠ no guard — Guarded at run-0005 for this model+build+backend at --parallel 1. llama-bench cells vary KV-quant (lossy) and flash-attn (binary), across which capability cannot be inherited, and llama-bench serves no endpoint to guard directly. Recorded as a known gap rather than an implied pass.
- energy window: eng-0002
run-0009 — llama-bench@1
- 2026-08-08 · config → cfg-0008
- decode_tps=21.676081
- ⚠ no guard — Guarded at run-0005 for this model+build+backend at --parallel 1. llama-bench cells vary KV-quant (lossy) and flash-attn (binary), across which capability cannot be inherited, and llama-bench serves no endpoint to guard directly. Recorded as a known gap rather than an implied pass.
- energy window: eng-0002
run-0010 — llama-bench@1
- 2026-08-08 · config → cfg-0008
- prefill_tps=283.473901
- ⚠ no guard — Guarded at run-0005 for this model+build+backend at --parallel 1. llama-bench cells vary KV-quant (lossy) and flash-attn (binary), across which capability cannot be inherited, and llama-bench serves no endpoint to guard directly. Recorded as a known gap rather than an implied pass.
- energy window: eng-0002
run-0011 — llama-bench@1
- 2026-08-08 · config → cfg-0008
- decode_tps=19.258807
- ⚠ no guard — Guarded at run-0005 for this model+build+backend at --parallel 1. llama-bench cells vary KV-quant (lossy) and flash-attn (binary), across which capability cannot be inherited, and llama-bench serves no endpoint to guard directly. Recorded as a known gap rather than an implied pass.
- energy window: eng-0002
run-0012 — llama-bench@1
- 2026-08-08 · config → cfg-0008
- prefill_tps=170.811067
- ⚠ no guard — Guarded at run-0005 for this model+build+backend at --parallel 1. llama-bench cells vary KV-quant (lossy) and flash-attn (binary), across which capability cannot be inherited, and llama-bench serves no endpoint to guard directly. Recorded as a known gap rather than an implied pass.
- energy window: eng-0002
run-0013 — llama-bench@1
- 2026-08-08 · config → cfg-0008
- decode_tps=9.867545
- ⚠ no guard — Guarded at run-0005 for this model+build+backend at --parallel 1. llama-bench cells vary KV-quant (lossy) and flash-attn (binary), across which capability cannot be inherited, and llama-bench serves no endpoint to guard directly. Recorded as a known gap rather than an implied pass.
- energy window: eng-0002
run-0014 — llama-bench@1
- 2026-08-08 · config → cfg-0009
- prefill_tps=321.722166
- ⚠ no guard — Guarded at run-0005 for this model+build+backend at --parallel 1. llama-bench cells vary KV-quant (lossy) and flash-attn (binary), across which capability cannot be inherited, and llama-bench serves no endpoint to guard directly. Recorded as a known gap rather than an implied pass.
- energy window: eng-0003
run-0015 — llama-bench@1
- 2026-08-08 · config → cfg-0009
- decode_tps=21.910456
- ⚠ no guard — Guarded at run-0005 for this model+build+backend at --parallel 1. llama-bench cells vary KV-quant (lossy) and flash-attn (binary), across which capability cannot be inherited, and llama-bench serves no endpoint to guard directly. Recorded as a known gap rather than an implied pass.
- energy window: eng-0003
run-0016 — llama-bench@1
- 2026-08-08 · config → cfg-0009
- prefill_tps=300.673419
- ⚠ no guard — Guarded at run-0005 for this model+build+backend at --parallel 1. llama-bench cells vary KV-quant (lossy) and flash-attn (binary), across which capability cannot be inherited, and llama-bench serves no endpoint to guard directly. Recorded as a known gap rather than an implied pass.
- energy window: eng-0003
run-0017 — llama-bench@1
- 2026-08-08 · config → cfg-0009
- decode_tps=21.393789
- ⚠ no guard — Guarded at run-0005 for this model+build+backend at --parallel 1. llama-bench cells vary KV-quant (lossy) and flash-attn (binary), across which capability cannot be inherited, and llama-bench serves no endpoint to guard directly. Recorded as a known gap rather than an implied pass.
- energy window: eng-0003
run-0018 — llama-bench@1
- 2026-08-08 · config → cfg-0009
- prefill_tps=208.771771
- ⚠ no guard — Guarded at run-0005 for this model+build+backend at --parallel 1. llama-bench cells vary KV-quant (lossy) and flash-attn (binary), across which capability cannot be inherited, and llama-bench serves no endpoint to guard directly. Recorded as a known gap rather than an implied pass.
- energy window: eng-0003
run-0019 — llama-bench@1
- 2026-08-08 · config → cfg-0009
- decode_tps=18.174796
- ⚠ no guard — Guarded at run-0005 for this model+build+backend at --parallel 1. llama-bench cells vary KV-quant (lossy) and flash-attn (binary), across which capability cannot be inherited, and llama-bench serves no endpoint to guard directly. Recorded as a known gap rather than an implied pass.
- energy window: eng-0003
run-0020 — llama-bench@1
- 2026-08-08 · config → cfg-0010
- prefill_tps=321.771492
- ⚠ no guard — Guarded at run-0005 for this model+build+backend at --parallel 1. llama-bench cells vary KV-quant (lossy) and flash-attn (binary), across which capability cannot be inherited, and llama-bench serves no endpoint to guard directly. Recorded as a known gap rather than an implied pass.
- energy window: eng-0004
run-0021 — llama-bench@1
- 2026-08-08 · config → cfg-0010
- decode_tps=21.63496
- ⚠ no guard — Guarded at run-0005 for this model+build+backend at --parallel 1. llama-bench cells vary KV-quant (lossy) and flash-attn (binary), across which capability cannot be inherited, and llama-bench serves no endpoint to guard directly. Recorded as a known gap rather than an implied pass.
- energy window: eng-0004
run-0022 — llama-bench@1
- 2026-08-08 · config → cfg-0010
- prefill_tps=302.605133
- ⚠ no guard — Guarded at run-0005 for this model+build+backend at --parallel 1. llama-bench cells vary KV-quant (lossy) and flash-attn (binary), across which capability cannot be inherited, and llama-bench serves no endpoint to guard directly. Recorded as a known gap rather than an implied pass.
- energy window: eng-0004
run-0023 — llama-bench@1
- 2026-08-08 · config → cfg-0010
- decode_tps=20.609235
- ⚠ no guard — Guarded at run-0005 for this model+build+backend at --parallel 1. llama-bench cells vary KV-quant (lossy) and flash-attn (binary), across which capability cannot be inherited, and llama-bench serves no endpoint to guard directly. Recorded as a known gap rather than an implied pass.
- energy window: eng-0004
run-0024 — llama-bench@1
- 2026-08-08 · config → cfg-0010
- prefill_tps=204.485278
- ⚠ no guard — Guarded at run-0005 for this model+build+backend at --parallel 1. llama-bench cells vary KV-quant (lossy) and flash-attn (binary), across which capability cannot be inherited, and llama-bench serves no endpoint to guard directly. Recorded as a known gap rather than an implied pass.
- energy window: eng-0004
run-0025 — llama-bench@1
- 2026-08-08 · config → cfg-0010
- decode_tps=15.256436
- ⚠ no guard — Guarded at run-0005 for this model+build+backend at --parallel 1. llama-bench cells vary KV-quant (lossy) and flash-attn (binary), across which capability cannot be inherited, and llama-bench serves no endpoint to guard directly. Recorded as a known gap rather than an implied pass.
- energy window: eng-0004
run-0026 — llama-bench@2
- 2026-08-08 · config → cfg-0011
- prefill_tps=322.369017
- ⚠ no guard — Patched build (3653e6d + ce7689f KV-dequant). Speculation/KV levers only; guard run-0005 covers the same model+backend at --parallel 1, but KV quant is a lossy lever so capability does not inherit. Recorded as a gap.
- energy window: eng-0010
run-0027 — llama-bench@2
- 2026-08-08 · config → cfg-0011
- decode_tps=21.83139
- ⚠ no guard — Patched build (3653e6d + ce7689f KV-dequant). Speculation/KV levers only; guard run-0005 covers the same model+backend at --parallel 1, but KV quant is a lossy lever so capability does not inherit. Recorded as a gap.
- energy window: eng-0010
run-0028 — llama-bench@2
- 2026-08-08 · config → cfg-0011
- prefill_tps=211.351215
- ⚠ no guard — Patched build (3653e6d + ce7689f KV-dequant). Speculation/KV levers only; guard run-0005 covers the same model+backend at --parallel 1, but KV quant is a lossy lever so capability does not inherit. Recorded as a gap.
- energy window: eng-0010
run-0029 — llama-bench@2
- 2026-08-08 · config → cfg-0011
- decode_tps=18.087085
- ⚠ no guard — Patched build (3653e6d + ce7689f KV-dequant). Speculation/KV levers only; guard run-0005 covers the same model+backend at --parallel 1, but KV quant is a lossy lever so capability does not inherit. Recorded as a gap.
- energy window: eng-0010
run-0030 — llama-bench@1
run-0031 — llama-bench@1
run-0032 — llama-bench@1
run-0033 — llama-bench@1
run-0034 — llama-bench@1
run-0035 — llama-bench@1
run-0036 — llama-bench@2
- 2026-08-08 · config → cfg-0013
- decode_tps=12.152979
- ⚠ no guard — Patched build at 131072 depth, decode-only (-p 0 -r 1). Guard run-0005 covers this model+backend at --parallel 1; KV quant is a lossy lever so capability does not inherit. Depth exceeds the guard's 8000-token probe — recorded as a gap.
- energy window: eng-0015
run-0037 — llama-bench@2
- 2026-08-08 · config → cfg-0014
- decode_tps=9.6866
- ⚠ no guard — Patched build at 204800 depth (production context), decode-only. Guard run-0005 covers this model+backend at --parallel 1; KV quant is a lossy lever so capability does not inherit, and depth far exceeds the guard's 8000-token probe. Recorded as a gap.
- energy window: eng-0018
run-0038 — llama-bench@1
run-0039 — llama-bench@1
run-0040 — llama-bench@1
run-0041 — llama-bench@1
run-0042 — llama-bench@1
run-0043 — llama-bench@1
run-0044 — llama-bench@1
- 2026-08-08 · config → cfg-0016
- prefill_tps=450.832446
- ⚠ no guard — GUARD FAILED on this model: retrieval at 8000 tokens returned a stale answer from a previous request. Coherence and tool-calling passed. Numbers recorded because throughput is unaffected by the anomaly, but they are NOT cleared by a guard — see clm-0025.
- energy window: eng-0022
run-0045 — llama-bench@1
- 2026-08-08 · config → cfg-0016
- decode_tps=55.454588
- ⚠ no guard — GUARD FAILED on this model: retrieval at 8000 tokens returned a stale answer from a previous request. Coherence and tool-calling passed. Numbers recorded because throughput is unaffected by the anomaly, but they are NOT cleared by a guard — see clm-0025.
- energy window: eng-0022
run-0046 — llama-bench@1
- 2026-08-08 · config → cfg-0016
- prefill_tps=421.546647
- ⚠ no guard — GUARD FAILED on this model: retrieval at 8000 tokens returned a stale answer from a previous request. Coherence and tool-calling passed. Numbers recorded because throughput is unaffected by the anomaly, but they are NOT cleared by a guard — see clm-0025.
- energy window: eng-0022
run-0047 — llama-bench@1
- 2026-08-08 · config → cfg-0016
- decode_tps=52.793591
- ⚠ no guard — GUARD FAILED on this model: retrieval at 8000 tokens returned a stale answer from a previous request. Coherence and tool-calling passed. Numbers recorded because throughput is unaffected by the anomaly, but they are NOT cleared by a guard — see clm-0025.
- energy window: eng-0022
run-0048 — llama-bench@1
- 2026-08-08 · config → cfg-0016
- prefill_tps=308.546401
- ⚠ no guard — GUARD FAILED on this model: retrieval at 8000 tokens returned a stale answer from a previous request. Coherence and tool-calling passed. Numbers recorded because throughput is unaffected by the anomaly, but they are NOT cleared by a guard — see clm-0025.
- energy window: eng-0022
run-0049 — llama-bench@1
- 2026-08-08 · config → cfg-0016
- decode_tps=41.698106
- ⚠ no guard — GUARD FAILED on this model: retrieval at 8000 tokens returned a stale answer from a previous request. Coherence and tool-calling passed. Numbers recorded because throughput is unaffected by the anomaly, but they are NOT cleared by a guard — see clm-0025.
- energy window: eng-0022
run-0050 — llama-bench@1
run-0051 — llama-bench@1
run-0052 — llama-bench@1
run-0053 — llama-bench@1
run-0054 — llama-bench@1
run-0055 — llama-bench@1
run-0056 — llama-bench@1
run-0057 — llama-bench@1
run-0058 — llama-bench@1
run-0059 — llama-bench@1
run-0060 — llama-bench@1
run-0061 — llama-bench@1
run-0062 — llama-bench@1
run-0063 — llama-bench@1
run-0064 — llama-bench@1
run-0065 — llama-bench@1
run-0066 — llama-bench@1
run-0067 — llama-bench@1
run-0068 — llama-bench@1
run-0069 — llama-bench@1
run-0070 — llama-bench@1
run-0071 — llama-bench@1
run-0072 — llama-bench@1
run-0073 — llama-bench@1
run-0074 — llama-bench@1
run-0075 — llama-bench@1
run-0076 — llama-bench@1
run-0077 — llama-bench@1
run-0078 — llama-bench@1
run-0079 — llama-bench@1
run-0080 — llama-bench@1
run-0081 — llama-bench@1
run-0082 — llama-bench@1
run-0083 — llama-bench@1
run-0084 — llama-bench@1
run-0085 — llama-bench@1
run-0086 — llama-bench@1
run-0087 — llama-bench@1
run-0088 — llama-bench@1
run-0089 — llama-bench@1
run-0090 — llama-bench@1
run-0091 — llama-bench@1
run-0092 — llama-bench@1
run-0093 — llama-bench@1
run-0094 — llama-bench@1
run-0095 — llama-bench@1
run-0096 — llama-bench@1
run-0097 — llama-bench@1
run-0112 — llama-bench@1
- 2026-08-12 · config → cfg-0032
- prefill_tps=1068.109863
- ⚠ no guard — No capability guard exists for Qwen3.6-35B on this backend/build/KV combination (run-0004's guard is a different suite and build; run-0005 guards the 122B, not this model). llama-bench measures throughput only and serves no endpoint for a guard to attach to. clm-0050/0051's methodology note records the closest available correctness check for these arms — 41/41 layers on GPU, CPU utilisation identical to stock, no stride bug — which is evidence against a gross correctness failure but is not a capability guard. Recorded as a gap.
run-0113 — llama-bench@1
- 2026-08-12 · config → cfg-0032
- decode_tps=51.056208
- ⚠ no guard — No capability guard exists for Qwen3.6-35B on this backend/build/KV combination (run-0004's guard is a different suite and build; run-0005 guards the 122B, not this model). llama-bench measures throughput only and serves no endpoint for a guard to attach to. clm-0050/0051's methodology note records the closest available correctness check for these arms — 41/41 layers on GPU, CPU utilisation identical to stock, no stride bug — which is evidence against a gross correctness failure but is not a capability guard. Recorded as a gap.
run-0114 — llama-bench@1
- 2026-08-12 · config → cfg-0032
- prefill_tps=253.623124
- ⚠ no guard — No capability guard exists for Qwen3.6-35B on this backend/build/KV combination (run-0004's guard is a different suite and build; run-0005 guards the 122B, not this model). llama-bench measures throughput only and serves no endpoint for a guard to attach to. clm-0050/0051's methodology note records the closest available correctness check for these arms — 41/41 layers on GPU, CPU utilisation identical to stock, no stride bug — which is evidence against a gross correctness failure but is not a capability guard. Recorded as a gap.
run-0115 — llama-bench@1
- 2026-08-12 · config → cfg-0032
- decode_tps=28.470427
- ⚠ no guard — No capability guard exists for Qwen3.6-35B on this backend/build/KV combination (run-0004's guard is a different suite and build; run-0005 guards the 122B, not this model). llama-bench measures throughput only and serves no endpoint for a guard to attach to. clm-0050/0051's methodology note records the closest available correctness check for these arms — 41/41 layers on GPU, CPU utilisation identical to stock, no stride bug — which is evidence against a gross correctness failure but is not a capability guard. Recorded as a gap.
run-0116 — llama-bench@1
- 2026-08-12 · config → cfg-0032
- prefill_tps=595.763132
- ⚠ no guard — No capability guard exists for Qwen3.6-35B on this backend/build/KV combination (run-0004's guard is a different suite and build; run-0005 guards the 122B, not this model). llama-bench measures throughput only and serves no endpoint for a guard to attach to. clm-0050/0051's methodology note records the closest available correctness check for these arms — 41/41 layers on GPU, CPU utilisation identical to stock, no stride bug — which is evidence against a gross correctness failure but is not a capability guard. Recorded as a gap.
run-0117 — llama-bench@1
- 2026-08-12 · config → cfg-0032
- decode_tps=42.326925
- ⚠ no guard — No capability guard exists for Qwen3.6-35B on this backend/build/KV combination (run-0004's guard is a different suite and build; run-0005 guards the 122B, not this model). llama-bench measures throughput only and serves no endpoint for a guard to attach to. clm-0050/0051's methodology note records the closest available correctness check for these arms — 41/41 layers on GPU, CPU utilisation identical to stock, no stride bug — which is evidence against a gross correctness failure but is not a capability guard. Recorded as a gap.
run-0118 — llama-bench@1
- 2026-08-12 · config → cfg-0032
- prefill_tps=962.643635
- ⚠ no guard — No capability guard exists for Qwen3.6-35B on this backend/build/KV combination (run-0004's guard is a different suite and build; run-0005 guards the 122B, not this model). llama-bench measures throughput only and serves no endpoint for a guard to attach to. clm-0050/0051's methodology note records the closest available correctness check for these arms — 41/41 layers on GPU, CPU utilisation identical to stock, no stride bug — which is evidence against a gross correctness failure but is not a capability guard. Recorded as a gap.
run-0119 — llama-bench@1
- 2026-08-12 · config → cfg-0032
- decode_tps=49.699174
- ⚠ no guard — No capability guard exists for Qwen3.6-35B on this backend/build/KV combination (run-0004's guard is a different suite and build; run-0005 guards the 122B, not this model). llama-bench measures throughput only and serves no endpoint for a guard to attach to. clm-0050/0051's methodology note records the closest available correctness check for these arms — 41/41 layers on GPU, CPU utilisation identical to stock, no stride bug — which is evidence against a gross correctness failure but is not a capability guard. Recorded as a gap.
run-0120 — llama-bench@1
- 2026-08-12 · config → cfg-0032
- prefill_tps=412.491129
- ⚠ no guard — No capability guard exists for Qwen3.6-35B on this backend/build/KV combination (run-0004's guard is a different suite and build; run-0005 guards the 122B, not this model). llama-bench measures throughput only and serves no endpoint for a guard to attach to. clm-0050/0051's methodology note records the closest available correctness check for these arms — 41/41 layers on GPU, CPU utilisation identical to stock, no stride bug — which is evidence against a gross correctness failure but is not a capability guard. Recorded as a gap.
run-0121 — llama-bench@1
- 2026-08-12 · config → cfg-0032
- decode_tps=36.432477
- ⚠ no guard — No capability guard exists for Qwen3.6-35B on this backend/build/KV combination (run-0004's guard is a different suite and build; run-0005 guards the 122B, not this model). llama-bench measures throughput only and serves no endpoint for a guard to attach to. clm-0050/0051's methodology note records the closest available correctness check for these arms — 41/41 layers on GPU, CPU utilisation identical to stock, no stride bug — which is evidence against a gross correctness failure but is not a capability guard. Recorded as a gap.
run-0122 — llama-bench@1
- 2026-08-12 · config → cfg-0033
- prefill_tps=251.750998
- ⚠ no guard — No capability guard exists for Qwen3.6-35B on this backend/build/KV combination (run-0004's guard is a different suite and build; run-0005 guards the 122B, not this model). llama-bench measures throughput only and serves no endpoint for a guard to attach to. clm-0050/0051's methodology note records the closest available correctness check for these arms — 41/41 layers on GPU, CPU utilisation identical to stock, no stride bug — which is evidence against a gross correctness failure but is not a capability guard. Recorded as a gap.
run-0123 — llama-bench@1
- 2026-08-12 · config → cfg-0033
- decode_tps=18.496346
- ⚠ no guard — No capability guard exists for Qwen3.6-35B on this backend/build/KV combination (run-0004's guard is a different suite and build; run-0005 guards the 122B, not this model). llama-bench measures throughput only and serves no endpoint for a guard to attach to. clm-0050/0051's methodology note records the closest available correctness check for these arms — 41/41 layers on GPU, CPU utilisation identical to stock, no stride bug — which is evidence against a gross correctness failure but is not a capability guard. Recorded as a gap.
run-0124 — llama-bench@1
- 2026-08-12 · config → cfg-0033
- prefill_tps=595.657338
- ⚠ no guard — No capability guard exists for Qwen3.6-35B on this backend/build/KV combination (run-0004's guard is a different suite and build; run-0005 guards the 122B, not this model). llama-bench measures throughput only and serves no endpoint for a guard to attach to. clm-0050/0051's methodology note records the closest available correctness check for these arms — 41/41 layers on GPU, CPU utilisation identical to stock, no stride bug — which is evidence against a gross correctness failure but is not a capability guard. Recorded as a gap.
run-0125 — llama-bench@1
- 2026-08-12 · config → cfg-0033
- decode_tps=35.215174
- ⚠ no guard — No capability guard exists for Qwen3.6-35B on this backend/build/KV combination (run-0004's guard is a different suite and build; run-0005 guards the 122B, not this model). llama-bench measures throughput only and serves no endpoint for a guard to attach to. clm-0050/0051's methodology note records the closest available correctness check for these arms — 41/41 layers on GPU, CPU utilisation identical to stock, no stride bug — which is evidence against a gross correctness failure but is not a capability guard. Recorded as a gap.
run-0126 — llama-bench@1
- 2026-08-12 · config → cfg-0033
- prefill_tps=407.836908
- ⚠ no guard — No capability guard exists for Qwen3.6-35B on this backend/build/KV combination (run-0004's guard is a different suite and build; run-0005 guards the 122B, not this model). llama-bench measures throughput only and serves no endpoint for a guard to attach to. clm-0050/0051's methodology note records the closest available correctness check for these arms — 41/41 layers on GPU, CPU utilisation identical to stock, no stride bug — which is evidence against a gross correctness failure but is not a capability guard. Recorded as a gap.
run-0127 — llama-bench@1
- 2026-08-12 · config → cfg-0033
- decode_tps=27.310116
- ⚠ no guard — No capability guard exists for Qwen3.6-35B on this backend/build/KV combination (run-0004's guard is a different suite and build; run-0005 guards the 122B, not this model). llama-bench measures throughput only and serves no endpoint for a guard to attach to. clm-0050/0051's methodology note records the closest available correctness check for these arms — 41/41 layers on GPU, CPU utilisation identical to stock, no stride bug — which is evidence against a gross correctness failure but is not a capability guard. Recorded as a gap.
run-0128 — llama-bench@1
- 2026-08-12 · config → cfg-0034
- prefill_tps=1110.521105
- ⚠ no guard — No capability guard exists for Qwen3.6-35B on this backend/build/KV combination (run-0004's guard is a different suite and build; run-0005 guards the 122B, not this model). llama-bench measures throughput only and serves no endpoint for a guard to attach to. clm-0050/0051's methodology note records the closest available correctness check for these arms — 41/41 layers on GPU, CPU utilisation identical to stock, no stride bug — which is evidence against a gross correctness failure but is not a capability guard. Recorded as a gap.
run-0129 — llama-bench@1
- 2026-08-12 · config → cfg-0034
- decode_tps=61.620189
- ⚠ no guard — No capability guard exists for Qwen3.6-35B on this backend/build/KV combination (run-0004's guard is a different suite and build; run-0005 guards the 122B, not this model). llama-bench measures throughput only and serves no endpoint for a guard to attach to. clm-0050/0051's methodology note records the closest available correctness check for these arms — 41/41 layers on GPU, CPU utilisation identical to stock, no stride bug — which is evidence against a gross correctness failure but is not a capability guard. Recorded as a gap.
run-0130 — llama-bench@1
- 2026-08-12 · config → cfg-0034
- prefill_tps=278.271174
- ⚠ no guard — No capability guard exists for Qwen3.6-35B on this backend/build/KV combination (run-0004's guard is a different suite and build; run-0005 guards the 122B, not this model). llama-bench measures throughput only and serves no endpoint for a guard to attach to. clm-0050/0051's methodology note records the closest available correctness check for these arms — 41/41 layers on GPU, CPU utilisation identical to stock, no stride bug — which is evidence against a gross correctness failure but is not a capability guard. Recorded as a gap.
run-0131 — llama-bench@1
- 2026-08-12 · config → cfg-0034
- decode_tps=34.24807
- ⚠ no guard — No capability guard exists for Qwen3.6-35B on this backend/build/KV combination (run-0004's guard is a different suite and build; run-0005 guards the 122B, not this model). llama-bench measures throughput only and serves no endpoint for a guard to attach to. clm-0050/0051's methodology note records the closest available correctness check for these arms — 41/41 layers on GPU, CPU utilisation identical to stock, no stride bug — which is evidence against a gross correctness failure but is not a capability guard. Recorded as a gap.
run-0132 — llama-bench@1
- 2026-08-12 · config → cfg-0034
- prefill_tps=697.733641
- ⚠ no guard — No capability guard exists for Qwen3.6-35B on this backend/build/KV combination (run-0004's guard is a different suite and build; run-0005 guards the 122B, not this model). llama-bench measures throughput only and serves no endpoint for a guard to attach to. clm-0050/0051's methodology note records the closest available correctness check for these arms — 41/41 layers on GPU, CPU utilisation identical to stock, no stride bug — which is evidence against a gross correctness failure but is not a capability guard. Recorded as a gap.
run-0133 — llama-bench@1
- 2026-08-12 · config → cfg-0034
- decode_tps=50.372403
- ⚠ no guard — No capability guard exists for Qwen3.6-35B on this backend/build/KV combination (run-0004's guard is a different suite and build; run-0005 guards the 122B, not this model). llama-bench measures throughput only and serves no endpoint for a guard to attach to. clm-0050/0051's methodology note records the closest available correctness check for these arms — 41/41 layers on GPU, CPU utilisation identical to stock, no stride bug — which is evidence against a gross correctness failure but is not a capability guard. Recorded as a gap.
run-0134 — llama-bench@1
- 2026-08-12 · config → cfg-0034
- prefill_tps=1020.010129
- ⚠ no guard — No capability guard exists for Qwen3.6-35B on this backend/build/KV combination (run-0004's guard is a different suite and build; run-0005 guards the 122B, not this model). llama-bench measures throughput only and serves no endpoint for a guard to attach to. clm-0050/0051's methodology note records the closest available correctness check for these arms — 41/41 layers on GPU, CPU utilisation identical to stock, no stride bug — which is evidence against a gross correctness failure but is not a capability guard. Recorded as a gap.
run-0135 — llama-bench@1
- 2026-08-12 · config → cfg-0034
- decode_tps=59.339747
- ⚠ no guard — No capability guard exists for Qwen3.6-35B on this backend/build/KV combination (run-0004's guard is a different suite and build; run-0005 guards the 122B, not this model). llama-bench measures throughput only and serves no endpoint for a guard to attach to. clm-0050/0051's methodology note records the closest available correctness check for these arms — 41/41 layers on GPU, CPU utilisation identical to stock, no stride bug — which is evidence against a gross correctness failure but is not a capability guard. Recorded as a gap.
run-0136 — llama-bench@1
- 2026-08-12 · config → cfg-0034
- prefill_tps=495.077024
- ⚠ no guard — No capability guard exists for Qwen3.6-35B on this backend/build/KV combination (run-0004's guard is a different suite and build; run-0005 guards the 122B, not this model). llama-bench measures throughput only and serves no endpoint for a guard to attach to. clm-0050/0051's methodology note records the closest available correctness check for these arms — 41/41 layers on GPU, CPU utilisation identical to stock, no stride bug — which is evidence against a gross correctness failure but is not a capability guard. Recorded as a gap.
run-0137 — llama-bench@1
- 2026-08-12 · config → cfg-0034
- decode_tps=43.464041
- ⚠ no guard — No capability guard exists for Qwen3.6-35B on this backend/build/KV combination (run-0004's guard is a different suite and build; run-0005 guards the 122B, not this model). llama-bench measures throughput only and serves no endpoint for a guard to attach to. clm-0050/0051's methodology note records the closest available correctness check for these arms — 41/41 layers on GPU, CPU utilisation identical to stock, no stride bug — which is evidence against a gross correctness failure but is not a capability guard. Recorded as a gap.
run-0138 — llama-bench@1
- 2026-08-12 · config → cfg-0035
- prefill_tps=197.414071
- ⚠ no guard — No capability guard exists for Qwen3.6-35B on this backend/build/KV combination (run-0004's guard is a different suite and build; run-0005 guards the 122B, not this model). llama-bench measures throughput only and serves no endpoint for a guard to attach to. clm-0050/0051's methodology note records the closest available correctness check for these arms — 41/41 layers on GPU, CPU utilisation identical to stock, no stride bug — which is evidence against a gross correctness failure but is not a capability guard. Recorded as a gap.
run-0139 — llama-bench@1
- 2026-08-12 · config → cfg-0035
- decode_tps=40.153708
- ⚠ no guard — No capability guard exists for Qwen3.6-35B on this backend/build/KV combination (run-0004's guard is a different suite and build; run-0005 guards the 122B, not this model). llama-bench measures throughput only and serves no endpoint for a guard to attach to. clm-0050/0051's methodology note records the closest available correctness check for these arms — 41/41 layers on GPU, CPU utilisation identical to stock, no stride bug — which is evidence against a gross correctness failure but is not a capability guard. Recorded as a gap.
run-0140 — llama-bench@1
- 2026-08-12 · config → cfg-0035
- prefill_tps=517.767921
- ⚠ no guard — No capability guard exists for Qwen3.6-35B on this backend/build/KV combination (run-0004's guard is a different suite and build; run-0005 guards the 122B, not this model). llama-bench measures throughput only and serves no endpoint for a guard to attach to. clm-0050/0051's methodology note records the closest available correctness check for these arms — 41/41 layers on GPU, CPU utilisation identical to stock, no stride bug — which is evidence against a gross correctness failure but is not a capability guard. Recorded as a gap.
run-0141 — llama-bench@1
- 2026-08-12 · config → cfg-0035
- decode_tps=53.803678
- ⚠ no guard — No capability guard exists for Qwen3.6-35B on this backend/build/KV combination (run-0004's guard is a different suite and build; run-0005 guards the 122B, not this model). llama-bench measures throughput only and serves no endpoint for a guard to attach to. clm-0050/0051's methodology note records the closest available correctness check for these arms — 41/41 layers on GPU, CPU utilisation identical to stock, no stride bug — which is evidence against a gross correctness failure but is not a capability guard. Recorded as a gap.
run-0142 — llama-bench@1
- 2026-08-12 · config → cfg-0035
- prefill_tps=334.453951
- ⚠ no guard — No capability guard exists for Qwen3.6-35B on this backend/build/KV combination (run-0004's guard is a different suite and build; run-0005 guards the 122B, not this model). llama-bench measures throughput only and serves no endpoint for a guard to attach to. clm-0050/0051's methodology note records the closest available correctness check for these arms — 41/41 layers on GPU, CPU utilisation identical to stock, no stride bug — which is evidence against a gross correctness failure but is not a capability guard. Recorded as a gap.
run-0143 — llama-bench@1
- 2026-08-12 · config → cfg-0035
- decode_tps=48.055912
- ⚠ no guard — No capability guard exists for Qwen3.6-35B on this backend/build/KV combination (run-0004's guard is a different suite and build; run-0005 guards the 122B, not this model). llama-bench measures throughput only and serves no endpoint for a guard to attach to. clm-0050/0051's methodology note records the closest available correctness check for these arms — 41/41 layers on GPU, CPU utilisation identical to stock, no stride bug — which is evidence against a gross correctness failure but is not a capability guard. Recorded as a gap.
run-0144 — llama-bench@10350
- 2026-08-12 · config → cfg-0036
- prefill_tps=1251.48551
- ⚠ no guard — No capability guard exists for Qwen3.6-35B on this backend/build/KV combination (run-0004's guard is a different suite and build; run-0005 guards the 122B, not this model). llama-bench measures throughput only and serves no endpoint for a guard to attach to. clm-0050/0051's methodology note records the closest available correctness check for these arms — 41/41 layers on GPU, CPU utilisation identical to stock, no stride bug — which is evidence against a gross correctness failure but is not a capability guard. Recorded as a gap.
run-0145 — llama-bench@10350
- 2026-08-12 · config → cfg-0036
- decode_tps=61.83462
- ⚠ no guard — No capability guard exists for Qwen3.6-35B on this backend/build/KV combination (run-0004's guard is a different suite and build; run-0005 guards the 122B, not this model). llama-bench measures throughput only and serves no endpoint for a guard to attach to. clm-0050/0051's methodology note records the closest available correctness check for these arms — 41/41 layers on GPU, CPU utilisation identical to stock, no stride bug — which is evidence against a gross correctness failure but is not a capability guard. Recorded as a gap.
run-0146 — llama-bench@10350
- 2026-08-12 · config → cfg-0036
- prefill_tps=321.598303
- ⚠ no guard — No capability guard exists for Qwen3.6-35B on this backend/build/KV combination (run-0004's guard is a different suite and build; run-0005 guards the 122B, not this model). llama-bench measures throughput only and serves no endpoint for a guard to attach to. clm-0050/0051's methodology note records the closest available correctness check for these arms — 41/41 layers on GPU, CPU utilisation identical to stock, no stride bug — which is evidence against a gross correctness failure but is not a capability guard. Recorded as a gap.
run-0147 — llama-bench@10350
- 2026-08-12 · config → cfg-0036
- decode_tps=34.463715
- ⚠ no guard — No capability guard exists for Qwen3.6-35B on this backend/build/KV combination (run-0004's guard is a different suite and build; run-0005 guards the 122B, not this model). llama-bench measures throughput only and serves no endpoint for a guard to attach to. clm-0050/0051's methodology note records the closest available correctness check for these arms — 41/41 layers on GPU, CPU utilisation identical to stock, no stride bug — which is evidence against a gross correctness failure but is not a capability guard. Recorded as a gap.
run-0148 — llama-bench@10350
- 2026-08-12 · config → cfg-0036
- prefill_tps=705.35212
- ⚠ no guard — No capability guard exists for Qwen3.6-35B on this backend/build/KV combination (run-0004's guard is a different suite and build; run-0005 guards the 122B, not this model). llama-bench measures throughput only and serves no endpoint for a guard to attach to. clm-0050/0051's methodology note records the closest available correctness check for these arms — 41/41 layers on GPU, CPU utilisation identical to stock, no stride bug — which is evidence against a gross correctness failure but is not a capability guard. Recorded as a gap.
run-0149 — llama-bench@10350
- 2026-08-12 · config → cfg-0036
- decode_tps=51.020735
- ⚠ no guard — No capability guard exists for Qwen3.6-35B on this backend/build/KV combination (run-0004's guard is a different suite and build; run-0005 guards the 122B, not this model). llama-bench measures throughput only and serves no endpoint for a guard to attach to. clm-0050/0051's methodology note records the closest available correctness check for these arms — 41/41 layers on GPU, CPU utilisation identical to stock, no stride bug — which is evidence against a gross correctness failure but is not a capability guard. Recorded as a gap.
run-0150 — llama-bench@10350
- 2026-08-12 · config → cfg-0036
- prefill_tps=1149.512533
- ⚠ no guard — No capability guard exists for Qwen3.6-35B on this backend/build/KV combination (run-0004's guard is a different suite and build; run-0005 guards the 122B, not this model). llama-bench measures throughput only and serves no endpoint for a guard to attach to. clm-0050/0051's methodology note records the closest available correctness check for these arms — 41/41 layers on GPU, CPU utilisation identical to stock, no stride bug — which is evidence against a gross correctness failure but is not a capability guard. Recorded as a gap.
run-0151 — llama-bench@10350
- 2026-08-12 · config → cfg-0036
- decode_tps=59.873755
- ⚠ no guard — No capability guard exists for Qwen3.6-35B on this backend/build/KV combination (run-0004's guard is a different suite and build; run-0005 guards the 122B, not this model). llama-bench measures throughput only and serves no endpoint for a guard to attach to. clm-0050/0051's methodology note records the closest available correctness check for these arms — 41/41 layers on GPU, CPU utilisation identical to stock, no stride bug — which is evidence against a gross correctness failure but is not a capability guard. Recorded as a gap.
run-0152 — llama-bench@10350
- 2026-08-12 · config → cfg-0036
- prefill_tps=507.325549
- ⚠ no guard — No capability guard exists for Qwen3.6-35B on this backend/build/KV combination (run-0004's guard is a different suite and build; run-0005 guards the 122B, not this model). llama-bench measures throughput only and serves no endpoint for a guard to attach to. clm-0050/0051's methodology note records the closest available correctness check for these arms — 41/41 layers on GPU, CPU utilisation identical to stock, no stride bug — which is evidence against a gross correctness failure but is not a capability guard. Recorded as a gap.
run-0153 — llama-bench@10350
- 2026-08-12 · config → cfg-0036
- decode_tps=43.934256
- ⚠ no guard — No capability guard exists for Qwen3.6-35B on this backend/build/KV combination (run-0004's guard is a different suite and build; run-0005 guards the 122B, not this model). llama-bench measures throughput only and serves no endpoint for a guard to attach to. clm-0050/0051's methodology note records the closest available correctness check for these arms — 41/41 layers on GPU, CPU utilisation identical to stock, no stride bug — which is evidence against a gross correctness failure but is not a capability guard. Recorded as a gap.
run-0154 — llama-bench@10350
- 2026-08-12 · config → cfg-0037
- prefill_tps=330.899852
- ⚠ no guard — No capability guard exists for Qwen3.6-35B on this backend/build/KV combination (run-0004's guard is a different suite and build; run-0005 guards the 122B, not this model). llama-bench measures throughput only and serves no endpoint for a guard to attach to. clm-0050/0051's methodology note records the closest available correctness check for these arms — 41/41 layers on GPU, CPU utilisation identical to stock, no stride bug — which is evidence against a gross correctness failure but is not a capability guard. Recorded as a gap.
run-0155 — llama-bench@10350
- 2026-08-12 · config → cfg-0037
- decode_tps=42.507696
- ⚠ no guard — No capability guard exists for Qwen3.6-35B on this backend/build/KV combination (run-0004's guard is a different suite and build; run-0005 guards the 122B, not this model). llama-bench measures throughput only and serves no endpoint for a guard to attach to. clm-0050/0051's methodology note records the closest available correctness check for these arms — 41/41 layers on GPU, CPU utilisation identical to stock, no stride bug — which is evidence against a gross correctness failure but is not a capability guard. Recorded as a gap.
run-0156 — llama-bench@10350
- 2026-08-12 · config → cfg-0037
- prefill_tps=701.138783
- ⚠ no guard — No capability guard exists for Qwen3.6-35B on this backend/build/KV combination (run-0004's guard is a different suite and build; run-0005 guards the 122B, not this model). llama-bench measures throughput only and serves no endpoint for a guard to attach to. clm-0050/0051's methodology note records the closest available correctness check for these arms — 41/41 layers on GPU, CPU utilisation identical to stock, no stride bug — which is evidence against a gross correctness failure but is not a capability guard. Recorded as a gap.
run-0157 — llama-bench@10350
- 2026-08-12 · config → cfg-0037
- decode_tps=55.004573
- ⚠ no guard — No capability guard exists for Qwen3.6-35B on this backend/build/KV combination (run-0004's guard is a different suite and build; run-0005 guards the 122B, not this model). llama-bench measures throughput only and serves no endpoint for a guard to attach to. clm-0050/0051's methodology note records the closest available correctness check for these arms — 41/41 layers on GPU, CPU utilisation identical to stock, no stride bug — which is evidence against a gross correctness failure but is not a capability guard. Recorded as a gap.
run-0158 — llama-bench@10350
- 2026-08-12 · config → cfg-0037
- prefill_tps=495.436605
- ⚠ no guard — No capability guard exists for Qwen3.6-35B on this backend/build/KV combination (run-0004's guard is a different suite and build; run-0005 guards the 122B, not this model). llama-bench measures throughput only and serves no endpoint for a guard to attach to. clm-0050/0051's methodology note records the closest available correctness check for these arms — 41/41 layers on GPU, CPU utilisation identical to stock, no stride bug — which is evidence against a gross correctness failure but is not a capability guard. Recorded as a gap.
run-0159 — llama-bench@10350
- 2026-08-12 · config → cfg-0037
- decode_tps=50.116692
- ⚠ no guard — No capability guard exists for Qwen3.6-35B on this backend/build/KV combination (run-0004's guard is a different suite and build; run-0005 guards the 122B, not this model). llama-bench measures throughput only and serves no endpoint for a guard to attach to. clm-0050/0051's methodology note records the closest available correctness check for these arms — 41/41 layers on GPU, CPU utilisation identical to stock, no stride bug — which is evidence against a gross correctness failure but is not a capability guard. Recorded as a gap.
run-0160 — llama-bench@1
- 2026-08-13 · config → cfg-0038
- prefill_tps=109.772077
- ⚠ no guard — No capability guard exists for this model on this backend/build/KV/fa combination. llama-bench measures throughput only and serves no endpoint for a guard to attach to. Recorded as a gap.
run-0161 — llama-bench@1
- 2026-08-13 · config → cfg-0038
- decode_tps=14.187799
- ⚠ no guard — No capability guard exists for this model on this backend/build/KV/fa combination. llama-bench measures throughput only and serves no endpoint for a guard to attach to. Recorded as a gap.
run-0162 — llama-bench@1
- 2026-08-13 · config → cfg-0039
- prefill_tps=144.299074
- ⚠ no guard — No capability guard exists for this model on this backend/build/KV/fa combination. llama-bench measures throughput only and serves no endpoint for a guard to attach to. Recorded as a gap.
run-0163 — llama-bench@1
- 2026-08-13 · config → cfg-0039
- decode_tps=23.968738
- ⚠ no guard — No capability guard exists for this model on this backend/build/KV/fa combination. llama-bench measures throughput only and serves no endpoint for a guard to attach to. Recorded as a gap.
run-0164 — llama-bench@1
- 2026-08-13 · config → cfg-0039
- prefill_tps=308.218394
- ⚠ no guard — No capability guard exists for this model on this backend/build/KV/fa combination. llama-bench measures throughput only and serves no endpoint for a guard to attach to. Recorded as a gap.
run-0165 — llama-bench@1
- 2026-08-13 · config → cfg-0039
- decode_tps=38.992642
- ⚠ no guard — No capability guard exists for this model on this backend/build/KV/fa combination. llama-bench measures throughput only and serves no endpoint for a guard to attach to. Recorded as a gap.
run-0166 — llama-bench@1
- 2026-08-13 · config → cfg-0039
- prefill_tps=226.741803
- ⚠ no guard — No capability guard exists for this model on this backend/build/KV/fa combination. llama-bench measures throughput only and serves no endpoint for a guard to attach to. Recorded as a gap.
run-0167 — llama-bench@1
- 2026-08-13 · config → cfg-0039
- decode_tps=33.010008
- ⚠ no guard — No capability guard exists for this model on this backend/build/KV/fa combination. llama-bench measures throughput only and serves no endpoint for a guard to attach to. Recorded as a gap.
run-0168 — llama-bench@1
- 2026-08-13 · config → cfg-0040
- prefill_tps=300.276211
- ⚠ no guard — No capability guard exists for this model on this backend/build/KV/fa combination. llama-bench measures throughput only and serves no endpoint for a guard to attach to. Recorded as a gap.
run-0169 — llama-bench@1
- 2026-08-13 · config → cfg-0040
- decode_tps=34.812484
- ⚠ no guard — No capability guard exists for this model on this backend/build/KV/fa combination. llama-bench measures throughput only and serves no endpoint for a guard to attach to. Recorded as a gap.
run-0170 — llama-bench@1
- 2026-08-13 · config → cfg-0041
- prefill_tps=486.963899
- ⚠ no guard — No capability guard exists for this model on this backend/build/KV/fa combination. llama-bench measures throughput only and serves no endpoint for a guard to attach to. Recorded as a gap.
run-0171 — llama-bench@1
- 2026-08-13 · config → cfg-0041
- decode_tps=22.530918
- ⚠ no guard — No capability guard exists for this model on this backend/build/KV/fa combination. llama-bench measures throughput only and serves no endpoint for a guard to attach to. Recorded as a gap.
run-0172 — llama-bench@1
- 2026-08-13 · config → cfg-0042
- prefill_tps=254.278593
- ⚠ no guard — No capability guard exists for this model on this backend/build/KV/fa combination. llama-bench measures throughput only and serves no endpoint for a guard to attach to. Recorded as a gap.
run-0173 — llama-bench@1
- 2026-08-13 · config → cfg-0042
- decode_tps=28.437726
- ⚠ no guard — No capability guard exists for this model on this backend/build/KV/fa combination. llama-bench measures throughput only and serves no endpoint for a guard to attach to. Recorded as a gap.
run-0174 — llama-bench@1
- 2026-08-13 · config → cfg-0042
- prefill_tps=593.167952
- ⚠ no guard — No capability guard exists for this model on this backend/build/KV/fa combination. llama-bench measures throughput only and serves no endpoint for a guard to attach to. Recorded as a gap.
run-0175 — llama-bench@1
- 2026-08-13 · config → cfg-0042
- decode_tps=42.322086
- ⚠ no guard — No capability guard exists for this model on this backend/build/KV/fa combination. llama-bench measures throughput only and serves no endpoint for a guard to attach to. Recorded as a gap.
run-0176 — llama-bench@1
- 2026-08-13 · config → cfg-0042
- prefill_tps=410.408958
- ⚠ no guard — No capability guard exists for this model on this backend/build/KV/fa combination. llama-bench measures throughput only and serves no endpoint for a guard to attach to. Recorded as a gap.
run-0177 — llama-bench@1
- 2026-08-13 · config → cfg-0042
- decode_tps=36.459817
- ⚠ no guard — No capability guard exists for this model on this backend/build/KV/fa combination. llama-bench measures throughput only and serves no endpoint for a guard to attach to. Recorded as a gap.
run-0178 — llama-bench@1
- 2026-08-13 · config → cfg-0043
- prefill_tps=590.823403
- ⚠ no guard — No capability guard exists for this model on this backend/build/KV/fa combination. llama-bench measures throughput only and serves no endpoint for a guard to attach to. Recorded as a gap.
run-0179 — llama-bench@1
- 2026-08-13 · config → cfg-0043
- decode_tps=35.238117
- ⚠ no guard — No capability guard exists for this model on this backend/build/KV/fa combination. llama-bench measures throughput only and serves no endpoint for a guard to attach to. Recorded as a gap.
run-0180 — llama-bench@1
- 2026-08-13 · config → cfg-0044
- prefill_tps=23.106136
- ⚠ no guard — No capability guard exists for this model on this backend/build/KV/fa combination. llama-bench measures throughput only and serves no endpoint for a guard to attach to. Recorded as a gap.
run-0181 — llama-bench@1
- 2026-08-13 · config → cfg-0044
- decode_tps=26.32025
- ⚠ no guard — No capability guard exists for this model on this backend/build/KV/fa combination. llama-bench measures throughput only and serves no endpoint for a guard to attach to. Recorded as a gap.
run-0182 — llama-bench@1
- 2026-08-13 · config → cfg-0044
- prefill_tps=294.772527
- ⚠ no guard — No capability guard exists for this model on this backend/build/KV/fa combination. llama-bench measures throughput only and serves no endpoint for a guard to attach to. Recorded as a gap.
run-0183 — llama-bench@1
- 2026-08-13 · config → cfg-0044
- decode_tps=45.133705
- ⚠ no guard — No capability guard exists for this model on this backend/build/KV/fa combination. llama-bench measures throughput only and serves no endpoint for a guard to attach to. Recorded as a gap.
run-0184 — llama-bench@1
- 2026-08-13 · config → cfg-0044
- prefill_tps=108.781763
- ⚠ no guard — No capability guard exists for this model on this backend/build/KV/fa combination. llama-bench measures throughput only and serves no endpoint for a guard to attach to. Recorded as a gap.
run-0185 — llama-bench@1
- 2026-08-13 · config → cfg-0044
- decode_tps=36.554433
- ⚠ no guard — No capability guard exists for this model on this backend/build/KV/fa combination. llama-bench measures throughput only and serves no endpoint for a guard to attach to. Recorded as a gap.
run-0186 — llama-bench@1
- 2026-08-13 · config → cfg-0045
- prefill_tps=290.91762
- ⚠ no guard — No capability guard exists for this model on this backend/build/KV/fa combination. llama-bench measures throughput only and serves no endpoint for a guard to attach to. Recorded as a gap.
run-0187 — llama-bench@1
- 2026-08-13 · config → cfg-0045
- decode_tps=46.880608
- ⚠ no guard — No capability guard exists for this model on this backend/build/KV/fa combination. llama-bench measures throughput only and serves no endpoint for a guard to attach to. Recorded as a gap.
run-0188 — llama-bench@1
- 2026-08-13 · config → cfg-0046
- prefill_tps=116.588386
- ⚠ no guard — No capability guard exists for this model on this backend/build/KV/fa combination. llama-bench measures throughput only and serves no endpoint for a guard to attach to. Recorded as a gap.
run-0189 — llama-bench@1
- 2026-08-13 · config → cfg-0046
- decode_tps=16.442597
- ⚠ no guard — No capability guard exists for this model on this backend/build/KV/fa combination. llama-bench measures throughput only and serves no endpoint for a guard to attach to. Recorded as a gap.
run-0190 — llama-bench@1
- 2026-08-13 · config → cfg-0046
- prefill_tps=139.810384
- ⚠ no guard — No capability guard exists for this model on this backend/build/KV/fa combination. llama-bench measures throughput only and serves no endpoint for a guard to attach to. Recorded as a gap.
run-0191 — llama-bench@1
- 2026-08-13 · config → cfg-0046
- decode_tps=17.559139
- ⚠ no guard — No capability guard exists for this model on this backend/build/KV/fa combination. llama-bench measures throughput only and serves no endpoint for a guard to attach to. Recorded as a gap.
run-0192 — llama-bench@1
- 2026-08-13 · config → cfg-0046
- prefill_tps=134.429305
- ⚠ no guard — No capability guard exists for this model on this backend/build/KV/fa combination. llama-bench measures throughput only and serves no endpoint for a guard to attach to. Recorded as a gap.
run-0193 — llama-bench@1
- 2026-08-13 · config → cfg-0046
- decode_tps=17.146404
- ⚠ no guard — No capability guard exists for this model on this backend/build/KV/fa combination. llama-bench measures throughput only and serves no endpoint for a guard to attach to. Recorded as a gap.
run-0194 — llama-bench@1
- 2026-08-13 · config → cfg-0047
- prefill_tps=132.935563
- ⚠ no guard — No capability guard exists for this model on this backend/build/KV/fa combination. llama-bench measures throughput only and serves no endpoint for a guard to attach to. Recorded as a gap.
run-0195 — llama-bench@1
- 2026-08-13 · config → cfg-0047
- decode_tps=17.71399
- ⚠ no guard — No capability guard exists for this model on this backend/build/KV/fa combination. llama-bench measures throughput only and serves no endpoint for a guard to attach to. Recorded as a gap.
run-0196 — llama-bench@1
- 2026-08-13 · config → cfg-0048
- prefill_tps=277.396421
- ⚠ no guard — No capability guard exists for this model on this backend/build/KV/fa combination. llama-bench measures throughput only and serves no endpoint for a guard to attach to. Recorded as a gap.
run-0197 — llama-bench@1
- 2026-08-13 · config → cfg-0048
- decode_tps=34.188152
- ⚠ no guard — No capability guard exists for this model on this backend/build/KV/fa combination. llama-bench measures throughput only and serves no endpoint for a guard to attach to. Recorded as a gap.
run-0198 — llama-bench@1
- 2026-08-13 · config → cfg-0048
- prefill_tps=698.048959
- ⚠ no guard — No capability guard exists for this model on this backend/build/KV/fa combination. llama-bench measures throughput only and serves no endpoint for a guard to attach to. Recorded as a gap.
run-0199 — llama-bench@1
- 2026-08-13 · config → cfg-0048
- decode_tps=50.313108
- ⚠ no guard — No capability guard exists for this model on this backend/build/KV/fa combination. llama-bench measures throughput only and serves no endpoint for a guard to attach to. Recorded as a gap.
run-0200 — llama-bench@1
- 2026-08-13 · config → cfg-0048
- prefill_tps=489.314523
- ⚠ no guard — No capability guard exists for this model on this backend/build/KV/fa combination. llama-bench measures throughput only and serves no endpoint for a guard to attach to. Recorded as a gap.
run-0201 — llama-bench@1
- 2026-08-13 · config → cfg-0048
- decode_tps=43.450871
- ⚠ no guard — No capability guard exists for this model on this backend/build/KV/fa combination. llama-bench measures throughput only and serves no endpoint for a guard to attach to. Recorded as a gap.
run-0202 — llama-bench@1
- 2026-08-13 · config → cfg-0049
- prefill_tps=517.392007
- ⚠ no guard — No capability guard exists for this model on this backend/build/KV/fa combination. llama-bench measures throughput only and serves no endpoint for a guard to attach to. Recorded as a gap.
run-0203 — llama-bench@1
- 2026-08-13 · config → cfg-0049
- decode_tps=53.870394
- ⚠ no guard — No capability guard exists for this model on this backend/build/KV/fa combination. llama-bench measures throughput only and serves no endpoint for a guard to attach to. Recorded as a gap.
run-0204 — llama-bench@50
- 2026-08-14 · config → cfg-0050
- prefill_tps=651.431862
- ⚠ no guard — No capability guard exists for this build (ed89854 / a94d563 are new to the fleet — see bench/protocol.json known_builds, added 2026-08-14). The anchor-pair calibration (matched pp1024/tg256 depth cells vs stock 3653e6d ROCm and Vulkan run in the same session, per protocol.json) is a throughput cross-check, not a capability guard. llama-bench serves no endpoint for a guard to attach to. Recorded as a gap.
run-0205 — llama-bench@50
- 2026-08-14 · config → cfg-0050
- decode_tps=47.310018
- ⚠ no guard — No capability guard exists for this build (ed89854 / a94d563 are new to the fleet — see bench/protocol.json known_builds, added 2026-08-14). The anchor-pair calibration (matched pp1024/tg256 depth cells vs stock 3653e6d ROCm and Vulkan run in the same session, per protocol.json) is a throughput cross-check, not a capability guard. llama-bench serves no endpoint for a guard to attach to. Recorded as a gap.
run-0206 — llama-bench@50
- 2026-08-14 · config → cfg-0051
- prefill_tps=1075.40898
- ⚠ no guard — No capability guard exists for this build (ed89854 / a94d563 are new to the fleet — see bench/protocol.json known_builds, added 2026-08-14). The anchor-pair calibration (matched pp1024/tg256 depth cells vs stock 3653e6d ROCm and Vulkan run in the same session, per protocol.json) is a throughput cross-check, not a capability guard. llama-bench serves no endpoint for a guard to attach to. Recorded as a gap.
run-0207 — llama-bench@50
- 2026-08-14 · config → cfg-0051
- decode_tps=54.001268
- ⚠ no guard — No capability guard exists for this build (ed89854 / a94d563 are new to the fleet — see bench/protocol.json known_builds, added 2026-08-14). The anchor-pair calibration (matched pp1024/tg256 depth cells vs stock 3653e6d ROCm and Vulkan run in the same session, per protocol.json) is a throughput cross-check, not a capability guard. llama-bench serves no endpoint for a guard to attach to. Recorded as a gap.
run-0208 — llama-bench@50
- 2026-08-14 · config → cfg-0051
- prefill_tps=257.931381
- ⚠ no guard — No capability guard exists for this build (ed89854 / a94d563 are new to the fleet — see bench/protocol.json known_builds, added 2026-08-14). The anchor-pair calibration (matched pp1024/tg256 depth cells vs stock 3653e6d ROCm and Vulkan run in the same session, per protocol.json) is a throughput cross-check, not a capability guard. llama-bench serves no endpoint for a guard to attach to. Recorded as a gap.
run-0209 — llama-bench@50
- 2026-08-14 · config → cfg-0051
- decode_tps=29.844321
- ⚠ no guard — No capability guard exists for this build (ed89854 / a94d563 are new to the fleet — see bench/protocol.json known_builds, added 2026-08-14). The anchor-pair calibration (matched pp1024/tg256 depth cells vs stock 3653e6d ROCm and Vulkan run in the same session, per protocol.json) is a throughput cross-check, not a capability guard. llama-bench serves no endpoint for a guard to attach to. Recorded as a gap.
run-0210 — llama-bench@50
- 2026-08-14 · config → cfg-0051
- prefill_tps=607.281338
- ⚠ no guard — No capability guard exists for this build (ed89854 / a94d563 are new to the fleet — see bench/protocol.json known_builds, added 2026-08-14). The anchor-pair calibration (matched pp1024/tg256 depth cells vs stock 3653e6d ROCm and Vulkan run in the same session, per protocol.json) is a throughput cross-check, not a capability guard. llama-bench serves no endpoint for a guard to attach to. Recorded as a gap.
run-0211 — llama-bench@50
- 2026-08-14 · config → cfg-0051
- decode_tps=44.957416
- ⚠ no guard — No capability guard exists for this build (ed89854 / a94d563 are new to the fleet — see bench/protocol.json known_builds, added 2026-08-14). The anchor-pair calibration (matched pp1024/tg256 depth cells vs stock 3653e6d ROCm and Vulkan run in the same session, per protocol.json) is a throughput cross-check, not a capability guard. llama-bench serves no endpoint for a guard to attach to. Recorded as a gap.
run-0212 — llama-bench@1
- 2026-08-14 · config → cfg-0052
- prefill_tps=587.393257
- ⚠ no guard — No capability guard exists for this build (ed89854 / a94d563 are new to the fleet — see bench/protocol.json known_builds, added 2026-08-14). The anchor-pair calibration (matched pp1024/tg256 depth cells vs stock 3653e6d ROCm and Vulkan run in the same session, per protocol.json) is a throughput cross-check, not a capability guard. llama-bench serves no endpoint for a guard to attach to. Recorded as a gap.
run-0213 — llama-bench@1
- 2026-08-14 · config → cfg-0052
- decode_tps=29.255685
- ⚠ no guard — No capability guard exists for this build (ed89854 / a94d563 are new to the fleet — see bench/protocol.json known_builds, added 2026-08-14). The anchor-pair calibration (matched pp1024/tg256 depth cells vs stock 3653e6d ROCm and Vulkan run in the same session, per protocol.json) is a throughput cross-check, not a capability guard. llama-bench serves no endpoint for a guard to attach to. Recorded as a gap.
run-0214 — llama-bench@1
- 2026-08-14 · config → cfg-0053
- prefill_tps=1083.28017
- ⚠ no guard — No capability guard exists for this build (ed89854 / a94d563 are new to the fleet — see bench/protocol.json known_builds, added 2026-08-14). The anchor-pair calibration (matched pp1024/tg256 depth cells vs stock 3653e6d ROCm and Vulkan run in the same session, per protocol.json) is a throughput cross-check, not a capability guard. llama-bench serves no endpoint for a guard to attach to. Recorded as a gap.
run-0215 — llama-bench@1
- 2026-08-14 · config → cfg-0053
- decode_tps=51.073067
- ⚠ no guard — No capability guard exists for this build (ed89854 / a94d563 are new to the fleet — see bench/protocol.json known_builds, added 2026-08-14). The anchor-pair calibration (matched pp1024/tg256 depth cells vs stock 3653e6d ROCm and Vulkan run in the same session, per protocol.json) is a throughput cross-check, not a capability guard. llama-bench serves no endpoint for a guard to attach to. Recorded as a gap.
run-0216 — llama-bench@1
- 2026-08-14 · config → cfg-0053
- prefill_tps=253.680696
- ⚠ no guard — No capability guard exists for this build (ed89854 / a94d563 are new to the fleet — see bench/protocol.json known_builds, added 2026-08-14). The anchor-pair calibration (matched pp1024/tg256 depth cells vs stock 3653e6d ROCm and Vulkan run in the same session, per protocol.json) is a throughput cross-check, not a capability guard. llama-bench serves no endpoint for a guard to attach to. Recorded as a gap.
run-0217 — llama-bench@1
- 2026-08-14 · config → cfg-0053
- decode_tps=28.468117
- ⚠ no guard — No capability guard exists for this build (ed89854 / a94d563 are new to the fleet — see bench/protocol.json known_builds, added 2026-08-14). The anchor-pair calibration (matched pp1024/tg256 depth cells vs stock 3653e6d ROCm and Vulkan run in the same session, per protocol.json) is a throughput cross-check, not a capability guard. llama-bench serves no endpoint for a guard to attach to. Recorded as a gap.
run-0218 — llama-bench@1
- 2026-08-14 · config → cfg-0053
- prefill_tps=595.008415
- ⚠ no guard — No capability guard exists for this build (ed89854 / a94d563 are new to the fleet — see bench/protocol.json known_builds, added 2026-08-14). The anchor-pair calibration (matched pp1024/tg256 depth cells vs stock 3653e6d ROCm and Vulkan run in the same session, per protocol.json) is a throughput cross-check, not a capability guard. llama-bench serves no endpoint for a guard to attach to. Recorded as a gap.
run-0219 — llama-bench@1
- 2026-08-14 · config → cfg-0053
- decode_tps=42.346651
- ⚠ no guard — No capability guard exists for this build (ed89854 / a94d563 are new to the fleet — see bench/protocol.json known_builds, added 2026-08-14). The anchor-pair calibration (matched pp1024/tg256 depth cells vs stock 3653e6d ROCm and Vulkan run in the same session, per protocol.json) is a throughput cross-check, not a capability guard. llama-bench serves no endpoint for a guard to attach to. Recorded as a gap.
run-0220 — llama-bench@27
- 2026-08-14 · config → cfg-0054
- prefill_tps=588.345904
- ⚠ no guard — No capability guard exists for this build (ed89854 / a94d563 are new to the fleet — see bench/protocol.json known_builds, added 2026-08-14). The anchor-pair calibration (matched pp1024/tg256 depth cells vs stock 3653e6d ROCm and Vulkan run in the same session, per protocol.json) is a throughput cross-check, not a capability guard. llama-bench serves no endpoint for a guard to attach to. Recorded as a gap.
run-0221 — llama-bench@27
- 2026-08-14 · config → cfg-0054
- decode_tps=29.123834
- ⚠ no guard — No capability guard exists for this build (ed89854 / a94d563 are new to the fleet — see bench/protocol.json known_builds, added 2026-08-14). The anchor-pair calibration (matched pp1024/tg256 depth cells vs stock 3653e6d ROCm and Vulkan run in the same session, per protocol.json) is a throughput cross-check, not a capability guard. llama-bench serves no endpoint for a guard to attach to. Recorded as a gap.
run-0222 — llama-bench@27
- 2026-08-14 · config → cfg-0055
- prefill_tps=1063.898925
- ⚠ no guard — No capability guard exists for this build (ed89854 / a94d563 are new to the fleet — see bench/protocol.json known_builds, added 2026-08-14). The anchor-pair calibration (matched pp1024/tg256 depth cells vs stock 3653e6d ROCm and Vulkan run in the same session, per protocol.json) is a throughput cross-check, not a capability guard. llama-bench serves no endpoint for a guard to attach to. Recorded as a gap.
run-0223 — llama-bench@27
- 2026-08-14 · config → cfg-0055
- decode_tps=50.733048
- ⚠ no guard — No capability guard exists for this build (ed89854 / a94d563 are new to the fleet — see bench/protocol.json known_builds, added 2026-08-14). The anchor-pair calibration (matched pp1024/tg256 depth cells vs stock 3653e6d ROCm and Vulkan run in the same session, per protocol.json) is a throughput cross-check, not a capability guard. llama-bench serves no endpoint for a guard to attach to. Recorded as a gap.
run-0224 — llama-bench@27
- 2026-08-14 · config → cfg-0055
- prefill_tps=254.950374
- ⚠ no guard — No capability guard exists for this build (ed89854 / a94d563 are new to the fleet — see bench/protocol.json known_builds, added 2026-08-14). The anchor-pair calibration (matched pp1024/tg256 depth cells vs stock 3653e6d ROCm and Vulkan run in the same session, per protocol.json) is a throughput cross-check, not a capability guard. llama-bench serves no endpoint for a guard to attach to. Recorded as a gap.
run-0225 — llama-bench@27
- 2026-08-14 · config → cfg-0055
- decode_tps=28.333051
- ⚠ no guard — No capability guard exists for this build (ed89854 / a94d563 are new to the fleet — see bench/protocol.json known_builds, added 2026-08-14). The anchor-pair calibration (matched pp1024/tg256 depth cells vs stock 3653e6d ROCm and Vulkan run in the same session, per protocol.json) is a throughput cross-check, not a capability guard. llama-bench serves no endpoint for a guard to attach to. Recorded as a gap.
run-0226 — llama-bench@27
- 2026-08-14 · config → cfg-0055
- prefill_tps=593.669237
- ⚠ no guard — No capability guard exists for this build (ed89854 / a94d563 are new to the fleet — see bench/protocol.json known_builds, added 2026-08-14). The anchor-pair calibration (matched pp1024/tg256 depth cells vs stock 3653e6d ROCm and Vulkan run in the same session, per protocol.json) is a throughput cross-check, not a capability guard. llama-bench serves no endpoint for a guard to attach to. Recorded as a gap.
run-0227 — llama-bench@27
- 2026-08-14 · config → cfg-0055
- decode_tps=42.1234
- ⚠ no guard — No capability guard exists for this build (ed89854 / a94d563 are new to the fleet — see bench/protocol.json known_builds, added 2026-08-14). The anchor-pair calibration (matched pp1024/tg256 depth cells vs stock 3653e6d ROCm and Vulkan run in the same session, per protocol.json) is a throughput cross-check, not a capability guard. llama-bench serves no endpoint for a guard to attach to. Recorded as a gap.
run-0228 — llama-bench@1
- 2026-08-14 · config → cfg-0056
- prefill_tps=1114.742346
- ⚠ no guard — No capability guard exists for this build. The anchor-pair calibration (matched pp1024/tg256 depth cells vs stock 3653e6d ROCm and Vulkan run in the same session, per protocol.json) is a throughput cross-check, not a capability guard. llama-bench serves no endpoint for a guard to attach to. Recorded as a gap.
run-0229 — llama-bench@1
- 2026-08-14 · config → cfg-0056
- decode_tps=61.650351
- ⚠ no guard — No capability guard exists for this build. The anchor-pair calibration (matched pp1024/tg256 depth cells vs stock 3653e6d ROCm and Vulkan run in the same session, per protocol.json) is a throughput cross-check, not a capability guard. llama-bench serves no endpoint for a guard to attach to. Recorded as a gap.
run-0230 — llama-bench@1
- 2026-08-14 · config → cfg-0056
- prefill_tps=267.616792
- ⚠ no guard — No capability guard exists for this build. The anchor-pair calibration (matched pp1024/tg256 depth cells vs stock 3653e6d ROCm and Vulkan run in the same session, per protocol.json) is a throughput cross-check, not a capability guard. llama-bench serves no endpoint for a guard to attach to. Recorded as a gap.
run-0231 — llama-bench@1
- 2026-08-14 · config → cfg-0056
- decode_tps=34.186336
- ⚠ no guard — No capability guard exists for this build. The anchor-pair calibration (matched pp1024/tg256 depth cells vs stock 3653e6d ROCm and Vulkan run in the same session, per protocol.json) is a throughput cross-check, not a capability guard. llama-bench serves no endpoint for a guard to attach to. Recorded as a gap.
run-0232 — llama-bench@1
- 2026-08-14 · config → cfg-0056
- prefill_tps=699.788866
- ⚠ no guard — No capability guard exists for this build. The anchor-pair calibration (matched pp1024/tg256 depth cells vs stock 3653e6d ROCm and Vulkan run in the same session, per protocol.json) is a throughput cross-check, not a capability guard. llama-bench serves no endpoint for a guard to attach to. Recorded as a gap.
run-0233 — llama-bench@1
- 2026-08-14 · config → cfg-0056
- decode_tps=50.363938
- ⚠ no guard — No capability guard exists for this build. The anchor-pair calibration (matched pp1024/tg256 depth cells vs stock 3653e6d ROCm and Vulkan run in the same session, per protocol.json) is a throughput cross-check, not a capability guard. llama-bench serves no endpoint for a guard to attach to. Recorded as a gap.
run-0234 — perf-matrix-gtt120@1
- 2026-08-14 · config → cfg-0057
- oom=true
- ⚠ no guard — No throughput was produced to guard — the process never reached a token. Recorded as a gap in the same sense as every other ungated performance row: llama-bench serves no endpoint for a guard to attach to, and here there is nothing behind it at all.
run-0235 — perf-matrix-gtt120@1
- 2026-08-14 · config → cfg-0058
- oom=true
- ⚠ no guard — No throughput was produced to guard — the process never reached a token. Recorded as a gap in the same sense as every other ungated performance row: llama-bench serves no endpoint for a guard to attach to, and here there is nothing behind it at all.
run-0236 — llama-bench@1
- 2026-08-14 · config → cfg-0059
- prefill_tps=244.051835
- ⚠ no guard — No capability guard exists for this model on this backend/build/KV/load-mode combination. llama-bench measures throughput only and serves no endpoint for a guard to attach to. Recorded as a gap.
run-0237 — llama-bench@1
- 2026-08-14 · config → cfg-0059
- decode_tps=16.434903
- ⚠ no guard — No capability guard exists for this model on this backend/build/KV/load-mode combination. llama-bench measures throughput only and serves no endpoint for a guard to attach to. Recorded as a gap.
run-0238 — llama-bench@1
- 2026-08-14 · config → cfg-0060
- prefill_tps=601.555743
- ⚠ no guard — No capability guard exists for Qwen3.6-35B on this backend/build/KV/load-mode combination. llama-bench measures throughput only and serves no endpoint for a guard to attach to. Recorded as a gap.
run-0239 — llama-bench@1
- 2026-08-14 · config → cfg-0060
- decode_tps=42.368806
- ⚠ no guard — No capability guard exists for Qwen3.6-35B on this backend/build/KV/load-mode combination. llama-bench measures throughput only and serves no endpoint for a guard to attach to. Recorded as a gap.
run-0240 — llama-bench@1
- 2026-08-14 · config → cfg-0032
- prefill_tps=597.271098
- ⚠ no guard — No capability guard exists for Qwen3.6-35B on this backend/build/KV combination (run-0004's guard is a different suite and build; run-0005 guards the 122B, not this model). llama-bench measures throughput only and serves no endpoint for a guard to attach to. Recorded as a gap.
run-0241 — llama-bench@1
- 2026-08-14 · config → cfg-0032
- decode_tps=42.41958
- ⚠ no guard — No capability guard exists for Qwen3.6-35B on this backend/build/KV combination (run-0004's guard is a different suite and build; run-0005 guards the 122B, not this model). llama-bench measures throughput only and serves no endpoint for a guard to attach to. Recorded as a gap.
run-0242 — memgate-test@1
- 2026-08-14 · config → cfg-0061
- oom=true
- ⚠ no guard — No throughput was produced to guard — the process never reached a token; it was killed by the queue's own bound, not by the model. Recorded as a gap in the same sense as run-0234/run-0235.
run-0243 — llama-bench@1
- 2026-08-14 · config → cfg-0059
- prefill_tps=248.788416
- ⚠ no guard — No capability guard exists for this model on this backend/build/KV/load-mode combination. llama-bench measures throughput only and serves no endpoint for a guard to attach to. Recorded as a gap.
run-0244 — llama-bench@1
- 2026-08-14 · config → cfg-0059
- decode_tps=16.390674
- ⚠ no guard — No capability guard exists for this model on this backend/build/KV/load-mode combination. llama-bench measures throughput only and serves no endpoint for a guard to attach to. Recorded as a gap.
run-0245 — llama-bench@1
run-0246 — llama-bench@1
run-0247 — llama-bench@1
run-0248 — llama-bench@1
run-0249 — llama-bench@1
run-0250 — llama-bench@1
run-0251 — llama-bench@1
run-0252 — llama-bench@1
run-0253 — llama-bench@1
run-0254 — llama-bench@1
run-0255 — llama-bench@1
run-0256 — llama-bench@1
run-0257 — llama-bench@1
run-0258 — llama-bench@1
run-0259 — llama-bench@1
run-0260 — llama-bench@1
run-0261 — llama-bench@1
run-0262 — llama-bench@1
run-0263 — llama-bench@1
run-0264 — llama-bench@1
run-0265 — llama-bench@1
- 2026-08-14 · config → cfg-0064
- device_lost=true
- ⚠ no guard — No throughput was produced to guard — both reps lost the GPU context before llama-bench emitted a row. Recorded as a gap in the same sense as every other ungated performance row: llama-bench serves no endpoint for a guard to attach to, and here there is nothing behind it at all.
- energy window: eng-0075
run-0266 — llama-bench@1
- 2026-08-14 · config → cfg-0065
- device_lost=true
- ⚠ no guard — No throughput was produced to guard — both reps lost the GPU context before llama-bench emitted a row. Recorded as a gap in the same sense as every other ungated performance row: llama-bench serves no endpoint for a guard to attach to, and here there is nothing behind it at all.
- energy window: eng-0076
run-0268 — mtp-probe@1
- 2026-08-15 · config → cfg-0067
- decode_tps=18.25 · decode_tps_no_speculation=7.82
- derived speculation_ratio: 2.33x — the ratio is the transferable number; the absolute is not comparable across configs
- ⚠ no guard — No capability guard exists for this model on this backend/build. The probe carries its own per-generation degeneracy check instead: every generation in both arms completed non-degenerate (0 of 15 in each arm, min-ok floor 250 of 500 tokens) — the specific hazard that invalidates n_max>=4 cells (clm-0055) is measured absent from BOTH numbers in this record.
- energy window: eng-0078
run-0269 — mtp-probe@1
- 2026-08-14 · config → cfg-0068
- decode_tps=17.75 · decode_tps_no_speculation=7.86
- derived speculation_ratio: 2.26x — the ratio is the transferable number; the absolute is not comparable across configs
- ⚠ no guard — No capability guard exists for this model on this backend/build. The probe carries its own per-generation degeneracy check instead: every generation in both arms completed non-degenerate (0 of 15 in each arm, min-ok floor 250 of 500 tokens) — the specific hazard that invalidates n_max>=4 cells (clm-0055) is measured absent from BOTH numbers in this record.
- energy window: eng-0079
run-0270 — mtp-probe@1
- 2026-08-14 · config → cfg-0069
- decode_tps=26.68 · decode_tps_no_speculation=11.98
- derived speculation_ratio: 2.23x — the ratio is the transferable number; the absolute is not comparable across configs
- ⚠ no guard — No capability guard exists for this model on this backend/build. The probe carries its own per-generation degeneracy check instead: every generation in both arms completed non-degenerate (0 of 15 in each arm, min-ok floor 250 of 500 tokens) — the specific hazard that invalidates n_max>=4 cells (clm-0055) is measured absent from BOTH numbers in this record.
- energy window: eng-0080
run-0288 — llama-bench@1
run-0289 — llama-bench@1
run-0290 — llama-bench@1
run-0291 — llama-bench@1
run-0292 — llama-bench@1
run-0293 — llama-bench@1
run-0294 — llama-bench@1
run-0295 — llama-bench@1
run-0296 — llama-bench@1
run-0297 — llama-bench@1
run-0298 — llama-bench@1
run-0299 — llama-bench@1
run-0300 — llama-bench@1
run-0301 — llama-bench@1
run-0302 — llama-bench@1
run-0303 — llama-bench@1
run-0308 — llama-bench@1
run-0309 — llama-bench@1
run-0310 — llama-bench@1
run-0311 — llama-bench@1
run-0312 — llama-bench@1
run-0313 — llama-bench@1
run-0314 — llama-bench@1
run-0315 — llama-bench@1
run-0317 — llama-bench@1
run-0318 — llama-bench@1
run-0319 — llama-bench@1
run-0320 — llama-bench@1
run-0321 — llama-bench@1
run-0322 — llama-bench@1
run-0323 — llama-bench@1
run-0324 — llama-bench@1
run-0325 — llama-bench@1
run-0326 — llama-bench@1
run-0327 — llama-bench@1
- 2026-08-16 · config → cfg-0091
- device_lost=true
- ⚠ no guard — No capability guard exists for this model on this backend/build/KV/load-mode combination. llama-bench measures throughput only and serves no endpoint for a guard to attach to. Recorded as a gap. This run also never produced a throughput number to guard, per the device_lost:true metric below.
- energy window: eng-0122
run-0332 — llama-bench@1
run-0333 — llama-bench@1
run-0334 — llama-bench@1
run-0335 — llama-bench@1
run-0336 — llama-bench@1
run-0337 — llama-bench@1
run-0339 — llama-bench@1
run-0340 — llama-bench@1
run-0341 — llama-bench@1
run-0342 — llama-bench@1
run-0343 — llama-bench@1
run-0344 — llama-bench@1
run-0346 — llama-bench@1
- 2026-08-16 · config → cfg-0097
- prefill_tps=357.8357
- ⚠ no guard — Matrix sanity gate (protocol §0 output-sanity precondition), not the full 4-item house guard: one real chat-endpoint generation on this exact ROCm binary, eyeballed for coherence before any throughput was recorded ("the quick brown fox" echoed correctly — sanity-rocm-server.log, aihydra ~/bench-results/qwen36-27b-mtp-fullbench/ matrix/). The full house guard (guard.sh, 4/4) is wired separately against this candidate's tau2 serving config — see run-0357/run-0359.
- energy window: eng-0135
run-0347 — llama-bench@1
run-0348 — llama-bench@1
run-0349 — llama-bench@1
run-0350 — llama-bench@1
run-0351 — llama-bench@1
run-0352 — llama-bench@1
- 2026-08-16 · config → cfg-0098
- prefill_tps=302.724
- ⚠ no guard — Matrix sanity gate (protocol §0 output-sanity precondition): one real chat-endpoint generation on this exact Vulkan binary, eyeballed for coherence before any throughput was recorded ("the quick brown fox" echoed correctly — sanity-vulkan-server.log, aihydra ~/bench-results/qwen36-27b-mtp-fullbench/matrix/). The full house guard (guard.sh, 4/4) is wired separately against this candidate's tau2 serving config (ROCm — see run-0357/run-0359), since Vulkan is not the config this bench serves.
- energy window: eng-0138
run-0353 — llama-bench@1
run-0354 — llama-bench@1
run-0355 — llama-bench@1
run-0356 — llama-bench@1
run-0364 — llama-bench@1
run-0365 — llama-bench@1
run-0366 — llama-bench@1
run-0367 — llama-bench@1
run-0368 — llama-bench@1
- 2026-08-17 · config → cfg-0102
- prefill_tps=12914.102
- ⚠ no guard — No capability guard was run against this candidate on the Vulkan backend — the house guard and the full 26-task tau2 run both served on ROCm (cfg-0101, run-0363), and a backend swap is a `binary` lever that is never inheritable (protocol §1a). This throughput number stands alone as a measured backend comparison, not as a capability-adjacent result.
- energy window: eng-0145
run-0369 — llama-bench@1
run-0370 — llama-bench@1
run-0371 — llama-bench@1
run-0373 — llama-bench@1
run-0374 — llama-bench@1
run-0375 — llama-bench@1
run-0376 — llama-bench@1
run-0377 — llama-bench@1
run-0378 — llama-bench@1
run-0379 — llama-bench@1
run-0380 — llama-bench@1
run-0381 — llama-bench@1
run-0382 — llama-bench@1
run-0383 — llama-bench@1
run-0384 — llama-bench@1
run-0385 — llama-bench@1
run-0386 — llama-bench@1
run-0387 — llama-bench@1
run-0388 — llama-bench@1
run-0389 — llama-bench@1
run-0390 — llama-bench@1
run-0391 — llama-bench@1
run-0392 — llama-bench@1
run-0396 — llama-bench@1
run-0397 — llama-bench@1
run-0398 — llama-bench@1
run-0399 — llama-bench@1
run-0400 — llama-bench@1
run-0401 — llama-bench@1
run-0402 — llama-bench@1
run-0403 — llama-bench@1
run-0404 — llama-bench@1
run-0405 — llama-bench@1
run-0406 — llama-bench@1
run-0407 — llama-bench@1
run-0408 — llama-bench@1
run-0409 — llama-bench@1
run-0410 — llama-bench@1
run-0411 — llama-bench@1
run-0412 — llama-bench@1
run-0413 — llama-bench@1
run-0414 — llama-bench@1
run-0415 — llama-bench@1
run-0416 — llama-bench@1
run-0417 — llama-bench@1
run-0418 — llama-bench@1
run-0419 — llama-bench@1
run-0420 — llama-bench@1
run-0421 — llama-bench@1
run-0422 — llama-bench@1
run-0423 — llama-bench@1
run-0424 — llama-bench@1
run-0425 — llama-bench@1
run-0426 — llama-bench@1
run-0427 — llama-bench@1
run-0428 — llama-bench@1
run-0429 — llama-bench@1
run-0430 — llama-bench@1
run-0431 — llama-bench@1
run-0435 — llama-bench@50
run-0436 — llama-bench@50
run-0437 — llama-bench@50
run-0438 — llama-bench@50
run-0441 — llama-bench@1
run-0442 — llama-bench@1
run-0443 — llama-bench@1
run-0444 — llama-bench@1
run-0445 — llama-bench@1
run-0446 — llama-bench@1
run-0447 — llama-bench@1
run-0448 — llama-bench@1
run-0449 — llama-bench@1
run-0450 — llama-bench@1
run-0453 — mtp-engagement-probe@1
run-0455 — mtp-engagement-probe@1
run-0457 — mtp-engagement-probe@1
run-0459 — mtp-engagement-probe@1
run-0461 — mtp-engagement-probe@2
run-0463 — mtp-engagement-probe@2
run-0465 — mtp-engagement-probe@2
run-0467 — mtp-engagement-probe@2
run-0474 — pr25494-vulkan-q8kv-stock-r7
- 2026-08-19 · config → cfg-0127
- decode_tps=50.471576 · prefill_tps=703.069152
- ⚠ no guard — Throughput-only current-stock Vulkan rerun admitted by hbreviewer t_ce7d961e; no capability guard was performed for this r7 boundary, so these performance numbers are record-backed throughput evidence only and must not imply a capability pass.
run-0475 — pr25494-vulkan-q8kv-stock-r7
- 2026-08-19 · config → cfg-0127
- decode_tps=50.370761 · prefill_tps=698.246577
- ⚠ no guard — Throughput-only current-stock Vulkan rerun admitted by hbreviewer t_ce7d961e; no capability guard was performed for this r7 boundary, so these performance numbers are record-backed throughput evidence only and must not imply a capability pass.
run-0476 — pr25494-vulkan-q8kv-stock-r7
- 2026-08-19 · config → cfg-0127
- decode_tps=50.490031 · prefill_tps=695.1861
- ⚠ no guard — Throughput-only current-stock Vulkan rerun admitted by hbreviewer t_ce7d961e; no capability guard was performed for this r7 boundary, so these performance numbers are record-backed throughput evidence only and must not imply a capability pass.
run-0477 — pr25494-vulkan-q8kv-stock-r7
- 2026-08-19 · config → cfg-0127
- decode_tps=43.460468 · prefill_tps=503.738149
- ⚠ no guard — Throughput-only current-stock Vulkan rerun admitted by hbreviewer t_ce7d961e; no capability guard was performed for this r7 boundary, so these performance numbers are record-backed throughput evidence only and must not imply a capability pass.
run-0478 — pr25494-vulkan-q8kv-stock-r7
- 2026-08-19 · config → cfg-0127
- decode_tps=43.505277 · prefill_tps=498.802738
- ⚠ no guard — Throughput-only current-stock Vulkan rerun admitted by hbreviewer t_ce7d961e; no capability guard was performed for this r7 boundary, so these performance numbers are record-backed throughput evidence only and must not imply a capability pass.
run-0479 — pr25494-vulkan-q8kv-stock-r7
- 2026-08-19 · config → cfg-0127
- decode_tps=43.465855 · prefill_tps=502.035791
- ⚠ no guard — Throughput-only current-stock Vulkan rerun admitted by hbreviewer t_ce7d961e; no capability guard was performed for this r7 boundary, so these performance numbers are record-backed throughput evidence only and must not imply a capability pass.
run-0480 — pr25494-vulkan-q8kv-stock-r7
- 2026-08-19 · config → cfg-0127
- decode_tps=34.255287 · prefill_tps=286.027664
- ⚠ no guard — Throughput-only current-stock Vulkan rerun admitted by hbreviewer t_ce7d961e; no capability guard was performed for this r7 boundary, so these performance numbers are record-backed throughput evidence only and must not imply a capability pass.
run-0481 — pr25494-vulkan-q8kv-stock-r7
- 2026-08-19 · config → cfg-0127
- decode_tps=34.246633 · prefill_tps=289.052184
- ⚠ no guard — Throughput-only current-stock Vulkan rerun admitted by hbreviewer t_ce7d961e; no capability guard was performed for this r7 boundary, so these performance numbers are record-backed throughput evidence only and must not imply a capability pass.
run-0482 — pr25494-vulkan-q8kv-stock-r7
- 2026-08-19 · config → cfg-0127
- decode_tps=34.24946 · prefill_tps=280.524123
- ⚠ no guard — Throughput-only current-stock Vulkan rerun admitted by hbreviewer t_ce7d961e; no capability guard was performed for this r7 boundary, so these performance numbers are record-backed throughput evidence only and must not imply a capability pass.
run-0483 — pr25494-vulkan-q8kv-stock-r7
- 2026-08-19 · config → cfg-0128
- decode_tps=53.792515 · prefill_tps=583.475104
- ⚠ no guard — Throughput-only current-stock Vulkan rerun admitted by hbreviewer t_ce7d961e; no capability guard was performed for this r7 boundary, so these performance numbers are record-backed throughput evidence only and must not imply a capability pass.
run-0484 — pr25494-vulkan-q8kv-stock-r7
- 2026-08-19 · config → cfg-0128
- decode_tps=53.812841 · prefill_tps=587.976532
- ⚠ no guard — Throughput-only current-stock Vulkan rerun admitted by hbreviewer t_ce7d961e; no capability guard was performed for this r7 boundary, so these performance numbers are record-backed throughput evidence only and must not imply a capability pass.
run-0485 — pr25494-vulkan-q8kv-stock-r7
- 2026-08-19 · config → cfg-0128
- decode_tps=53.927361 · prefill_tps=585.73268
- ⚠ no guard — Throughput-only current-stock Vulkan rerun admitted by hbreviewer t_ce7d961e; no capability guard was performed for this r7 boundary, so these performance numbers are record-backed throughput evidence only and must not imply a capability pass.
run-0486 — pr25494-vulkan-q8kv-stock-r7
- 2026-08-19 · config → cfg-0128
- decode_tps=48.256168 · prefill_tps=399.005192
- ⚠ no guard — Throughput-only current-stock Vulkan rerun admitted by hbreviewer t_ce7d961e; no capability guard was performed for this r7 boundary, so these performance numbers are record-backed throughput evidence only and must not imply a capability pass.
run-0487 — pr25494-vulkan-q8kv-stock-r7
- 2026-08-19 · config → cfg-0128
- decode_tps=47.961215 · prefill_tps=398.333257
- ⚠ no guard — Throughput-only current-stock Vulkan rerun admitted by hbreviewer t_ce7d961e; no capability guard was performed for this r7 boundary, so these performance numbers are record-backed throughput evidence only and must not imply a capability pass.
run-0488 — pr25494-vulkan-q8kv-stock-r7
- 2026-08-19 · config → cfg-0128
- decode_tps=48.060848 · prefill_tps=398.363037
- ⚠ no guard — Throughput-only current-stock Vulkan rerun admitted by hbreviewer t_ce7d961e; no capability guard was performed for this r7 boundary, so these performance numbers are record-backed throughput evidence only and must not imply a capability pass.
run-0489 — pr25494-vulkan-q8kv-stock-r7
- 2026-08-19 · config → cfg-0128
- decode_tps=40.050751 · prefill_tps=246.449896
- ⚠ no guard — Throughput-only current-stock Vulkan rerun admitted by hbreviewer t_ce7d961e; no capability guard was performed for this r7 boundary, so these performance numbers are record-backed throughput evidence only and must not imply a capability pass.
run-0490 — pr25494-vulkan-q8kv-stock-r7
- 2026-08-19 · config → cfg-0128
- decode_tps=39.995436 · prefill_tps=246.044834
- ⚠ no guard — Throughput-only current-stock Vulkan rerun admitted by hbreviewer t_ce7d961e; no capability guard was performed for this r7 boundary, so these performance numbers are record-backed throughput evidence only and must not imply a capability pass.
run-0491 — pr25494-vulkan-q8kv-stock-r7
- 2026-08-19 · config → cfg-0128
- decode_tps=40.051932 · prefill_tps=246.12419
- ⚠ no guard — Throughput-only current-stock Vulkan rerun admitted by hbreviewer t_ce7d961e; no capability guard was performed for this r7 boundary, so these performance numbers are record-backed throughput evidence only and must not imply a capability pass.
run-0492 — llama-bench@1
run-0493 — llama-bench@1
run-0496 — llama-bench@1
run-0497 — llama-bench@1
run-0524 — stageA-c32768
run-0525 — stageA-c32768
run-0526 — stageA-c32768
run-0527 — stageA-c32768
run-0530 — llama-bench@1
run-0531 — llama-bench@1
run-0532 — llama-bench@1
run-0533 — llama-bench@1
run-0534 — llama-bench@1
run-0535 — llama-server-completion@1
run-0536 — llama-server-completion@1
run-0537 — llama-server-completion@1
run-0538 — llama-server-completion@1
run-0539 — llama-server-completion@1
run-0540 — llama-bench@1
run-0557 — llama-bench@1
- 2026-08-22 · config → cfg-0009
- decode_tps=21.919
- ⚠ no guard — llama-bench throughput ladder cell. No semantic/capability claim is made from this row: it records a fixed-depth decode throughput sample (terminal: this scalar is a hardware+KV-context-performance reading, not a model-capability result). The original control-guard record is no longer independently addressable in canonical content, so this row is an explicit guard gap and must not establish guard status or any comparison. control f16 d0
run-0558 — llama-bench@1
- 2026-08-22 · config → cfg-0009
- decode_tps=17.971
- ⚠ no guard — llama-bench throughput ladder cell. No semantic/capability claim is made from this row: it records a fixed-depth decode throughput sample (terminal: this scalar is a hardware+KV-context-performance reading, not a model-capability result). The original control-guard record is no longer independently addressable in canonical content, so this row is an explicit guard gap and must not establish guard status or any comparison. control f16 d32768
run-0559 — llama-bench@1
- 2026-08-22 · config → cfg-0009
- decode_tps=11.981
- ⚠ no guard — llama-bench throughput ladder cell. No semantic/capability claim is made from this row: it records a fixed-depth decode throughput sample (terminal: this scalar is a hardware+KV-context-performance reading, not a model-capability result). The originally named control-guard identifier is now assigned in the canonical record to an unrelated LLaDA capability run, so no independently addressable exact-fingerprint guard record remains. This row is an explicit guard gap and must not establish guard status or any comparison. control f16 d131072
run-0560 — llama-bench@1
- 2026-08-22 · config → cfg-0009
- decode_tps=9.341
- ⚠ no guard — llama-bench throughput ladder cell. No semantic/capability claim is made from this row: it records a fixed-depth decode throughput sample (terminal: this scalar is a hardware+KV-context-performance reading, not a model-capability result). The original control-guard record is no longer independently addressable in canonical content, so this row is an explicit guard gap and must not establish guard status or any comparison. control f16 d204800
run-0561 — llama-bench@1
- 2026-08-22 · config → cfg-0009
- decode_tps=8.17
- ⚠ no guard — llama-bench throughput ladder cell. No semantic/capability claim is made from this row: it records a fixed-depth decode throughput sample (terminal: this scalar is a hardware+KV-context-performance reading, not a model-capability result). The original control-guard record is no longer independently addressable in canonical content, so this row is an explicit guard gap and must not establish guard status or any comparison. control f16 d262144
run-0562 — llama-bench@1
- 2026-08-22 · config → cfg-0011
- decode_tps=21.806
- ⚠ no guard — llama-bench throughput ladder cell. No semantic/capability claim is made from this row: it records a fixed-depth decode throughput sample (terminal: this scalar is a hardware+KV-context-performance reading, not a model-capability result). Patched-arm capability guard evidence is run-0554. patched q8_0 d0
run-0563 — llama-bench@1
- 2026-08-22 · config → cfg-0011
- decode_tps=17.994
- ⚠ no guard — llama-bench throughput ladder cell. No semantic/capability claim is made from this row: it records a fixed-depth decode throughput sample (terminal: this scalar is a hardware+KV-context-performance reading, not a model-capability result). Patched-arm capability guard evidence is run-0554. patched q8_0 d32768
run-0564 — llama-bench@1
- 2026-08-22 · config → cfg-0011
- decode_tps=12.18
- ⚠ no guard — llama-bench throughput ladder cell. No semantic/capability claim is made from this row: it records a fixed-depth decode throughput sample (terminal: this scalar is a hardware+KV-context-performance reading, not a model-capability result). Patched-arm capability guard evidence is run-0554. patched q8_0 d131072
run-0565 — llama-bench@1
- 2026-08-22 · config → cfg-0011
- decode_tps=9.586
- ⚠ no guard — llama-bench throughput ladder cell. No semantic/capability claim is made from this row: it records a fixed-depth decode throughput sample (terminal: this scalar is a hardware+KV-context-performance reading, not a model-capability result). Patched-arm capability guard evidence is run-0554. patched q8_0 d204800
run-0566 — llama-bench@1
- 2026-08-22 · config → cfg-0011
- decode_tps=8.337
- ⚠ no guard — llama-bench throughput ladder cell. No semantic/capability claim is made from this row: it records a fixed-depth decode throughput sample (terminal: this scalar is a hardware+KV-context-performance reading, not a model-capability result). Patched-arm capability guard evidence is run-0554. patched q8_0 d262144
run-0579 — ho012-122b-prod-matrix-r2.cell1
run-0580 — ho012-122b-prod-matrix-r2.cell1
run-0581 — ho012-122b-prod-matrix-r2.cell1
run-0582 — ho012-122b-prod-matrix-r2.cell1
run-0583 — ho012-122b-prod-matrix-r2.cell1
run-0584 — ho012-122b-prod-matrix-r2.cell1
run-0585 — ho012-122b-prod-matrix-r2.cell2-depth
run-0586 — ho012-122b-prod-matrix-r2.cell3-f16
run-0587 — ho012-122b-prod-matrix-r2.cell3-q4_0
run-0592 — ho005-revised-ornith-q4kxl-batch-r3.control-b512-ub512
run-0593 — ho005-revised-ornith-q4kxl-batch-r3.b1024-ub512
run-0594 — ho005-revised-ornith-q4kxl-batch-r3.b2048-ub1024
run-0595 — ho005-revised-ornith-q4kxl-batch-r3.b4096-ub2048
run-0596 — ho005-revised-ornith-q4kxl-batch-r3.anchor-d262144
run-0602 — hk001-fullbench-fit@c32768
run-0603 — hk001-fullbench-depth-r3@llama-bench
run-0604 — hk001-fullbench-depth-r3@llama-bench
run-0606 — ho014a-fill-r1.cell1-latency
run-0607 — ho014a-fill-r1.cell1-latency
run-0608 — ho014a-fill-r1.cell1-latency
run-0609 — ho014a-fill-r1.cell1-latency
run-0610 — ho014a-fill-r1.cell1-latency
run-0611 — ho014a-fill-r1.cell1-latency
run-0612 — ho014a-fill-r1.cell2-depth
run-0613 — ho014a-fill-r1.cell2-depth
run-0614 — ho014a-fill-r1.cell2-depth
run-0616 — ho014-run-kv-overlap.cell3b-bench-pair
run-0617 — ho014-run-kv-overlap.cell3b-bench-pair
run-0619 — hk001-rerun-reasonoff.specsweep-spec-on
run-0620 — hk001-rerun-reasonoff.specsweep-spec-off
run-0626 — qwen38-prod-tuning.server-screen-code@1
- 2026-08-24 · config → cfg-0174
- decode_tps=52.2
- ⚠ no guard — No standard house guard, llama-bench floor, or matched plain/spec-off served-path control was supplied in the archived handoff. This is an intended-config server screening sample only; it must not be represented as production throughput.
run-0628 — qwen38-prod-tuning.cache-reuse-trace@45-turn
- 2026-08-24 · config → cfg-0175
- turns_before_degradation=45
- ⚠ no guard — No standard house guard or independent answer grader was supplied. The source shows successful returned snippets and no logged DeviceLost/lockup signature, but this is a cache-reuse trace, not a capability or production-throughput certification.
run-0629 — qwen38-prod-tuning.server-depth@target-32000
- 2026-08-24 · config → cfg-0176
- decode_tps=24.1 · prefill_tps=253.8
- ⚠ no guard — No standard house guard, llama-bench floor, or matched plain/spec-off served-path control was supplied. One served-path sample is retained at its actual prompt depth; it is not a production figure or a 32k cell.
run-0630 — qwen38-prod-tuning.server-depth@target-200000
- 2026-08-24 · config → cfg-0176
- decode_tps=13.8 · prefill_tps=125
- ⚠ no guard — No standard house guard, llama-bench floor, or matched plain/spec-off served-path control was supplied. One served-path sample is retained at its actual prompt depth; it is not a production figure or a 200k cell.
run-0631 — qwen38-prod-tuning.server-screen-prose@1
- 2026-08-24 · config → cfg-0174
- decode_tps=31.5
- ⚠ no guard — No standard house guard, llama-bench floor, or matched plain/spec-off served-path control was supplied in the archived handoff. This is an intended-config server screening sample only; it must not be represented as production throughput.
run-0633 — llama-bench-decode-depth
- 2026-08-27 · config → cfg-0178
- decode_tps=22.82 · prefill_tps=401.8
- ⚠ no guard — Throughput ladder; output-sanity inherited from run-0632 (tau2 capability, same binary/model/quant, valid outputs). No separate 4/4 needle guard for the perf cells.
run-0634 — llama-bench-decode-depth
- 2026-08-27 · config → cfg-0178
- decode_tps=14.09 · prefill_tps=201.3
- ⚠ no guard — Throughput ladder; output-sanity inherited from run-0632 (tau2 capability, same binary/model/quant, valid outputs). No separate 4/4 needle guard for the perf cells.
run-0635 — llama-bench-decode-depth
- 2026-08-27 · config → cfg-0178
- decode_tps=7.09 · prefill_tps=86.9
- ⚠ no guard — Throughput ladder; output-sanity inherited from run-0632 (tau2 capability, same binary/model/quant, valid outputs). No separate 4/4 needle guard for the perf cells.
run-0636 — llama-bench-decode-depth
- 2026-08-27 · config → cfg-0178
- decode_tps=5.73 · prefill_tps=63.9
- ⚠ no guard — Throughput ladder; output-sanity inherited from run-0632 (tau2 capability, same binary/model/quant, valid outputs). No separate 4/4 needle guard for the perf cells.
run-0640 — qsa-cache-false-prompt-processing
- 2026-08-30 · config → cfg-0180
- prefill_tps=405.750925
- ⚠ no guard — Same exact artifact/runtime/config as valid Tau2 trajectories; this record is limited to server-native prompt processing and makes no generation claim.
- ⚠ contaminated by: inc-0010
run-0641 — qsa-cache-false-prompt-processing
- 2026-08-30 · config → cfg-0180
- decode_tps=20.908478 · prefill_tps=336.149588
- ⚠ no guard — Same exact artifact/runtime/config as valid Tau2 trajectories; server-native nonzero-output prompt and generation timings are retained without transfer.
- ⚠ contaminated by: inc-0010
run-0642 — qsa-cache-false-prompt-processing
- 2026-08-30 · config → cfg-0180
- decode_tps=17.822403 · prefill_tps=266.788943
- ⚠ no guard — Same exact artifact/runtime/config as valid Tau2 trajectories; deep cells are single-pass under protocol and retain server-native nonzero-output timings.
- ⚠ contaminated by: inc-0010
run-0643 — qsa-cache-false-prompt-processing
- 2026-08-30 · config → cfg-0180
- decode_tps=14.942665 · prefill_tps=194.085064
- ⚠ no guard — Same exact artifact/runtime/config as valid Tau2 trajectories; deep cells are single-pass under protocol and retain server-native nonzero-output timings.
- ⚠ contaminated by: inc-0010
run-0644 — qsa-cache-false-prompt-processing
- 2026-08-30 · config → cfg-0180
- prefill_tps=121.838179
- ⚠ no guard — Same exact artifact/runtime/config as valid Tau2 trajectories; this record is limited to server-native prompt processing and makes no generation claim.
- ⚠ contaminated by: inc-0010
run-0645 — qsa-full-window-feasibility
- 2026-08-30 · config → cfg-0180
- context_refused=true
- ⚠ no guard — Feasibility outcome only; the server refused the full-window request before a timing payload or prompt_n, so no throughput metric exists.
- ⚠ contaminated by: inc-0010
run-0646 — qsa-prefix-reuse-trace
- 2026-08-30 · config → cfg-0180
- decode_tps=21.958091 · prefill_tps=48.228216
- ⚠ no guard — Controlled in-slot cache trace on the same admitted served config; deliberately separate from the cache_prompt:false depth ladder.
- ⚠ contaminated by: inc-0010
run-0647 — qsa-prefix-reuse-trace
- 2026-08-30 · config → cfg-0180
- decode_tps=18.0764 · prefill_tps=37.444068
- ⚠ no guard — Controlled in-slot cache trace on the same admitted served config; deliberately separate from the cache_prompt:false depth ladder.
- ⚠ contaminated by: inc-0010
run-0648 — llama-bench-source-config-mmap-raw
- 2026-08-30 · config → cfg-0181
- decode_tps=20.699823 · prefill_tps=97.658939
- ⚠ no guard — Raw engine-only source-config/default-mmap row; same artifact/runtime produced valid Tau2 outputs, but this row is not a served QSA or aggregate-floor claim.
- ⚠ contaminated by: inc-0010
run-0649 — llama-bench-source-config-mmap-raw
- 2026-08-30 · config → cfg-0181
- decode_tps=20.415509 · prefill_tps=208.431375
- ⚠ no guard — Raw engine-only source-config/default-mmap row; same artifact/runtime produced valid Tau2 outputs, but this row is not a served QSA or aggregate-floor claim.
- ⚠ contaminated by: inc-0010
run-0650 — llama-bench-source-config-mmap-raw
- 2026-08-30 · config → cfg-0181
- decode_tps=14.799038 · prefill_tps=125.882976
- ⚠ no guard — Raw N=1 engine-only source-config/default-mmap row; same artifact/runtime produced valid Tau2 outputs, but this is not a served QSA or aggregate-floor claim.
- ⚠ contaminated by: inc-0010
run-0651 — llama-bench-source-config-mmap-raw
- 2026-08-30 · config → cfg-0181
- decode_tps=12.101294 · prefill_tps=85.159391
- ⚠ no guard — Raw N=1 engine-only source-config/default-mmap row; same artifact/runtime produced valid Tau2 outputs, but this is not a served QSA or aggregate-floor claim.
- ⚠ contaminated by: inc-0010
run-0652 — llama-bench-source-config-mmap-raw
- 2026-08-30 · config → cfg-0181
- decode_tps=11.941786 · prefill_tps=74.734674
- ⚠ no guard — Raw N=1 engine-only source-config/default-mmap retry of the interrupted first attempt; same artifact/runtime produced valid Tau2 outputs, but this is not a served QSA or aggregate-floor claim.
- ⚠ contaminated by: inc-0010
run-0654 — decode-depth-served
- 2026-09-02 · config → cfg-0182
- decode_tps=16.1 · prefill_tps=298
- ⚠ no guard — Throughput cell; output-sanity inherited from run-0653 (tau2 capability, same binary/model/quant, valid outputs). No separate 4/4 needle guard for the perf cells.
run-0655 — decode-depth-served
- 2026-09-02 · config → cfg-0182
- decode_tps=16 · prefill_tps=219
- ⚠ no guard — Throughput cell; output-sanity inherited from run-0653 (tau2 capability, same binary/model/quant, valid outputs). No separate 4/4 needle guard for the perf cells.
run-0656 — decode-depth-served
- 2026-09-02 · config → cfg-0182
- decode_tps=22.8 · prefill_tps=169
- ⚠ no guard — Throughput cell; output-sanity inherited from run-0653 (tau2 capability, same binary/model/quant, valid outputs). No separate 4/4 needle guard for the perf cells.
run-0657 — decode-depth-served
- 2026-09-02 · config → cfg-0182
- decode_tps=26.1 · prefill_tps=127
- ⚠ no guard — Throughput cell; output-sanity inherited from run-0653 (tau2 capability, same binary/model/quant, valid outputs). No separate 4/4 needle guard for the perf cells.
run-0659 — slotpin-proxy-metrics@tau2-window
run-0661 — decode-depth-served@drive_depth_served
- 2026-09-05 · config → cfg-0184
- decode_tps=19.23 · prefill_tps=238.9
- ⚠ no guard — Single-slot depth cell on the pr27311 build. The #25992 multi-slot leak is absent by construction at --parallel 1 (one slot cannot leak across slots); self-spec MTP is a neutral lever; the worker's capability is measured on the sibling cfg-0183 (run-0658, 91.7%); and this cell produced a full non-empty 256-token generation, its own empty-output cliff check. Recorded per the guard-substitution note.
- energy window: eng-0279
run-0662 — decode-depth-served@drive_depth_served
- 2026-09-05 · config → cfg-0184
- decode_tps=16.99 · prefill_tps=141.4
- ⚠ no guard — Single-slot depth cell on the pr27311 build; #25992 absent at --parallel 1, self-spec MTP neutral, capability measured on cfg-0183 (run-0658), full non-empty 256-token generation produced. Same waiver as run-0661.
- energy window: eng-0279
run-0663 — decode-depth-served@drive_depth_served
- 2026-09-05 · config → cfg-0184
- decode_tps=14.71 · prefill_tps=74.4
- ⚠ no guard — Single-slot depth cell on the pr27311 build; #25992 absent at --parallel 1, self-spec MTP neutral, capability measured on cfg-0183 (run-0658), full non-empty 256-token generation produced. Same waiver as run-0661.
- energy window: eng-0279
run-0664 — decode-depth-served@drive_depth_served
- 2026-09-05 · config → cfg-0184
- decode_tps=12.49 · prefill_tps=49.7
- ⚠ no guard — Single-slot depth cell on the pr27311 build; #25992 absent at --parallel 1, self-spec MTP neutral, capability measured on cfg-0183 (run-0658), full non-empty 256-token generation produced. Same waiver as run-0661.
- energy window: eng-0279
run-0667 — slotpin-mslot@dflash-pm4
run-0668 — decode-depth-served@drive_depth_served
- 2026-09-06 · config → cfg-0186
- decode_tps=27.98 · prefill_tps=319
- ⚠ no guard — Single-slot depth cell, served via slotpin. Capability of this candidate is measured on the sibling cfg-0185 (run-0665); the DFlash2 empty-output cliff is gated by run-0666 (0 EMPTY / 66 min), and this cell produced a full non-empty 256-token generation.
run-0669 — decode-depth-served@drive_depth_served
- 2026-09-06 · config → cfg-0186
- decode_tps=23.18 · prefill_tps=198.7
- ⚠ no guard — Single-slot depth cell; capability on cfg-0185 (run-0665), cliff gated by run-0666, full non-empty 256-token generation. Same waiver as run-0668.
run-0670 — decode-depth-served@drive_depth_served
- 2026-09-06 · config → cfg-0186
- decode_tps=22.12 · prefill_tps=127.4
- ⚠ no guard — Single-slot depth cell; capability on cfg-0185 (run-0665), cliff gated by run-0666, full non-empty 256-token generation. Same waiver as run-0668.
run-0671 — decode-depth-served@drive_depth_served
- 2026-09-06 · config → cfg-0186
- decode_tps=17.07 · prefill_tps=83.9
- ⚠ no guard — Single-slot depth cell; capability on cfg-0185 (run-0665), cliff gated by run-0666, full non-empty 256-token generation. Same waiver as run-0668.
run-0672 — pm4-isolation@retained-on
- 2026-09-06 · config → cfg-0187
- decode_tps=20.17 · prefill_tps=262
- ⚠ no guard — Isolation probe of OUR worker model/spec (UD-Q4_K_XL + self-spec MTP, capability measured on cfg-0183/run-0658) run on the pwilkin build. Single-slot, self-spec MTP confirmed against the target; full non-empty 256-token generation.
run-0673 — pm4-isolation@retained-off
- 2026-09-06 · config → cfg-0187
- decode_tps=19.45 · prefill_tps=264.6
- ⚠ no guard — Same isolation probe as run-0672 with retained-PM4 OFF (identical binary/model/flags, only the env toggle differs). Capability of this model/spec is on cfg-0183/run-0658.
Energy (280)
Wh-denominated, wall-metered, read as a cumulative-counter difference over the window. delta_w over the idle baseline is the only figure that means anything.
idle-aihydra-2026-08 — idle baseline, aihydra
- period 2026-08-07 → 2026-08-10 · method wall-meter
- box-idle-all-empty: 10.1 W · 14 h
- igpu-active: 168 W · 53 h
- igpu-resident-quiet: 13.6 W · 0.17 h
- npu-resident-quiet: unmeasured — an honest gap beats an invented number
- standing: 7.4 kWh/mo · £2.24/mo · duty cycle 0.79
- First idle_baseline record for aihydra, covering its first days of service. The box-idle floor of 10.1 W is remarkably low for a 128 GB machine and is the figure every delta_w in the energy records subtracts from. igpu-resident-quiet MEASURED 2026-08-11 (eng-0035: 10 minutes, 73 GB model loaded, zero requests): **13.6 W mean, peak 35.1 W** — just +3.5 W over the empty-box floor. The keep_alive question is answered: holding the 122B resident costs ~2.5 kWh/month, ~76p at 30.3 p/kWh. This also anchors energy attribution for cloud-simulator arms: with resident-quiet this low, a high mean during such arms means genuine computation, not wait-state draw. `npu-resident-quiet` remains unmeasured — still the question an always-on NPU triage loop turns on. Duty cycle is 0.79 — this box has been working, not idling, across the period. That is unrepresentative of steady state and will fall once the benchmark programme ends; the standing-cost figures should be re-derived then rather than quoted from this window. Standing cost assumes 30.3 p/kWh grid import. With solar and battery the realised cost is lower and sometimes zero, which is what the energy records' tariff_window field exists to capture.
eng-0001
- 2026-08-08 · wall-meter · TP-Link smart plug via Home Assistant · window reconstructed
- wh_total 4.977 · mean 59.7 W · baseline 10.1 W (box-idle-all-empty) · delta 49.6 W
- lead-in assumed 5min (no prior invocation within 1h). invocation JSON is 0 bytes (empty output) — the run failed to record results, so no run can be linked.
eng-0002
- 2026-08-08 · wall-meter · TP-Link smart plug via Home Assistant · window reconstructed
- wh_total 26.164 · mean 88.1 W · baseline 10.1 W (box-idle-all-empty) · delta 78 W
- covers: run-0008 run-0009 run-0010 run-0011 run-0012 run-0013
- window = previous invocation end. matched to cfg-0008 (commit 3653e6d, fa=0, ctk/ctv=f16) by exact avg_ts value match against all 6 rows.
eng-0003
- 2026-08-08 · wall-meter · TP-Link smart plug via Home Assistant · window reconstructed
- wh_total 12.32 · mean 168.5 W · baseline 10.1 W (box-idle-all-empty) · delta 158.4 W
- covers: run-0014 run-0015 run-0016 run-0017 run-0018 run-0019
- window = previous invocation end. matched to cfg-0009 (commit 3653e6d, fa=1, ctk/ctv=f16) by exact avg_ts value match against all 6 rows.
eng-0004
- 2026-08-08 · wall-meter · TP-Link smart plug via Home Assistant · window reconstructed
- wh_total 13.087 · mean 129.3 W · baseline 10.1 W (box-idle-all-empty) · delta 119.2 W
- covers: run-0020 run-0021 run-0022 run-0023 run-0024 run-0025
- window = previous invocation end. matched to cfg-0010 (commit 3653e6d, fa=1, ctk/ctv=q8_0) by exact avg_ts value match against all 6 rows.
eng-0005
- 2026-08-08 · wall-meter · TP-Link smart plug via Home Assistant · window reconstructed
- wh_total 29.199 · mean 175.5 W · baseline 10.1 W (box-idle-all-empty) · delta 165.4 W
- window = previous invocation end. model_filename is Qwen3.6-27B-Q4_K_M.gguf — no content/configs record exists for this model at all; appears to be an exploratory flash-attn cliff-detector sweep (depths 0/32768/65536) that was never promoted to a run.
eng-0006
- 2026-08-08 · wall-meter · TP-Link smart plug via Home Assistant · window reconstructed
- wh_total 38.37 · mean 173.4 W · baseline 10.1 W (box-idle-all-empty) · delta 163.3 W
- window = previous invocation end. model_filename is Qwen3.6-27B-Q4_K_M.gguf — no content/configs record exists for this model at all; appears to be an exploratory flash-attn cliff-detector sweep (depths 0/32768/65536) that was never promoted to a run.
eng-0007
- 2026-08-08 · wall-meter · TP-Link smart plug via Home Assistant · window reconstructed
- wh_total 15.225 · mean 139.6 W · baseline 10.1 W (box-idle-all-empty) · delta 129.5 W
- window = previous invocation end. signature (commit 3653e6d, fa=1, ctk/ctv=f16, model qwen35-122b) matches cfg-0009's config, but avg_ts values (0-depth prefill 326.36 vs run-0014's 321.72) don't exactly match any promoted run — looks like a preliminary A/B rerun superseded by the sweep2 sweep, not itself promoted.
eng-0008
- 2026-08-08 · wall-meter · TP-Link smart plug via Home Assistant · window reconstructed
- wh_total 10.678 · mean 160.6 W · baseline 10.1 W (box-idle-all-empty) · delta 150.5 W
- window = previous invocation end. signature (commit 3653e6d, fa=1, ctk/ctv=q8_0, model qwen35-122b) matches cfg-0010's config, but avg_ts values don't exactly match any promoted run — preliminary A/B rerun, not itself promoted.
eng-0009
- 2026-08-08 · wall-meter · TP-Link smart plug via Home Assistant · window reconstructed
- wh_total 10.22 · mean 158.8 W · baseline 10.1 W (box-idle-all-empty) · delta 148.7 W
- window = previous invocation end. signature is commit ce7689f, fa=1, ctk/ctv=f16 — no content/configs record combines that commit with f16 KV (cfg-0011/0013/0014 on ce7689f are all q8_0), so no run to link.
eng-0010
- 2026-08-08 · wall-meter · TP-Link smart plug via Home Assistant · window reconstructed
- wh_total 10.167 · mean 151.2 W · baseline 10.1 W (box-idle-all-empty) · delta 141.1 W
- covers: run-0026 run-0027 run-0028 run-0029
- window = previous invocation end. matched to cfg-0011 (commit ce7689f, fa=1, ctk/ctv=q8_0) by exact avg_ts value match against all 4 rows.
eng-0011
- 2026-08-08 · wall-meter · TP-Link smart plug via Home Assistant · window reconstructed
- wh_total 9.025 · mean 127 W · baseline 10.1 W (box-idle-all-empty) · delta 116.9 W
- covers: run-0030 run-0031 run-0032 run-0033 run-0034 run-0035
- window = previous invocation end. matched to cfg-0012 by exact avg_ts value match against all 6 rows.
eng-0012
- 2026-08-08 · wall-meter · TP-Link smart plug via Home Assistant · window reconstructed
- wh_total 45.193 · mean 169.4 W · baseline 10.1 W (box-idle-all-empty) · delta 159.3 W
- window = previous invocation end. commit 3653e6d + ctk/ctv f16 signature, single row (depth 131072) — no cfg record combines 3653e6d with a 131072-depth f16 test, and its avg_ts (12.031) matches no run.
eng-0013
- 2026-08-08 · wall-meter · TP-Link smart plug via Home Assistant · window reconstructed
- wh_total 45.466 · mean 168.8 W · baseline 10.1 W (box-idle-all-empty) · delta 158.7 W
- window = previous invocation end. commit 3653e6d + ctk/ctv q8_0, single row (depth 131072) — avg_ts (7.812) matches no run; appears superseded by the deepkv200-082221 rerun at the larger 204800 depth.
eng-0014
- 2026-08-08 · wall-meter · TP-Link smart plug via Home Assistant · window reconstructed
- wh_total 44.91 · mean 175.1 W · baseline 10.1 W (box-idle-all-empty) · delta 165 W
- window = previous invocation end. commit ce7689f + ctk/ctv f16, single row (depth 131072) — no cfg record combines that commit with f16 KV, and avg_ts (12.043) matches no run.
eng-0015
- 2026-08-08 · wall-meter · TP-Link smart plug via Home Assistant · window reconstructed
- wh_total 43.512 · mean 174.4 W · baseline 10.1 W (box-idle-all-empty) · delta 164.3 W
- covers: run-0036
- window = previous invocation end. matched to cfg-0013 (commit ce7689f, ctk/ctv q8_0, depth 131072) by exact avg_ts value match.
eng-0016
- 2026-08-08 · wall-meter · TP-Link smart plug via Home Assistant · window reconstructed
- wh_total 0.842 · mean 10.1 W · baseline 10.1 W (box-idle-all-empty) · delta 0 W
- lead-in assumed 5min (no prior invocation within 1h). invocation JSON is 0 bytes (empty output) — the run failed to record results, so no run can be linked.
eng-0017
- 2026-08-08 · wall-meter · TP-Link smart plug via Home Assistant · window reconstructed
- wh_total 146.082 · mean 171.7 W · baseline 10.1 W (box-idle-all-empty) · delta 161.6 W
- window = previous invocation end. commit 3653e6d + ctk/ctv q8_0, single row (depth 204800) — avg_ts (5.686) matches no run; no cfg record combines 3653e6d with a 204800-depth q8_0 test.
eng-0018
- 2026-08-08 · wall-meter · TP-Link smart plug via Home Assistant · window reconstructed
- wh_total 85.248 · mean 178.5 W · baseline 10.1 W (box-idle-all-empty) · delta 168.4 W
- covers: run-0037
- window = previous invocation end. matched to cfg-0014 (commit ce7689f, ctk/ctv q8_0, depth 204800) by exact avg_ts value match.
eng-0019
- 2026-08-08 · wall-meter · TP-Link smart plug via Home Assistant · window reconstructed
- wh_total 89.553 · mean 177.7 W · baseline 10.1 W (box-idle-all-empty) · delta 167.6 W
- window = previous invocation end. commit 3653e6d + ctk/ctv f16, single row (depth 204800) — avg_ts (9.599) matches no run.
eng-0020
- 2026-08-08 · wall-meter · TP-Link smart plug via Home Assistant · window reconstructed
- wh_total 90.427 · mean 177.7 W · baseline 10.1 W (box-idle-all-empty) · delta 167.6 W
- window = previous invocation end. commit ce7689f + ctk/ctv f16, single row (depth 204800) — no cfg record combines that commit with f16 KV, and avg_ts (9.539) matches no run.
eng-0021
- 2026-08-08 · wall-meter · TP-Link smart plug via Home Assistant · window reconstructed
- wh_total 19.658 · mean 46.9 W · baseline 10.1 W (box-idle-all-empty) · delta 36.8 W
- window = previous invocation end. cold.json/warm.json/diverged.json are empty and slots.json is a raw /slots endpoint dump ({n_ctx, speculative, is_processing}), not a llama-bench result — this is a disk-warmth/slot-restore diagnostic, not a benchmark run, so there is no run collection entry to link.
eng-0022
- 2026-08-08 · wall-meter · TP-Link smart plug via Home Assistant · window reconstructed
- wh_total 7.714 · mean 121 W · baseline 10.1 W (box-idle-all-empty) · delta 110.9 W
- covers: run-0044 run-0045 run-0046 run-0047 run-0048 run-0049
- window = previous invocation end. matched to cfg-0016 by exact avg_ts value match against all 6 rows.
eng-0023
- 2026-08-08 · wall-meter · TP-Link smart plug via Home Assistant · window reconstructed
- wh_total 14.858 · mean 151.1 W · baseline 10.1 W (box-idle-all-empty) · delta 141 W
- covers: run-0038 run-0039 run-0040 run-0041 run-0042 run-0043
- window = previous invocation end. matched to cfg-0015 by exact avg_ts value match against 5 rows; the depth-0 decode row (17.506 t/s) is shared verbatim with cfg-0020's depth-0 decode row, disambiguated by the other 5 rows in this file all keying to cfg-0015.
eng-0024
- 2026-08-08 · wall-meter · TP-Link smart plug via Home Assistant · window reconstructed
- wh_total 16.943 · mean 37.7 W · baseline 10.1 W (box-idle-all-empty) · delta 27.6 W
- window = previous invocation end. cold.json/warm.json/diverged.json are empty and slots.json is a raw /slots endpoint dump, not a llama-bench result — disk-warmth/slot-restore diagnostic, no run collection entry to link.
eng-0025
- 2026-08-08 · wall-meter · TP-Link smart plug via Home Assistant · window reconstructed
- wh_total 44.231 · mean 103.8 W · baseline 10.1 W (box-idle-all-empty) · delta 93.7 W
- covers: run-0050 run-0051 run-0052 run-0053 run-0054 run-0055
- window = previous invocation end. matched to cfg-0017 by exact avg_ts value match against all 6 rows.
eng-0026
- 2026-08-08 · wall-meter · TP-Link smart plug via Home Assistant · window reconstructed
- wh_total 142.65 · mean 173.3 W · baseline 10.1 W (box-idle-all-empty) · delta 163.2 W
- covers: run-0056 run-0057 run-0058 run-0059 run-0060 run-0061
- window = previous invocation end. matched to cfg-0018 by exact avg_ts value match against all 6 rows.
eng-0027
- 2026-08-08 · wall-meter · TP-Link smart plug via Home Assistant · window reconstructed
- wh_total 22.565 · mean 159.7 W · baseline 10.1 W (box-idle-all-empty) · delta 149.6 W
- covers: run-0086 run-0087 run-0088 run-0089 run-0090 run-0091
- window = previous invocation end. matched to cfg-0023 by exact avg_ts value match against all 6 rows.
eng-0028
- 2026-08-08 · wall-meter · TP-Link smart plug via Home Assistant · window reconstructed
- wh_total 28.854 · mean 179.2 W · baseline 10.1 W (box-idle-all-empty) · delta 169.1 W
- covers: run-0080 run-0081 run-0082 run-0083 run-0084 run-0085
- window = previous invocation end. matched to cfg-0022 by exact avg_ts value match against all 6 rows.
eng-0029
- 2026-08-08 · wall-meter · TP-Link smart plug via Home Assistant · window reconstructed
- wh_total 20.58 · mean 176.5 W · baseline 10.1 W (box-idle-all-empty) · delta 166.4 W
- covers: run-0092 run-0093 run-0094 run-0095 run-0096 run-0097
- window = previous invocation end. matched to cfg-0024 by exact avg_ts value match against all 6 rows.
eng-0030
- 2026-08-08 · wall-meter · TP-Link smart plug via Home Assistant · window reconstructed
- wh_total 48.097 · mean 172.2 W · baseline 10.1 W (box-idle-all-empty) · delta 162.1 W
- covers: run-0068 run-0069 run-0070 run-0071 run-0072 run-0073
- window = previous invocation end. matched to cfg-0020 by exact avg_ts value match against 5 rows; the depth-0 decode row (17.506 t/s) is shared verbatim with cfg-0015's depth-0 decode row, disambiguated by the other 5 rows in this file all keying to cfg-0020.
eng-0031
- 2026-08-08 · wall-meter · TP-Link smart plug via Home Assistant · window reconstructed
- wh_total 63.404 · mean 178 W · baseline 10.1 W (box-idle-all-empty) · delta 167.9 W
- covers: run-0062 run-0063 run-0064 run-0065 run-0066 run-0067
- window = previous invocation end. matched to cfg-0019 by exact avg_ts value match against all 6 rows.
eng-0032
- 2026-08-08 · wall-meter · TP-Link smart plug via Home Assistant · window reconstructed
- wh_total 48.764 · mean 178.3 W · baseline 10.1 W (box-idle-all-empty) · delta 168.2 W
- covers: run-0074 run-0075 run-0076 run-0077 run-0078 run-0079
- window = previous invocation end. matched to cfg-0021 by exact avg_ts value match against all 6 rows.
eng-0033
- 2026-08-09 · wall-meter · TP-Link smart plug via Home Assistant · window reconstructed
- wh_total 38.246 · mean 156.5 W · baseline 10.1 W (box-idle-all-empty) · delta 146.4 W
- per task: 12.749 Wh · 0.3863 p (tariff 30.3 p/kWh)
- covers: run-0098
- tau2 sim dir 20260809_144711 (airline, n=3, reasoning_content present -> think); picked as best- effort match for run-0098 (tasks_total=3, max_steps=40 matching the nothink group's config) out of 3 candidate reruns with identical reward pattern.
eng-0034
- 2026-08-09 · wall-meter · TP-Link smart plug via Home Assistant · window reconstructed
- wh_total 28.564 · mean 169.2 W · baseline 10.1 W (box-idle-all-empty) · delta 159.1 W
- per task: 5.713 Wh · 0.1731 p (tariff 30.3 p/kWh)
- covers: run-0099
- tau2 sim dir 20260809_161821 (airline, n=5, no reasoning_content -> nothink); picked as best- effort match for run-0099 (tasks_total=5, all rewards 1.0) as the earliest fully-passing nothink rerun following the chosen think run, out of 5 candidate fully-passing reruns.
eng-0035
- 2026-08-11 · wall-meter · AI Hydra energy/power sensors via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy; peak/mean drawn from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 2.275 · mean 13.6 W · baseline 10.1 W (box-idle-all-empty) · delta 3.5 W
- aihydra run-meta.jsonl entry: suite=idle-baseline, model=qwen35-122b, variant="resident- quiet". no run-NNNN file exists yet for this window; run left unlinked ([]).
eng-0036
- 2026-08-11 · wall-meter · AI Hydra energy/power sensors via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy; peak/mean drawn from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 0.494 · mean 29.6 W · baseline 10.1 W (box-idle-all-empty) · delta 19.5 W
- aihydra run-meta.jsonl entry: suite=tau2-kvpatch, model=qwen35-122b, variant="kvfix-f16". rc=2 (failed fast, 53s) — early-exit failure, window kept as it is a real recorded draw, not a zero-duration preflight guard. no run-NNNN file exists yet for this window; run left unlinked ([]).
eng-0037
- 2026-08-11 · wall-meter · AI Hydra energy/power sensors via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy; peak/mean drawn from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 0.434 · mean 31.3 W · baseline 10.1 W (box-idle-all-empty) · delta 21.2 W
- aihydra run-meta.jsonl entry: suite=tau2-kvpatch, model=qwen35-122b, variant="kvfix-q8". rc=2 (failed fast, 55s) — early-exit failure, window kept as it is a real recorded draw, not a zero-duration preflight guard. no run-NNNN file exists yet for this window; run left unlinked ([]).
eng-0038
- 2026-08-11 · wall-meter · AI Hydra energy/power sensors via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy; peak/mean drawn from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 116.723 · mean 65.3 W · baseline 10.1 W (box-idle-all-empty) · delta 55.2 W
- aihydra run-meta.jsonl entry: suite=llama-bench, model=nemotron3-super, variant="UD-Q4_K_M-first-look". no run-NNNN file exists yet for this window; run left unlinked ([]).
eng-0039
- 2026-08-11 · wall-meter · AI Hydra energy/power sensors via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy; peak/mean drawn from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 620.052 · mean 154.5 W · baseline 10.1 W (box-idle-all-empty) · delta 144.4 W
- covers: run-0103
- aihydra run-meta.jsonl entry: suite=tau2-kvpatch, model=qwen35-122b, variant="kvfix-f16". rc=124 (timed out) — window is real, arm did not complete but drew power the whole time. linked to run-0103 (ingested 2026-08-12 via scripts/ingest-tau2.mjs).
eng-0040
- 2026-08-11 · wall-meter · AI Hydra energy/power sensors via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy; peak/mean drawn from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 170.496 · mean 164 W · baseline 10.1 W (box-idle-all-empty) · delta 153.9 W
- covers: run-0104
- aihydra run-meta.jsonl entry: suite=tau2-kvpatch, model=qwen35-122b, variant="kvfix-q8". linked to run-0104 (ingested 2026-08-12 via scripts/ingest-tau2.mjs).
eng-0041
- 2026-08-11 · wall-meter · AI Hydra energy/power sensors via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy; peak/mean drawn from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 97.594 · mean 63.1 W · baseline 10.1 W (box-idle-all-empty) · delta 53 W
- aihydra run-meta.jsonl entry: suite=anchor-bench, model=qwen35-122b, variant="OLD-3653e6d". no run-NNNN file exists yet for this window; run left unlinked ([]).
eng-0042
- 2026-08-11 · wall-meter · AI Hydra energy/power sensors via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy; peak/mean drawn from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 55.131 · mean 162.7 W · baseline 10.1 W (box-idle-all-empty) · delta 152.6 W
- covers: run-0105
- aihydra run-meta.jsonl entry: suite=anchor-tau2, model=qwen35-122b, variant="OLD-3653e6d". linked to run-0105 (ingested 2026-08-12 via scripts/ingest-tau2.mjs).
eng-0043
- 2026-08-11 · wall-meter · AI Hydra energy/power sensors via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy; peak/mean drawn from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 118.131 · mean 61.1 W · baseline 10.1 W (box-idle-all-empty) · delta 51 W
- aihydra run-meta.jsonl entry: suite=anchor-bench, model=qwen35-122b, variant="NEW-62bf73d". no run-NNNN file exists yet for this window; run left unlinked ([]).
eng-0044
- 2026-08-11 · wall-meter · AI Hydra energy/power sensors via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy; peak/mean drawn from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 55.132 · mean 158.8 W · baseline 10.1 W (box-idle-all-empty) · delta 148.7 W
- covers: run-0106
- aihydra run-meta.jsonl entry: suite=anchor-tau2, model=qwen35-122b, variant="NEW-62bf73d". linked to run-0106 (ingested 2026-08-12 via scripts/ingest-tau2.mjs).
eng-0045
- 2026-08-11 · wall-meter · AI Hydra energy/power sensors via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy; peak/mean drawn from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 0.153 · mean 27.5 W · baseline 10.1 W (box-idle-all-empty) · delta 17.4 W
- aihydra run-meta.jsonl entry: suite=muse-fit, model=muse-glimmer-30b, variant="min-62bf73d". no run-NNNN file exists yet for this window; run left unlinked ([]).
eng-0046
- 2026-08-11 · wall-meter · AI Hydra energy/power sensors via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy; peak/mean drawn from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 1.861 · mean 167.5 W · baseline 10.1 W (box-idle-all-empty) · delta 157.4 W
- aihydra run-meta.jsonl entry: suite=muse-guard, model=muse-glimmer-30b, variant="min-62bf73d". rc=1 (failed) — window kept, real recorded draw. no run-NNNN file exists yet for this window; run left unlinked ([]).
eng-0047
- 2026-08-11 · wall-meter · AI Hydra energy/power sensors via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy; peak/mean drawn from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 86.118 · mean 172.2 W · baseline 10.1 W (box-idle-all-empty) · delta 162.1 W
- aihydra run-meta.jsonl entry: suite=muse-smoke, model=muse-glimmer-30b, variant="min-62bf73d". rc=124 (timed out) — window is real, arm did not complete but drew power the whole time. no run-NNNN file exists yet for this window; run left unlinked ([]).
eng-0048
- 2026-08-11 · wall-meter · AI Hydra energy/power sensors via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy; peak/mean drawn from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 1.061 · mean 54.6 W · baseline 10.1 W (box-idle-all-empty) · delta 44.5 W
- aihydra run-meta.jsonl entry: suite=laguna-fit, model=laguna-s21, variant="min-62bf73d". no run-NNNN file exists yet for this window; run left unlinked ([]).
eng-0049
- 2026-08-11 · wall-meter · AI Hydra energy/power sensors via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy; peak/mean drawn from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 0.995 · mean 89.6 W · baseline 10.1 W (box-idle-all-empty) · delta 79.5 W
- aihydra run-meta.jsonl entry: suite=bf16pr-matrix, model=qwen36-35b, variant="ROCm0-f16f16-p1024n256". no run-NNNN file exists yet for this window; run left unlinked ([]).
eng-0050
- 2026-08-11 · wall-meter · AI Hydra energy/power sensors via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy; peak/mean drawn from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 15.113 · mean 183.5 W · baseline 10.1 W (box-idle-all-empty) · delta 173.4 W
- aihydra run-meta.jsonl entry: suite=bf16pr-matrix, model=qwen36-35b, variant="ROCm0-f16f16-p32768proxy". no run-NNNN file exists yet for this window; run left unlinked ([]).
eng-0051
- 2026-08-11 · wall-meter · AI Hydra energy/power sensors via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy; peak/mean drawn from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 1.492 · mean 107.4 W · baseline 10.1 W (box-idle-all-empty) · delta 97.3 W
- aihydra run-meta.jsonl entry: suite=bf16pr-matrix, model=qwen36-35b, variant="ROCm0-bf16bf16-p1024n256". no run-NNNN file exists yet for this window; run left unlinked ([]).
eng-0052
- 2026-08-11 · wall-meter · AI Hydra energy/power sensors via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy; peak/mean drawn from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 14.041 · mean 179.1 W · baseline 10.1 W (box-idle-all-empty) · delta 169 W
- aihydra run-meta.jsonl entry: suite=bf16pr-matrix, model=qwen36-35b, variant="ROCm0-bf16bf16-p32768proxy". no run-NNNN file exists yet for this window; run left unlinked ([]).
eng-0053
- 2026-08-11 · wall-meter · AI Hydra energy/power sensors via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy; peak/mean drawn from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 1.449 · mean 104.3 W · baseline 10.1 W (box-idle-all-empty) · delta 94.2 W
- aihydra run-meta.jsonl entry: suite=bf16pr-matrix, model=qwen36-35b, variant="Vulkan0-f16f16-p1024n256". no run-NNNN file exists yet for this window; run left unlinked ([]).
eng-0054
- 2026-08-11 · wall-meter · AI Hydra energy/power sensors via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy; peak/mean drawn from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 13.352 · mean 176.8 W · baseline 10.1 W (box-idle-all-empty) · delta 166.7 W
- aihydra run-meta.jsonl entry: suite=bf16pr-matrix, model=qwen36-35b, variant="Vulkan0-f16f16-p32768proxy". no run-NNNN file exists yet for this window; run left unlinked ([]).
eng-0055
- 2026-08-11 · wall-meter · AI Hydra energy/power sensors via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy; peak/mean drawn from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 1.097 · mean 98.7 W · baseline 10.1 W (box-idle-all-empty) · delta 88.6 W
- aihydra run-meta.jsonl entry: suite=bf16pr-matrix, model=qwen36-35b, variant="Vulkan0-bf16bf16-p1024n256". no run-NNNN file exists yet for this window; run left unlinked ([]).
eng-0056
- 2026-08-11 · wall-meter · AI Hydra energy/power sensors via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy; peak/mean drawn from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 17.168 · mean 176.5 W · baseline 10.1 W (box-idle-all-empty) · delta 166.4 W
- aihydra run-meta.jsonl entry: suite=bf16pr-matrix, model=qwen36-35b, variant="Vulkan0-bf16bf16-p32768proxy". no run-NNNN file exists yet for this window; run left unlinked ([]).
eng-0057
- 2026-08-11 · wall-meter · AI Hydra energy/power sensors via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy; peak/mean drawn from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 4.157 · mean 149.7 W · baseline 10.1 W (box-idle-all-empty) · delta 139.6 W
- aihydra run-meta.jsonl entry: suite=bf16pr-matrix-depthfix, model=qwen36-35b, variant="ROCm0-f16f16-p1024n256d32768-realdepth". no run-NNNN file exists yet for this window; run left unlinked ([]).
eng-0058
- 2026-08-11 · wall-meter · AI Hydra energy/power sensors via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy; peak/mean drawn from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 3.916 · mean 156.6 W · baseline 10.1 W (box-idle-all-empty) · delta 146.5 W
- aihydra run-meta.jsonl entry: suite=bf16pr-matrix-depthfix, model=qwen36-35b, variant="ROCm0-bf16bf16-p1024n256d32768-realdepth". no run-NNNN file exists yet for this window; run left unlinked ([]).
eng-0059
- 2026-08-11 · wall-meter · AI Hydra energy/power sensors via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy; peak/mean drawn from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 3.781 · mean 151.2 W · baseline 10.1 W (box-idle-all-empty) · delta 141.1 W
- aihydra run-meta.jsonl entry: suite=bf16pr-matrix-depthfix, model=qwen36-35b, variant="Vulkan0-f16f16-p1024n256d32768-realdepth". no run-NNNN file exists yet for this window; run left unlinked ([]).
eng-0060
- 2026-08-11 · wall-meter · AI Hydra energy/power sensors via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy; peak/mean drawn from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 4.834 · mean 158.2 W · baseline 10.1 W (box-idle-all-empty) · delta 148.1 W
- aihydra run-meta.jsonl entry: suite=bf16pr-matrix-depthfix, model=qwen36-35b, variant="Vulkan0-bf16bf16-p1024n256d32768-realdepth". no run-NNNN file exists yet for this window; run left unlinked ([]).
eng-0061
- 2026-08-11 · wall-meter · AI Hydra energy/power sensors via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy; peak/mean drawn from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 427.416 · mean 165.2 W · baseline 10.1 W (box-idle-all-empty) · delta 155.1 W
- covers: run-0107
- aihydra run-meta.jsonl entry: suite=nemotron-quant, model=nemotron3-super, variant="E1". linked to run-0107 (ingested 2026-08-12 via scripts/ingest-tau2.mjs).
eng-0062
- 2026-08-11 · wall-meter · AI Hydra energy/power sensors via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy; peak/mean drawn from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 234.142 · mean 168 W · baseline 10.1 W (box-idle-all-empty) · delta 157.9 W
- covers: run-0108
- aihydra run-meta.jsonl entry: suite=nemotron-quant, model=nemotron3-super, variant="E2". linked to run-0108 (ingested 2026-08-12 via scripts/ingest-tau2.mjs).
eng-0063
- 2026-08-11 · wall-meter · AI Hydra energy/power sensors via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy; peak/mean drawn from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 240.779 · mean 152.7 W · baseline 10.1 W (box-idle-all-empty) · delta 142.6 W
- covers: run-0109
- aihydra run-meta.jsonl entry: suite=sim-sensitivity, model=qwen35-122b, variant="F1". linked to run-0109 (ingested 2026-08-12 via scripts/ingest-tau2.mjs).
eng-0064
- 2026-08-11 · wall-meter · AI Hydra energy/power sensors via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy; peak/mean drawn from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 65.027 · mean 137.6 W · baseline 10.1 W (box-idle-all-empty) · delta 127.5 W
- covers: run-0110
- aihydra run-meta.jsonl entry: suite=sim-sensitivity, model=qwen35-122b, variant="F2". linked to run-0110 (ingested 2026-08-12 via scripts/ingest-tau2.mjs).
eng-0065
- 2026-08-14 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy) · window recorded
- wh_total 5.564 · mean 120.7 W · baseline 10.1 W (box-idle-all-empty) · delta 110.6 W
- covers: run-0245 run-0246
- aihydra run-meta.jsonl entries rocm-q8-d0-rep1/2/3 (qwen38-screen), suite qwen38-screen, cfg-0062, depth 0. 3 fresh-process llama-bench reps merged into one window (page cache dropped between reps, per run-0245/run-0246 comment). Retroactive join computed 2026-08-15 — the qwen38-27b bench published 2026-08-14 night with no energy join run against it; this closes that gap.
eng-0066
- 2026-08-14 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy) · window recorded
- wh_total 32.461 · mean 169.6 W · baseline 10.1 W (box-idle-all-empty) · delta 159.5 W
- covers: run-0247 run-0248
- aihydra run-meta.jsonl entries rocm-q8-d32768-rep1/2/3, suite qwen38-screen, cfg-0062, depth 32768. Retroactive join computed 2026-08-15.
eng-0067
- 2026-08-14 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy) · window recorded
- wh_total 170.252 · mean 177.7 W · baseline 10.1 W (box-idle-all-empty) · delta 167.6 W
- covers: run-0249 run-0250
- aihydra run-meta.jsonl entries rocm-q8-d131072-rep1/2/3, suite qwen38-screen, cfg-0062, depth 131072. Retroactive join computed 2026-08-15.
eng-0068
- 2026-08-14 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy) · window recorded
- wh_total 4.365 · mean 152.6 W · baseline 10.1 W (box-idle-all-empty) · delta 142.5 W
- covers: run-0251 run-0252
- aihydra run-meta.jsonl entries rocm-q4kxl-d0-rep1/2/3, suite qwen38-screen, cfg-0063, depth 0. Retroactive join computed 2026-08-15.
eng-0069
- 2026-08-14 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy) · window recorded
- wh_total 25.232 · mean 182 W · baseline 10.1 W (box-idle-all-empty) · delta 171.9 W
- covers: run-0253 run-0254
- aihydra run-meta.jsonl entries rocm-q4kxl-d32768-rep1/2/3, suite qwen38-screen, cfg-0063, depth 32768. Retroactive join computed 2026-08-15.
eng-0070
- 2026-08-14 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy) · window recorded
- wh_total 141.072 · mean 179.6 W · baseline 10.1 W (box-idle-all-empty) · delta 169.5 W
- covers: run-0255 run-0256
- aihydra run-meta.jsonl entries rocm-q4kxl-d131072-rep1/2/3, suite qwen38-screen, cfg-0063, depth 131072. Retroactive join computed 2026-08-15.
eng-0071
- 2026-08-14 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy) · window recorded
- wh_total 5.737 · mean 140.5 W · baseline 10.1 W (box-idle-all-empty) · delta 130.4 W
- covers: run-0257 run-0258
- aihydra run-meta.jsonl entries vulkan-q8-d0-rep1/2/3, suite qwen38-screen, cfg-0064, depth 0. Retroactive join computed 2026-08-15.
eng-0072
- 2026-08-14 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy) · window recorded
- wh_total 33.096 · mean 157.6 W · baseline 10.1 W (box-idle-all-empty) · delta 147.5 W
- covers: run-0259 run-0260
- aihydra run-meta.jsonl entries vulkan-q8-d32768-rep1/2/3, suite qwen38-screen, cfg-0064, depth 32768. Retroactive join computed 2026-08-15.
eng-0073
- 2026-08-14 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy) · window recorded
- wh_total 4.096 · mean 150.5 W · baseline 10.1 W (box-idle-all-empty) · delta 140.4 W
- covers: run-0261 run-0262
- aihydra run-meta.jsonl entries vulkan-q4kxl-d0-rep1/2/3, suite qwen38-screen, cfg-0065, depth 0. Retroactive join computed 2026-08-15.
eng-0074
- 2026-08-14 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy) · window recorded
- wh_total 28.497 · mean 175.7 W · baseline 10.1 W (box-idle-all-empty) · delta 165.6 W
- covers: run-0263 run-0264
- aihydra run-meta.jsonl entries vulkan-q4kxl-d32768-rep1/2/3, suite qwen38-screen, cfg-0065, depth 32768. Retroactive join computed 2026-08-15.
eng-0075
- 2026-08-14 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy) · window recorded
- wh_total 86.761 · mean 98.3 W · baseline 10.1 W (box-idle-all-empty) · delta 88.2 W
- covers: run-0265
- aihydra run-meta.jsonl entries vulkan-q8-d131072-rep1/2 (rc=134, device lost both reps — see run-0265 metrics.device_lost). Mean power (98.3 W, delta 88.2 W) sits well below the working cells at this depth (cfg-0062/0063 d131072 ran 168-180 W delta) — consistent with the GPU hanging partway through rather than completing useful decode for the full window; the failure itself is part of what this energy figure measures. Retroactive join computed 2026-08-15.
eng-0076
- 2026-08-14 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy) · window recorded
- wh_total 86.366 · mean 101.2 W · baseline 10.1 W (box-idle-all-empty) · delta 91.1 W
- covers: run-0266
- aihydra run-meta.jsonl entries vulkan-q4kxl-d131072-rep1/2 (rc=134, device lost both reps — see run-0266 metrics.device_lost). Same device-lost pattern as eng-0075: mean power (101.2 W) well below the working d131072 cells, consistent with the hang. Retroactive join computed 2026-08-15.
eng-0077
- 2026-08-15 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy) · window recorded
- wh_total 37.087 · mean 183.4 W · baseline 10.1 W (box-idle-all-empty) · delta 173.3 W
- covers: run-0267
- aihydra run-meta.jsonl entry qwen38-screen/tau2-full-cliffguard-n3 (second occurrence, 2026-08-15T01:59:23Z rc=0 — the first occurrence at 01:44:21Z was a discarded rehearsal, not this run). 25 sequential full-length generations at cfg-0066 (n_max=3), the EOS-cliff guard for the tau2 full arm (run-0271). Retroactive join computed 2026-08-15.
eng-0078
- 2026-08-15 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy) · window recorded
- wh_total 62.054 · mean 157.2 W · baseline 10.1 W (box-idle-all-empty) · delta 147.1 W
- covers: run-0268
- aihydra run-meta.jsonl entries mtp-rocm-Q8_0-nmaxoff-c32768 (01:15:33-01:31:48) + mtp-rocm-Q8_0-nmax3-c32768 (01:31:50-01:39:14), suite qwen38-screen, cfg-0067. Window spans both arms of run-0268 (floor + n_max=3, ROCm). See clm-0055 for the per-n_max Wh/1000-tokens breakdown computed from these same sub-windows. Retroactive join computed 2026-08-15.
eng-0079
- 2026-08-14 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy) · window recorded
- wh_total 81.261 · mean 147.2 W · baseline 10.1 W (box-idle-all-empty) · delta 137.1 W
- covers: run-0269
- aihydra run-meta.jsonl entries mtp-vulkan-Q8_0-nmaxoff-c32768 (18:27:11-18:43:44) + mtp-vulkan-Q8_0-nmax2-c32768 (18:43:47-18:52:12) + mtp-vulkan-Q8_0-nmax3-c32768 (18:52:15-19:00:18), suite qwen38-screen, cfg-0068 — the first sweep pass, matching run-0269s stated span floor..n3. The n_max>=4 cells that follow in the raw sweep are EXCLUDED from this window: they are the degenerate EOS-cliff cells (clm-0055) and are not part of the published run. See clm-0055 for the per-n_max Wh/1000-tokens breakdown computed from the nmaxoff/nmax3 sub-windows. Retroactive join computed 2026-08-15.
eng-0080
- 2026-08-14 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy) · window recorded
- wh_total 58.346 · mean 160.6 W · baseline 10.1 W (box-idle-all-empty) · delta 150.5 W
- covers: run-0270
- aihydra run-meta.jsonl entries mtp-vulkan-UD-Q4_K_XL-nmaxoff-c32768 (19:08:50-19:19:52) + mtp-vulkan-UD-Q4_K_XL-nmax2-c32768 (19:19:54-19:25:35) + mtp-vulkan-UD-Q4_K_XL-nmax3-c32768 (19:25:37-19:30:38), suite qwen38-screen, cfg-0069 — the first sweep pass, matching run-0270s stated span floor..n3. The n_max>=4 cells are EXCLUDED (degenerate, clm-0055). See clm-0055 for the per-n_max Wh/1000-tokens breakdown computed from the nmaxoff/nmax3 sub-windows. Retroactive join computed 2026-08-15.
eng-0081
- 2026-08-15 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy) · window recorded
- wh_total 574.37 · mean 181.4 W · baseline 10.1 W (box-idle-all-empty) · delta 171.3 W
- per task: 38.2913 Wh · 1.1602 p (tariff 30.3 p/kWh)
- covers: run-0271
- aihydra run-meta.jsonl entry qwen38-screen/tau2-full (rc=124 — the 11400s wall bound cut the run at 26 of the requested 50 tasks; run-0271 records the 26 completed and scored). wh_per_task is Wh-per-CORRECT-ANSWER: wh_total / tasks_passed (15 of 26, run-0271.metrics). This is WHOLE-SESSION energy, not agent-only — the window includes simulator (user_llm) wait time and a ~19-minute gap between task 9 and task 10 (see run-0271s per-task log), so it should not be read as pure inference cost. Retroactive join computed 2026-08-15, the day after the bench published without its energy figure.
eng-0082
- 2026-08-15 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy) · window recorded
- wh_total 163.696 · mean 137 W · baseline 10.1 W (box-idle-all-empty) · delta 126.9 W
- per task: 14.8814 Wh · 0.4509 p (tariff 30.3 p/kWh)
- covers: run-0272
- aihydra run-meta.jsonl entry kv-quality-patched/f16 (2026-08-15T05:22:15Z, rc=124 — the hard-split design caps this arm at half the total time budget by construction, per bench/queue-kv-quality-patched.sh; the cut is the intended stopping point, not a failure). wh_per_task is Wh-per-CORRECT-ANSWER: wh_total / tasks_passed (11 of 14 scored, run-0272.metrics; 2 of the 16 attempted sims hit infrastructure_error and are excluded from tasks_total per run-0272). Paired against eng-0083 (q8_0, same 14 task-matched ids): see clm-0058 for the quality comparison and the Wh-per-correct reading together. Retroactive join computed 2026-08-15.
eng-0083
- 2026-08-15 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy) · window recorded
- wh_total 103.874 · mean 135.7 W · baseline 10.1 W (box-idle-all-empty) · delta 125.6 W
- per task: 8.6562 Wh · 0.2623 p (tariff 30.3 p/kWh)
- covers: run-0273
- aihydra run-meta.jsonl entry kv-quality-patched/q8_0 (2026-08-15T06:34:00Z, rc=0). wh_per_task is Wh-per-CORRECT-ANSWER: wh_total / tasks_passed (12 of 14, run-0273.metrics). Task-matched against eng-0082s f16 arm — identical 14 task ids, same seed. Both wall time (2755s vs 4303s) and energy (103.874 Wh vs 163.696 Wh) are well below the f16 arm despite scoring HIGHER (12/14 vs 11/14): q8_0 KV is not merely quality-neutral on the patched build (clm-0058) — it is also the cheaper arm to run, on both counts, over the identical task set. See clm-0058 for the full comparison. Retroactive join computed 2026-08-15.
eng-0084
- 2026-08-15 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy) · window recorded
- wh_total 76.0596 · mean 157.63 W · baseline 10.1 W (box-idle-all-empty) · delta 147.53 W
- per task: 19.0149 Wh · 0.5761 p (tariff 30.3 p/kWh)
- covers: run-0274
- Nemotron-3-Super-120B-A12B Q4_K_M SMOKE arm (run-0274, cfg-0072), the fair-quant re-screen this candidate had been waiting on since 2026-08-09. wh_per_task is Wh-per-CORRECT-ANSWER: wh_total / tasks_passed (4 of 5, run-0274.metrics). Counter samples nearest the window bounds: start 19.1516851633787076 kWh (2026-08-15T10:48:09Z, 3s before window open — nearest recorded tick), end 19.2277447432279476 kWh (2026-08-15T11:17:09Z, exact). At 157.6 W mean active draw this is a notably heavier arm than the 35B-class models in the corpus — consistent with Nemotron-3-Super's own throughput being the slowest measured on this box (clm-0035). Retroactive join computed 2026-08-15, same day as the run.
eng-0085
- 2026-08-15 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy) · window recorded
- wh_total 19.2072 · mean 193.15 W · baseline 10.1 W (box-idle-all-empty) · delta 183.05 W
- covers: run-0275
- EOS-cliff x q8_0-KV probe, n_max=3 (run-0275, cfg-0073) — a cliff-detector guard run (15 generation requests, 0 degenerate), not a reward-scored capability run, so wh_per_task/pence_per_correct_answer do not apply here (left null rather than forced into the wrong denominator); Wh/request = 19.2072/15 = 1.28. Counter samples: start 19.2282059639692276 kWh (2026-08-15T11:19:19Z, 2s before window open — nearest recorded tick), end 19.2474132031202276 kWh (2026-08-15T11:25:19Z, exact). Mean active draw 193.2 W is the highest of the three joins in this batch — draft-mtp speculative decode at n_max=3 draws more power than plain serving, consistent with clm-0055's energy section (speculation's power draw rises, not flat). Retroactive join computed 2026-08-15, same day as the run.
eng-0086
- 2026-08-15 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy) · window recorded
- wh_total 27.5745 · mean 183.49 W · baseline 10.1 W (box-idle-all-empty) · delta 173.39 W
- covers: run-0276
- EOS-cliff x q8_0-KV probe, n_max=7 (run-0276, cfg-0074) — a cliff-detector guard run (15 generation requests, 0 degenerate), not a reward-scored capability run; wh_per_task/pence_per_correct_answer left null for the same reason as eng-0085. Wh/request = 27.5745/15 = 1.84 — higher than n_max=3's 1.28 Wh/request (eng-0085), consistent with a larger draft window costing more compute per accepted step. Counter samples: start 19.2474132031202276 kWh (2026-08-15T11:25:19Z, exact — shared boundary with eng-0085's end, the two runs are back-to-back on one server session), end 19.2749877423047976 kWh (2026-08-15T11:34:19Z, 1s before window close — nearest recorded tick). Retroactive join computed 2026-08-15, same day as the run.
eng-0087
- 2026-08-15 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy) · window recorded
- wh_total 6.177 · mean 160 W · baseline 10.1 W (box-idle-all-empty) · delta 149.9 W
- covers: run-0277
- EOS-cliff cause-isolation matrix, cell 0 (run-0277, cfg-0075) — a cliff-detector guard run (15 generation requests, 11 degenerate), not a reward-scored capability run, so wh_per_task/pence_per_correct_answer do not apply (left null). Counter samples: start 19.3267277926206576 kWh (13:14:49 local / 12:14:49Z, exact), end 19.3329050987958876 kWh (13:17:09 local, 1s after window close — nearest recorded tick). Retroactive join computed 2026-08-15, same session as the run.
eng-0088
- 2026-08-15 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy) · window recorded
- wh_total 22.012 · mean 184.3 W · baseline 10.1 W (box-idle-all-empty) · delta 174.2 W
- covers: run-0278
- EOS-cliff cause-isolation matrix, cell B / backend-flip-only (run-0278, cfg-0076) — a cliff-detector guard run (15 generation requests, 0 degenerate), not a reward-scored capability run; wh_per_task/ pence_per_correct_answer left null for the same reason as eng-0087. ROCm draws visibly more power than the Vulkan baseline cell over the identical workload shape (184.3 W mean vs eng-0087's 160.0 W) — consistent with clm-0050's per-model/per-phase backend note, not itself evidence about the cliff mechanism. Counter samples: start 19.3329050987958876 kWh (shared boundary with eng-0087's end), end 19.3549172133207276 kWh (13:24:19 local, 1s after window close). Retroactive join computed 2026-08-15.
eng-0089
- 2026-08-15 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy) · window recorded
- wh_total 21.991 · mean 180.3 W · baseline 10.1 W (box-idle-all-empty) · delta 170.2 W
- covers: run-0279
- EOS-cliff cause-isolation matrix, cell C / KV-flip-only (run-0279, cfg-0077) — a cliff-detector guard run (15 generation requests, 1 short early-stop at 18 tokens, 14 counted as non-degenerate); wh_per_task/ pence_per_correct_answer left null. Counter samples: start 19.3549172133207276 kWh (shared boundary with eng-0088's end), end 19.3769087046384776 kWh (13:31:39 local, 2s after window close). Retroactive join computed 2026-08-15.
eng-0090
- 2026-08-15 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy) · window recorded
- wh_total 21.469 · mean 181.9 W · baseline 10.1 W (box-idle-all-empty) · delta 171.8 W
- covers: run-0280
- EOS-cliff cause-isolation matrix, cell E / new-evidence combo at n_max=5 (run-0280, cfg-0078) — a cliff-detector guard run (15 generation requests, 3 short early-stops at 271/267/267 tokens, 12 counted as non-degenerate); wh_per_task/pence_per_correct_answer left null. Counter samples: start 19.3769087046384776 kWh (shared boundary with eng-0089's end), end 19.3983780592679876 kWh (13:38:39 local, 3s before window close — nearest recorded tick). Retroactive join computed 2026-08-15.
eng-0091
- 2026-08-15 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy) · window recorded
- wh_total 65.206 · mean 185.1 W · baseline 10.1 W (box-idle-all-empty) · delta 175 W
- covers: run-0281
- EOS-cliff cause-isolation matrix, extended-volume confirmation of cell B (run-0281, cfg-0076, 45 requests, 0 degenerate) — a cliff-detector guard run, not a reward-scored capability run; wh_per_task/pence_per_correct_answer left null. Wh/request = 65.206/45 = 1.45, close to eng-0088's 15-request cell B rate (22.012/15 = 1.47) — power draw stayed flat across 3x the session volume, consistent with the clean result not being a transient. Counter samples: start 19.3990371674299176 kWh (13:40:59 local, 1s after window open), end 19.4642435759305876 kWh (14:02:09 local, 3s after window close). Retroactive join computed 2026-08-15.
eng-0092
- 2026-08-15 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy) · window recorded
- wh_total 82.561 · mean 184.3 W · baseline 10.1 W (box-idle-all-empty) · delta 174.2 W
- covers: run-0282
- EOS-cliff cause-isolation matrix, extended-volume confirmation of run-0276 (run-0282, cfg-0081, 45 requests, 0 strict-degenerate, periodic early-stops — see run note); wh_per_task/pence_per_correct_answer left null (guard run). Wh/request = 82.561/45 = 1.83, close to eng-0086's 15-request n_max=7 rate (27.5745/15 = 1.84) — power draw held flat at 3x the session volume. Counter samples: start 19.4642435759305876 kWh (shared boundary with eng-0091's end), end 19.5468044728040676 kWh (14:28:59 local, exact). Retroactive join computed 2026-08-15.
eng-0093
- 2026-08-15 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy) · window recorded
- wh_total 21.194 · mean 161 W · baseline 10.1 W (box-idle-all-empty) · delta 150.9 W
- covers: run-0283
- EOS-cliff cause-isolation matrix, avoidance probe 1 / n_max=3 on the failing baseline config (run-0283, cfg-0079, 15 requests, 0 degenerate); wh_per_task/pence_per_correct_answer left null (guard run). Counter samples: start 19.5468044728040676 kWh (shared boundary with eng-0092's end), end 19.5679985731840076 kWh (14:36:49 local, 4s before window close — nearest recorded tick). Retroactive join computed 2026-08-15.
eng-0094
- 2026-08-15 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy) · window recorded
- wh_total 26.557 · mean 177.1 W · baseline 10.1 W (box-idle-all-empty) · delta 167 W
- covers: run-0284
- EOS-cliff cause-isolation matrix, avoidance probe 2 / spec-draft-p-min 0.9 on the failing baseline config (run-0284, cfg-0080, 15 requests, 1 degenerate); wh_per_task/pence_per_correct_answer left null (guard run). Counter samples: start 19.5679985731840076 kWh (shared boundary with eng-0093's end), end 19.5945554226636776 kWh (14:45:49 local, 4s before window close — nearest recorded tick). Retroactive join computed 2026-08-15.
eng-0095
- 2026-08-15 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy) · window recorded
- wh_total 5.255 · mean 176.8 W · baseline 10.1 W (box-idle-all-empty) · delta 166.7 W
- covers: run-0285
- EOS-cliff cause-isolation matrix, avoidance probe 3 / cache_prompt:true on the failing baseline config (run-0285, cfg-0075, 15 requests, 12 degenerate — WORSE than the same config's own 11/15 without cache_prompt, run-0277); wh_per_task/pence_per_correct_answer left null (guard run). The short 107 s window reflects how quickly this cell collapses into 1-token generations. Counter samples: start 19.5945554226636776 kWh (shared boundary with eng-0094's end), end 19.5998108834028176 kWh (14:47:39 local, 1s before window close — nearest recorded tick). Retroactive join computed 2026-08-15.
eng-0096
- 2026-08-15 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy) · window recorded
- wh_total 22.72 · mean 172.6 W · baseline 10.1 W (box-idle-all-empty) · delta 162.5 W
- covers: run-0286
- EOS-cliff cause-isolation matrix, avoidance probe 4 / server restarted every 3 requests on the failing baseline config (run-0286, cfg-0075, 15 requests across 5 fresh server processes, 1 degenerate — down sharply from 11/15 continuous); wh_per_task/pence_per_correct_answer left null (guard run). This window includes 5 model-load overheads (~9-13s each observed in the isolation log) inside the wall-meter total, so it is not directly comparable per-request to the continuous-session cells (run-0277, run-0283) despite similar total Wh — reload overhead is part of what restart-as-mitigation costs operationally. Counter samples: start 19.5998108834028176 kWh (shared boundary with eng-0095's end), end 19.6225305050611476 kWh (14:55:39 local, 5s after window close — nearest recorded tick). Retroactive join computed 2026-08-15.
eng-0097
- 2026-08-15 · wall-meter · aihydra energy sensor via Home Assistant (5-minute statistics on sensor.hardware_ai_hydra_energy, linearly interpolated to the run window) · window recorded
- wh_total 10.99 · mean 144.4 W · baseline 10.1 W (box-idle-all-empty) · delta 134.3 W
- covers: run-0288 run-0289
- aihydra deepseek-v4-flash-fullbench throughput matrix cell. Wall-metered via HA sensor.hardware_ai_hydra_energy, joined from the permanent 5-minute statistics series (linearly interpolated to the exact run-meta window boundaries), not raw 10s history -- the matrix's ~4h span exceeded the history query's row cap by the time of ingest. Baseline 10.1 W is the repo's dated idle-aihydra-2026-08 box-idle-all-empty record.
eng-0098
- 2026-08-15 · wall-meter · aihydra energy sensor via Home Assistant (5-minute statistics on sensor.hardware_ai_hydra_energy, linearly interpolated to the run window) · window recorded
- wh_total 59.51 · mean 166.7 W · baseline 10.1 W (box-idle-all-empty) · delta 156.6 W
- covers: run-0290 run-0291
- aihydra deepseek-v4-flash-fullbench throughput matrix cell. Wall-metered via HA sensor.hardware_ai_hydra_energy, joined from the permanent 5-minute statistics series (linearly interpolated to the exact run-meta window boundaries), not raw 10s history -- the matrix's ~4h span exceeded the history query's row cap by the time of ingest. Baseline 10.1 W is the repo's dated idle-aihydra-2026-08 box-idle-all-empty record.
eng-0099
- 2026-08-15 · wall-meter · aihydra energy sensor via Home Assistant (5-minute statistics on sensor.hardware_ai_hydra_energy, linearly interpolated to the run window) · window recorded
- wh_total 43.08 · mean 171.6 W · baseline 10.1 W (box-idle-all-empty) · delta 161.5 W
- covers: run-0292 run-0293
- aihydra deepseek-v4-flash-fullbench throughput matrix cell. Wall-metered via HA sensor.hardware_ai_hydra_energy, joined from the permanent 5-minute statistics series (linearly interpolated to the exact run-meta window boundaries), not raw 10s history -- the matrix's ~4h span exceeded the history query's row cap by the time of ingest. Baseline 10.1 W is the repo's dated idle-aihydra-2026-08 box-idle-all-empty record.
eng-0100
- 2026-08-15 · wall-meter · aihydra energy sensor via Home Assistant (5-minute statistics on sensor.hardware_ai_hydra_energy, linearly interpolated to the run window) · window recorded
- wh_total 117.48 · mean 177.9 W · baseline 10.1 W (box-idle-all-empty) · delta 167.8 W
- covers: run-0294 run-0295
- aihydra deepseek-v4-flash-fullbench throughput matrix cell. Wall-metered via HA sensor.hardware_ai_hydra_energy, joined from the permanent 5-minute statistics series (linearly interpolated to the exact run-meta window boundaries), not raw 10s history -- the matrix's ~4h span exceeded the history query's row cap by the time of ingest. Baseline 10.1 W is the repo's dated idle-aihydra-2026-08 box-idle-all-empty record.
eng-0101
- 2026-08-15 · wall-meter · aihydra energy sensor via Home Assistant (5-minute statistics on sensor.hardware_ai_hydra_energy, linearly interpolated to the run window) · window recorded
- wh_total 349.87 · mean 175.3 W · baseline 10.1 W (box-idle-all-empty) · delta 165.2 W
- covers: run-0296 run-0297
- aihydra deepseek-v4-flash-fullbench throughput matrix cell. Wall-metered via HA sensor.hardware_ai_hydra_energy, joined from the permanent 5-minute statistics series (linearly interpolated to the exact run-meta window boundaries), not raw 10s history -- the matrix's ~4h span exceeded the history query's row cap by the time of ingest. Baseline 10.1 W is the repo's dated idle-aihydra-2026-08 box-idle-all-empty record.
eng-0102
- 2026-08-15 · wall-meter · aihydra energy sensor via Home Assistant (5-minute statistics on sensor.hardware_ai_hydra_energy, linearly interpolated to the run window) · window recorded
- wh_total 7.4 · mean 87 W · baseline 10.1 W (box-idle-all-empty) · delta 76.9 W
- covers: run-0298 run-0299
- aihydra deepseek-v4-flash-fullbench throughput matrix cell. Wall-metered via HA sensor.hardware_ai_hydra_energy, joined from the permanent 5-minute statistics series (linearly interpolated to the exact run-meta window boundaries), not raw 10s history -- the matrix's ~4h span exceeded the history query's row cap by the time of ingest. Baseline 10.1 W is the repo's dated idle-aihydra-2026-08 box-idle-all-empty record.
eng-0103
- 2026-08-15 · wall-meter · aihydra energy sensor via Home Assistant (5-minute statistics on sensor.hardware_ai_hydra_energy, linearly interpolated to the run window) · window recorded
- wh_total 30.53 · mean 91.2 W · baseline 10.1 W (box-idle-all-empty) · delta 81.1 W
- covers: run-0300
- aihydra deepseek-v4-flash-fullbench throughput matrix cell. Wall-metered via HA sensor.hardware_ai_hydra_energy, joined from the permanent 5-minute statistics series (linearly interpolated to the exact run-meta window boundaries), not raw 10s history -- the matrix's ~4h span exceeded the history query's row cap by the time of ingest. Baseline 10.1 W is the repo's dated idle-aihydra-2026-08 box-idle-all-empty record. Window covers 3 rep(s) x 2 device-loss attempts each (6 crash+kernel-ring-reset cycles), not steady compute — this is NOT a throughput-comparable power figure.
eng-0104
- 2026-08-15 · wall-meter · aihydra energy sensor via Home Assistant (5-minute statistics on sensor.hardware_ai_hydra_energy, linearly interpolated to the run window) · window recorded
- wh_total 10.45 · mean 95.3 W · baseline 10.1 W (box-idle-all-empty) · delta 85.2 W
- covers: run-0301
- aihydra deepseek-v4-flash-fullbench throughput matrix cell. Wall-metered via HA sensor.hardware_ai_hydra_energy, joined from the permanent 5-minute statistics series (linearly interpolated to the exact run-meta window boundaries), not raw 10s history -- the matrix's ~4h span exceeded the history query's row cap by the time of ingest. Baseline 10.1 W is the repo's dated idle-aihydra-2026-08 box-idle-all-empty record. Window covers 1 rep(s) x 2 device-loss attempts each (2 crash+kernel-ring-reset cycles), not steady compute — this is NOT a throughput-comparable power figure.
eng-0105
- 2026-08-15 · wall-meter · aihydra energy sensor via Home Assistant (5-minute statistics on sensor.hardware_ai_hydra_energy, linearly interpolated to the run window) · window recorded
- wh_total 8.72 · mean 81.5 W · baseline 10.1 W (box-idle-all-empty) · delta 71.4 W
- covers: run-0302
- aihydra deepseek-v4-flash-fullbench throughput matrix cell. Wall-metered via HA sensor.hardware_ai_hydra_energy, joined from the permanent 5-minute statistics series (linearly interpolated to the exact run-meta window boundaries), not raw 10s history -- the matrix's ~4h span exceeded the history query's row cap by the time of ingest. Baseline 10.1 W is the repo's dated idle-aihydra-2026-08 box-idle-all-empty record. Window covers 1 rep(s) x 2 device-loss attempts each (2 crash+kernel-ring-reset cycles), not steady compute — this is NOT a throughput-comparable power figure.
eng-0106
- 2026-08-15 · wall-meter · aihydra energy sensor via Home Assistant (5-minute statistics on sensor.hardware_ai_hydra_energy, linearly interpolated to the run window) · window recorded
- wh_total 4.54 · mean 38.7 W · baseline 10.1 W (box-idle-all-empty) · delta 28.6 W
- covers: run-0303
- aihydra deepseek-v4-flash-fullbench throughput matrix cell. Wall-metered via HA sensor.hardware_ai_hydra_energy, joined from the permanent 5-minute statistics series (linearly interpolated to the exact run-meta window boundaries), not raw 10s history -- the matrix's ~4h span exceeded the history query's row cap by the time of ingest. Baseline 10.1 W is the repo's dated idle-aihydra-2026-08 box-idle-all-empty record. Window covers 1 rep(s) x 2 device-loss attempts each (2 crash+kernel-ring-reset cycles), not steady compute — this is NOT a throughput-comparable power figure.
eng-0107
- 2026-08-15 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, 5-minute statistics interpolated to the run-meta window) · window recorded
- wh_total 534.28 · mean 171.9 W · baseline 10.1 W (box-idle-all-empty) · delta 161.8 W
- per task: 24.2853 Wh · 0.7358 p (tariff 30.3 p/kWh)
- covers: run-0287
- aihydra run-meta.jsonl entry deepseek-v4-flash-fullbench/tau2-full-26task (rc=0 — the run completed all 26 requested tasks cleanly, NOT a wall-bound cut; contrast qwen38-27b's eng-0081, which was cut at rc=124). wh_per_task is Wh-per-CORRECT-ANSWER: wh_total / tasks_passed (22 of 26, run-0287.metrics). This is WHOLE-SESSION energy, not agent-only — the window includes simulator (user_llm) wait time end to end. Joined same-session (not retroactive): HA statistics pulled directly against the exact run-meta window immediately after the run completed.
eng-0108
- 2026-08-16 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy) · window recorded
- wh_total 9.8273 · mean 147.41 W · baseline 10.1 W (box-idle-all-empty) · delta 137.31 W
- covers: run-0304
- ITEM A of the 2026-08-16 bounded-box filing-gap job (eos-cliff isolation matrix quant-only flip, run-0304, cfg-0085) — a cliff-detector guard run (15 generation requests, 7 degenerate), not a reward-scored capability run, so wh_per_task/pence_per_correct_answer do not apply (left null). Counter samples read live via the Home Assistant MCP at each boundary, not retroactively: start 20.8945578187704056 kWh (02:26:29 local / 01:26:29Z), end 20.9043851345777486 kWh (02:30:29 local / 01:30:29Z). Window is the boundary-read span (240s), slightly wider than the script's own run-meta window (184s) by the polling gap either side.
eng-0109
- 2026-08-16 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy) · window recorded
- wh_total 6.4163 · mean 144.37 W · baseline 10.1 W (box-idle-all-empty) · delta 134.27 W
- covers: run-0305
- ITEM B of the 2026-08-16 bounded-box filing-gap job (still-present-on-master check, run-0305, cfg-0086, llama.cpp master ece963f) — a cliff-detector guard run (15 generation requests, 11 degenerate, bit-identical sequence to run-0277's 3653e6d result), not a reward-scored capability run, so wh_per_task/pence_per_correct_answer do not apply (left null). Counter samples read live via the Home Assistant MCP at each boundary: start 20.9082643836736656 kWh (02:35:49 local / 01:35:49Z), end 20.9146807044744466 kWh (02:38:29 local / 01:38:29Z). Window is the boundary-read span (160s), close to the script's own run-meta window (145s).
eng-0110
- 2026-08-16 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy) · window recorded
- wh_total 1.8744 · mean 96.4 W · baseline 10.1 W (box-idle-all-empty) · delta 86.3 W
- covers: run-0306
- ITEM C of the 2026-08-16 bounded-box filing-gap job (house 4-item capability guard backfill against cfg-0066, run-0306) — a guard run (4 short probes: coherence/toolcall/needle/isolation), not a reward-scored capability run, so wh_per_task/pence_per_correct_answer do not apply (left null). Counter samples read live via the Home Assistant MCP at each boundary: start 20.9148347228765456 kWh (02:39:09 local / 01:39:09Z), end 20.9167091697454426 kWh (02:40:19 local / 01:40:19Z). Window is the boundary-read span (70s), close to the script's own run-meta window (36-76s across the server-start-plus-guard span).
eng-0111
- 2026-08-16 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s history samples interpolated to the run-meta window) · window recorded
- wh_total 1.6578 · mean 64.87 W · baseline 10.1 W (box-idle-all-empty) · delta 54.77 W
- covers: run-0307
- A guard run (4 short probes: coherence/toolcall/needle/isolation), not a reward-scored capability run, so wh_per_task/pence_per_correct_answer do not apply (left null). Counter interpolated from raw ~10s-resolution HA history samples (source=history, not the 5-minute statistics bucket) since the window (92s) is shorter than statistics granularity: start 20.9887294913 kWh, end 20.9903873286 kWh (linear interpolation between bracketing raw samples at each boundary). Window spans the server-load-plus-guard span (matches run-0307's own started_at/ended_at), same convention as eng-0110 (qwen38-27b's guard backfill).
eng-0112
- 2026-08-16 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s history samples interpolated to the run-meta window) · window recorded
- wh_total 4.5594 · mean 72.31 W · baseline 10.1 W (box-idle-all-empty) · delta 62.21 W
- covers: run-0308 run-0309
- aihydra nemotron3-super-fullbench-matrix cell rocm-d0 (3 fresh-process reps, pp1024/tg256). Counter interpolated from raw ~10s-resolution HA history samples: start 20.9203543295 kWh, end 20.9249136998 kWh. Window includes model load (this candidate's ~78.7 GiB weights, cold each rep) plus the brief llama-bench workload at d0, hence the modest mean power vs the d32768 cells.
eng-0113
- 2026-08-16 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s history samples interpolated to the run-meta window) · window recorded
- wh_total 27.6194 · mean 147.74 W · baseline 10.1 W (box-idle-all-empty) · delta 137.64 W
- covers: run-0310 run-0311
- aihydra nemotron3-super-fullbench-matrix cell rocm-d32768 (3 fresh-process reps, pp1024/tg256, prefill to depth 32768 each rep). Counter interpolated from raw ~10s-resolution HA history samples: start 20.9249136998 kWh, end 20.9525331191 kWh. This is the depth that OOM'd under mmap pre-fix (perf-matrix-gtt120, 2026-08-14) — the window here reflects three clean completions under --load-mode none.
eng-0114
- 2026-08-16 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s history samples interpolated to the run-meta window) · window recorded
- wh_total 5.1557 · mean 64 W · baseline 10.1 W (box-idle-all-empty) · delta 53.9 W
- covers: run-0312 run-0313
- aihydra nemotron3-super-fullbench-matrix cell vulkan-d0 (3 fresh-process reps, pp1024/tg256). Counter interpolated from raw ~10s-resolution HA history samples: start 20.9525663038 kWh, end 20.9577220116 kWh.
eng-0115
- 2026-08-16 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s history samples interpolated to the run-meta window) · window recorded
- wh_total 30.5166 · mean 149.47 W · baseline 10.1 W (box-idle-all-empty) · delta 139.37 W
- covers: run-0314 run-0315
- aihydra nemotron3-super-fullbench-matrix cell vulkan-d32768 (3 fresh-process reps, pp1024/tg256, prefill to depth 32768 each rep, zero device-loss). Counter interpolated from raw ~10s-resolution HA history samples: start 20.9577220116 kWh, end 20.9882386300 kWh. Highest mean power of the four matrix cells — the longest-running, deepest, most compute-dense arm.
eng-0116
- 2026-08-16 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s history samples interpolated to the run-meta window) · window recorded
- wh_total 648.6 · mean 146.23 W · baseline 10.1 W (box-idle-all-empty) · delta 136.13 W
- per task: 32.43 Wh · 0.9826 p (tariff 30.3 p/kWh)
- covers: run-0316
- aihydra run-meta.jsonl entry nemotron3-super-fullbench/tau2-full-26task (rc=0 — all 26 requested tasks completed cleanly, not a wall-bound cut). Joined same-session: counter interpolated from raw ~10s-resolution HA history samples at each boundary — start 20.99128201 kWh (03:32:56+01:00 local / 02:32:56Z), end 21.63988203 kWh (07:59:04+01:00 local / 06:59:04Z). wh_per_task is Wh-per-CORRECT-ANSWER: wh_total / tasks_passed (20 of 26, run-0316.metrics). WHOLE-SESSION energy, not agent-only — the window includes simulator (user_llm) wait time end to end, directly comparable on that denominator to qwen38-27b's eng-0081 and deepseek-v4-flash's eng-0107. Mean power (146.23 W) sits between the throughput matrix's rocm-d32768 cell (147.74 W, eng-0113) and vulkan-d32768 cell (149.47 W, eng-0115) — consistent with a mostly-decode-bound agentic workload at the same serving depth band.
eng-0117
- 2026-08-16 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the run-meta window) · window recorded
- wh_total 5.5944 · mean 40.77 W · baseline 10.1 W (box-idle-all-empty) · delta 30.67 W
- covers: run-0317 run-0318
- aihydra laguna-s21-fullbench throughput matrix cell rocm-d0 (3 fresh-process reps, pp1024/tg256). Counter interpolated from raw ~10s-resolution HA history samples pulled fresh this session: start 21.644490603467226 kWh, end 21.650084960847323 kWh. Recomputed independently from HA raw history (not copied from the delegate's dead-drop report) per house policy against trusting prose over raw data; this window's figure matches the MODEL-PAGE.md draft to within rounding (5.59 Wh / 40.8 W).
eng-0118
- 2026-08-16 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the run-meta window) · window recorded
- wh_total 25.6989 · mean 114.5 W · baseline 10.1 W (box-idle-all-empty) · delta 104.4 W
- covers: run-0319 run-0320
- aihydra laguna-s21-fullbench throughput matrix cell rocm-d32768 (3 fresh-process reps, pp1024/tg256 at depth 32768). Counter interpolated from raw ~10s-resolution HA history samples pulled fresh this session: start 21.650084960847323 kWh, end 21.675783821154234 kWh. Recomputed independently from HA raw history; matches the MODEL-PAGE.md draft to within rounding (25.70 Wh / 114.5 W).
eng-0119
- 2026-08-16 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the run-meta window) · window recorded
- wh_total 41.3597 · mean 165.44 W · baseline 10.1 W (box-idle-all-empty) · delta 155.34 W
- covers: run-0321 run-0322
- aihydra laguna-s21-fullbench throughput matrix cell rocm-d131072 (N=1 rep, pp1024/tg256 at depth 131072 -- the deepest cell either backend completed this job). Counter interpolated from raw ~10s-resolution HA history samples pulled fresh this session: start 21.675783821154234 kWh, end 21.717143484197067 kWh. Recomputed independently from HA raw history; matches the MODEL-PAGE.md draft to within rounding (41.36 Wh / 165.4 W). Highest mean power in the matrix -- this cell's 900s window is dominated by prefill at depth on a 118B/~8B-active hybrid model.
eng-0120
- 2026-08-16 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the run-meta window) · window recorded
- wh_total 5.4354 · mean 53.9 W · baseline 10.1 W (box-idle-all-empty) · delta 43.8 W
- covers: run-0323 run-0324
- aihydra laguna-s21-fullbench throughput matrix cell vulkan-d0 (3 fresh-process reps, pp1024/tg256). Counter interpolated from raw ~10s-resolution HA history samples pulled fresh this session: start 21.71715612107426 kWh, end 21.72259148846168 kWh. Recomputed independently from HA raw history; matches the MODEL-PAGE.md draft to within rounding (5.44 Wh / 53.9 W).
eng-0121
- 2026-08-16 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the run-meta window) · window recorded
- wh_total 35.4253 · mean 101.05 W · baseline 10.1 W (box-idle-all-empty) · delta 90.95 W
- covers: run-0325 run-0326
- aihydra laguna-s21-fullbench throughput matrix cell vulkan-d32768 (3 fresh-process reps, pp1024/tg256 at depth 32768). Counter interpolated from raw ~10s-resolution HA history samples pulled fresh this session: start 21.722610139605965 kWh, end 21.75803541361878 kWh. Recomputed independently from HA raw history; matches the MODEL-PAGE.md draft to within rounding (35.43 Wh / 101.1 W).
eng-0122
- 2026-08-16 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the run-meta window) · window recorded
- wh_total 48.5675 · mean 91.25 W · baseline 10.1 W (box-idle-all-empty) · delta 81.15 W
- covers: run-0327
- aihydra laguna-s21-fullbench throughput matrix cell vulkan-d131072 -- CELL FAILED, both allowed attempts lost the GPU device (rc=134, vk::DeviceLostError; run-0327). This is WASTED energy, not a throughput result: 48.57 Wh spent across two attempts (1035s + 881s) that each aborted before producing a decode/prefill number. Counter interpolated from raw ~10s-resolution HA history samples pulled fresh this session: start 21.75803541361878 kWh, end 21.80660287926063 kWh. Recomputed independently from HA raw history; matches the MODEL-PAGE.md draft to within rounding (48.57 Wh / 91.3 W). Recorded per protocol's "publish the frontier, not the winner" -- a failed cell carries its own energy cost and that cost is part of the record.
eng-0123
- 2026-08-16 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the run-meta window) · window recorded
- wh_total 0.8021 · mean 137.51 W · baseline 10.1 W (box-idle-all-empty) · delta 127.41 W
- covers: run-0328
- aihydra laguna-s21-fullbench house 4-item guard against cfg-0092 (the tau2 serving config). 21s window per run-meta.jsonl's own "guard-tau2-config-rocm-c32768" entry (09:04:48Z-09:05:09Z; the preceding 72s server boot is not in this window). Counter interpolated from raw ~10s-resolution HA history samples pulled fresh this session: start 21.80933683257663 kWh, end 21.810138981607825 kWh -- delta 0.802149 kWh*1e-3 = 0.8021 Wh. DISCREPANCY FLAGGED: MODEL-PAGE.md's draft reports this window as 0.84 Wh / 143.8 W. Recomputing the same window (09:04:48Z-09:05:09Z) directly from the raw ~10s HA samples straddling it (09:04:39.9Z/49.9Z/59.9Z/09:05:09.9Z local) gives 0.80 Wh / 137.5 W instead -- a ~5% difference on this one short window only. Every other window in this job's energy join (matrix cells, the 26-task tau2 run itself) reproduced the draft to within rounding; this guard window is the one exception found. Recorded per house policy: recompute from raw data rather than copy the delegate's number: this record uses the recomputed 0.80 Wh / 137.5 W figure, not the draft's 0.84 / 143.8.
eng-0124
- 2026-08-16 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the run-meta window) · window recorded
- wh_total 130.6589 · mean 136.58 W · baseline 10.1 W (box-idle-all-empty) · delta 126.48 W
- per task: 7.2588 Wh · 0.2199 p (tariff 30.3 p/kWh)
- covers: run-0329
- aihydra laguna-s21-fullbench tau2 26-task run window (09:11:47Z-10:09:11Z, 3444s), matching run-0329 exactly (excludes the preceding 73s server-load window, 0.64 Wh / 31.7 W, per the deepseek-v4-flash convention -- see run-0329's note). Counter interpolated from raw ~10s-resolution HA history samples pulled fresh this session: start 21.811893394767356 kWh, end 21.94255229834369 kWh. Recomputed independently from HA raw history (not copied from the delegate's dead-drop report), reproducing its headline arithmetic: 130.66 Wh / 136.6 W mean. wh_per_task is Wh-per-CORRECT-ANSWER: wh_total / tasks_passed (18 of 26, run-0329.metrics) = 7.2588 Wh, 0.2199 pence at the standing 30.3 p/kWh tariff. WHOLE-SESSION energy -- includes simulator (user_llm) wait time end to end, not an agent-only or decode-only figure, directly comparable to deepseek-v4-flash's eng-0107 and nemotron3-super's eng-0116 on the same denominator.
eng-0125
- 2026-08-16 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the run-meta window) · window recorded
- wh_total 1.241 · mean 78.38 W · baseline 10.1 W (box-idle-all-empty) · delta 68.28 W
- covers: run-0332 run-0333
- aihydra ornith-35b-fullbench throughput matrix cell rocm-d0 (3 fresh-process reps, pp1024/tg256 at depth 0). Counter interpolated from raw ~10s-resolution HA history samples pulled fresh this session: start 22.0200983842 kWh, end 22.0213393864 kWh.
eng-0126
- 2026-08-16 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the run-meta window) · window recorded
- wh_total 9.2021 · mean 154.8 W · baseline 10.1 W (box-idle-all-empty) · delta 144.7 W
- covers: run-0334 run-0335
- aihydra ornith-35b-fullbench throughput matrix cell rocm-d32768 (3 fresh-process reps, pp1024/tg256 at depth 32768). Counter interpolated from raw ~10s-resolution HA history samples pulled fresh this session: start 22.0213707742 kWh, end 22.0305728758 kWh.
eng-0127
- 2026-08-16 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the run-meta window) · window recorded
- wh_total 18.5202 · mean 178.75 W · baseline 10.1 W (box-idle-all-empty) · delta 168.65 W
- covers: run-0336 run-0337
- aihydra ornith-35b-fullbench throughput matrix cell rocm-d131072 (N=1 rep, house convention above 32K -- deepseek-v4-flash/laguna-s-21 precedent: the number does not move at this scatter level). Counter interpolated from raw ~10s-resolution HA history samples pulled fresh this session: start 22.0305947666 kWh, end 22.0491149191 kWh.
eng-0128
- 2026-08-16 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the run-meta window) · window recorded
- wh_total 1.2183 · mean 82.76 W · baseline 10.1 W (box-idle-all-empty) · delta 72.66 W
- covers: run-0339 run-0340
- aihydra ornith-35b-fullbench throughput matrix cell vulkan-d0 (3 fresh-process reps, pp1024/tg256 at depth 0). Counter interpolated from raw ~10s-resolution HA history samples pulled fresh this session: start 22.0493805279 kWh, end 22.0505988742 kWh.
eng-0129
- 2026-08-16 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the run-meta window) · window recorded
- wh_total 8.1174 · mean 165.1 W · baseline 10.1 W (box-idle-all-empty) · delta 155 W
- covers: run-0341 run-0342
- aihydra ornith-35b-fullbench throughput matrix cell vulkan-d32768 (3 fresh-process reps, pp1024/tg256 at depth 32768). Counter interpolated from raw ~10s-resolution HA history samples pulled fresh this session: start 22.0505988742 kWh, end 22.0587162559 kWh.
eng-0130
- 2026-08-16 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the run-meta window) · window recorded
- wh_total 15.705 · mean 178.35 W · baseline 10.1 W (box-idle-all-empty) · delta 168.25 W
- covers: run-0343 run-0344
- aihydra ornith-35b-fullbench throughput matrix cell vulkan-d131072 (N=1 rep, house convention above 32K). Unlike laguna-s-21's Vulkan arm, this cell completed cleanly -- no device loss, only one attempt needed. Counter interpolated from raw ~10s-resolution HA history samples pulled fresh this session: start 22.0587365430 kWh, end 22.0744415281 kWh.
eng-0131
- 2026-08-16 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the run-meta window) · window recorded
- wh_total 0.1999 · mean 102.79 W · baseline 10.1 W (box-idle-all-empty) · delta 92.69 W
- covers: run-0338
- aihydra ornith-35b-fullbench house 4-item capability guard against the Vulkan tau2 serving config (cfg-0095), immediately preceding the full 26-task tau2 run. Counter interpolated from raw ~10s-resolution HA history samples pulled fresh this session: start 22.0767107464 kWh, end 22.0769106092 kWh.
eng-0132
- 2026-08-16 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the run-meta window) · window recorded
- wh_total 0.2524 · mean 100.97 W · baseline 10.1 W (box-idle-all-empty) · delta 90.87 W
- covers: run-0330
- Backfilled batch join for the 2026-08-16 ornith-35b screen's house 4-item guard (ROCm, cfg-0093) -- this was flagged energy:null in the candidate record and clm-0068, joined within the ~10-day HA retention window during this job's follow-up fullbench session. Counter interpolated from raw ~10s-resolution HA history samples: start 21.9841945303 kWh, end 21.9844469438 kWh.
eng-0133
- 2026-08-16 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the run-meta window) · window recorded
- wh_total 17.4428 · mean 123.85 W · baseline 10.1 W (box-idle-all-empty) · delta 113.75 W
- per task: 3.4886 Wh · 0.1057 p (tariff 30.3 p/kWh)
- covers: run-0331
- Backfilled batch join for the 2026-08-16 ornith-35b screen's 5-task tau2 SMOKE (ROCm, cfg-0093) -- this was flagged energy:null in the candidate record and clm-0068, joined within the ~10-day HA retention window during this job's follow-up fullbench session. wh_per_task is Wh-per-correct-answer: wh_total / tasks_passed (5 of 5) = 3.4886 Wh, 0.1057 pence at the standing 30.3 p/kWh tariff -- a 5-task SMOKE window, not comparable to the full 26-task standard (see eng-0134). Counter interpolated from raw ~10s-resolution HA history samples: start 21.9845072601 kWh, end 22.0019500298 kWh.
eng-0134
- 2026-08-16 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the run-meta window) · window recorded
- wh_total 149.1174 · mean 128.03 W · baseline 10.1 W (box-idle-all-empty) · delta 117.93 W
- per task: 6.4834 Wh · 0.1964 p (tariff 30.3 p/kWh)
- covers: run-0345
- aihydra ornith-35b-fullbench tau2 26-task run window (12:58:58Z-14:08:51Z, 4193s), matching run-0345 exactly (excludes the preceding 13s server-load window, 0.1246 Wh / 34.5 W, per the deepseek-v4-flash/laguna-s-21 convention -- including it moves the figure by <0.1%: 149.2420 Wh total, 6.4888 Wh/correct). Counter interpolated from raw ~10s-resolution HA history samples pulled fresh this session: start 22.0771553905 kWh, end 22.2262727988 kWh. wh_per_task is Wh-per-CORRECT-ANSWER: wh_total / tasks_passed (23 of 26, run-0345.metrics) = 6.4834 Wh, 0.1964 pence at the standing 30.3 p/kWh tariff. WHOLE-SESSION energy -- includes simulator (user_llm) wait time end to end, not an agent-only or decode-only figure, directly comparable to deepseek-v4-flash's eng-0107, nemotron3-super's eng-0116 and laguna-s-21's eng-0124 on the same denominator. THIS IS THE NEW LEADER among full standard-26-task tau2 comparisons on this box, ahead of laguna-s-21's 7.2588 Wh, deepseek-v4-flash's 24.29 Wh, nemotron3-super's 32.43 Wh and qwen38-27b's 38.29 Wh -- screened against every other benched model's Wh-per-correct-answer before being claimed (the qwen35-122b 5.71 Wh/correct figure is a 5-task SMOKE window predating the standard-26-task convention and remains not a like-for-like comparator, same caveat laguna-s-21's page already carries).
eng-0135
- 2026-08-16 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the run-meta window) · window recorded
- wh_total 4.3489 · mean 109.48 W · baseline 10.1 W (box-idle-all-empty) · delta 99.38 W
- covers: run-0346 run-0347
- aihydra qwen36-27b-mtp-fullbench throughput matrix cell rocm-d0 (3 fresh-process reps, pp1024/tg256 at depth 0). Counter interpolated from raw ~10s-resolution HA history samples pulled fresh this session: start 22.2315131320 kWh, end 22.2358619935 kWh.
eng-0136
- 2026-08-16 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the run-meta window) · window recorded
- wh_total 25.4693 · mean 178.04 W · baseline 10.1 W (box-idle-all-empty) · delta 167.94 W
- covers: run-0348 run-0349
- aihydra qwen36-27b-mtp-fullbench throughput matrix cell rocm-d32768 (3 fresh-process reps, pp1024/tg256 at depth 32768). Counter interpolated from raw ~10s-resolution HA history samples pulled fresh this session: start 22.2358968413 kWh, end 22.2613661774 kWh.
eng-0137
- 2026-08-16 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the run-meta window) · window recorded
- wh_total 46.973 · mean 178.94 W · baseline 10.1 W (box-idle-all-empty) · delta 168.84 W
- covers: run-0350 run-0351
- aihydra qwen36-27b-mtp-fullbench throughput matrix cell rocm-d131072 (1 rep, pp1024/tg256 at depth 131072). Counter interpolated from raw ~10s-resolution HA history samples pulled fresh this session: start 22.2613661774 kWh, end 22.3083391897 kWh.
eng-0138
- 2026-08-16 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the run-meta window) · window recorded
- wh_total 3.981 · mean 115.58 W · baseline 10.1 W (box-idle-all-empty) · delta 105.48 W
- covers: run-0352 run-0353
- aihydra qwen36-27b-mtp-fullbench throughput matrix cell vulkan-d0 (3 fresh-process reps, pp1024/tg256 at depth 0). Counter interpolated from raw ~10s-resolution HA history samples pulled fresh this session: start 22.3092380594 kWh, end 22.3132190570 kWh.
eng-0139
- 2026-08-16 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the run-meta window) · window recorded
- wh_total 30.656 · mean 167.21 W · baseline 10.1 W (box-idle-all-empty) · delta 157.11 W
- covers: run-0354 run-0355
- aihydra qwen36-27b-mtp-fullbench throughput matrix cell vulkan-d32768 (3 fresh-process reps, pp1024/tg256 at depth 32768). Counter interpolated from raw ~10s-resolution HA history samples pulled fresh this session: start 22.3132190570 kWh, end 22.3438750970 kWh.
eng-0140
- 2026-08-16 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the run-meta window) · window recorded
- wh_total 0.8499 · mean 145.7 W · baseline 10.1 W (box-idle-all-empty) · delta 135.6 W
- covers: run-0357
- aihydra qwen36-27b-mtp-fullbench house guard (run-0357), against the tau2 serving config (cfg-0099). Counter interpolated from raw ~10s-resolution HA history samples: start 22.4306878575 kWh, end 22.4315377590 kWh.
eng-0141
- 2026-08-16 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the run-meta window) · window recorded
- wh_total 1446.5888 · mean 180.82 W · baseline 10.1 W (box-idle-all-empty) · delta 170.72 W
- per task: 103.3278 Wh · 3.1308 p (tariff 30.3 p/kWh)
- covers: run-0358
- aihydra qwen36-27b-mtp-fullbench tau2 arm1 window (16:14:57Z-00:14:57Z, 28800s, the full 8h TAU2_TIMEOUT safety ceiling — a BOUND-CUT arm, 20/26 tasks scored, 14 passed). wh_per_task is Wh-per-CORRECT-ANSWER for THIS ARM ONLY: 1446.5888 / 14 = 103.3278 Wh, 3.1308 pence at 30.3 p/kWh. WHOLE-SESSION energy (includes simulator wait time end to end), same denominator convention as every other board entry. See clm-0076 for the COMBINED figure across this arm and the continuation (run-0360 / eng-0143): 2055.19 Wh / 16 correct = 128.45 Wh per correct answer — the figure that belongs on the model page, since 26 scored/25-scored is the comparable unit, not this arm alone. Counter interpolated from raw ~10s-resolution HA history samples pulled fresh this session: start 22.4315377590 kWh, end 23.8781265856 kWh.
eng-0142
- 2026-08-17 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the run-meta window) · window recorded
- wh_total 0.9452 · mean 154.68 W · baseline 10.1 W (box-idle-all-empty) · delta 144.58 W
- covers: run-0359
- aihydra qwen36-27b-mtp-fullbench house guard (run-0359), re-run fresh against the continuation server before the remaining 5 tau2 task-ids. Counter interpolated from raw ~10s-resolution HA history samples: start 23.8790832519 kWh, end 23.8800284894 kWh.
eng-0143
- 2026-08-17 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the run-meta window) · window recorded
- wh_total 608.6042 · mean 182.04 W · baseline 10.1 W (box-idle-all-empty) · delta 171.94 W
- per task: 304.3021 Wh · 9.2203 p (tariff 30.3 p/kWh)
- covers: run-0360
- aihydra qwen36-27b-mtp-fullbench tau2 arm2 (continuation, task-ids 21-25) window (00:18:28Z-03:39:04Z, 12036s / 3.34h). wh_per_task is Wh-per-CORRECT-ANSWER for THIS ARM ONLY: 608.6042 / 2 = 304.3021 Wh, 9.2203 pence at 30.3 p/kWh — a small-N figure (only 2 of 5 passed, including task 23's ~90-minute cumulative wall time) that is NOT representative alone. See clm-0076 and eng-0141 for the COMBINED 26-task figure (2055.19 Wh / 16 correct = 128.45 Wh per correct answer), the one that belongs on the model page. Counter interpolated from raw ~10s-resolution HA history samples pulled fresh this session: start 23.8800284894 kWh, end 24.4886327086 kWh.
eng-0144
- 2026-08-16 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to each run-meta window; two windows summed, not a single counter-diff over the outer span) · window recorded
- wh_total 2055.193 · mean 181.18 W · baseline 10.1 W (box-idle-all-empty) · delta 171.08 W
- per task: 128.4496 Wh · 3.892 p (tariff 30.3 p/kWh)
- covers: run-0358 run-0360
- COMBINED energy across the full standard 26-task tau2 airline set, summed from its two constituent arm windows (eng-0141: arm1, 1446.5888 Wh over 28,800s; eng-0143: arm2/continuation, 608.6042 Wh over 12,036s) = 2055.1930 Wh total. This is a SUM of two measured windows, NOT a fresh counter-difference over the full outer span (16:14:57Z-03:39:04Z, which would also include a ~3-minute server-stop/guard/ server-restart gap between the two sessions — matching the exclude-overhead convention every other board entry's tau2 energy figure already uses, e.g. ornith-35b eng-0134 excluding its preceding 13s server-load window). mean_w_active (181.18 W) and delta_w (171.08 W) are computed over the SUM of the two active durations (28,800s + 12,036s = 40,836s), not the outer wall-clock span — a reasonable single figure since both arms ran the same config at similar load (arm1 mean 180.82 W, arm2 mean 182.04 W, within 1.2 W of each other); see eng-0141/eng-0143 for the per-arm mean/delta wattage individually. wh_per_task is Wh-per-CORRECT-ANSWER across BOTH arms: 2055.1930 / 16 (14 arm1 + 2 arm2 passed, of 25 scored) = 128.4496 Wh, 3.8920 pence at 30.3 p/kWh. THE LOWEST-RANKED Wh-per-correct-answer of any full-bench candidate's Wh-per-correct-answer figure on this board (clm-0076): trails qwen38-27b's previous-worst 38.29 Wh/correct by >3x, and ornith-35b's leading 6.48 Wh/correct by ~20x.
eng-0145
- 2026-08-17 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the run-meta window) · window recorded
- wh_total 1.3051 · mean 38.2 W · baseline 10.1 W (box-idle-all-empty) · delta 28.1 W
- covers: run-0363 run-0364 run-0365 run-0366 run-0367 run-0368 run-0369 run-0370 run-0371
- aihydra deep-thought-posttrain throughput matrix (run-0364..run-0371, 8 cells x 2 backends) + house 4-item guard (run-0363), combined window 06:31:48Z-06:33:51Z (123s) -- matches the matrix+guard runs exactly (excludes the preceding sanity-check server starts and the earlier two-point kvprobe, none of which are referenced by a run id, per the deepseek-v4-flash/laguna-s-21/ornith-35b convention of excluding unreferenced windows). Counter interpolated from raw ~10s-resolution HA history samples pulled fresh this session: window-start bracket 24.8110571354627676/24.8111601322889376 kWh (06:31:40.557Z/06:31:50.518Z) -> interpolated 24.8111341 kWh at 06:31:48Z; window-end bracket 24.8124376982450476/24.8124698847532276 kWh (06:33:50.538Z/06:34:00.528Z) -> interpolated 24.8124392 kWh at 06:33:51Z. Delta 0.0013051 kWh = 1.3051 Wh over 123s = 38.20 W mean, 28.10 W over the 10.1 W box-idle baseline -- a small, genuinely measured delta consistent with a 361.8M model doing twelve few-second llama-bench cells and one 8-second guard call, not an artefact: the ROCm cells (06:31:48-06:32:09, slower prefill but the two most compute-dense cells in the matrix) visibly carry the counter's steepest climb in the raw samples above (24.8111601 -> 24.8120493 kWh across 06:31:50Z-06:32:30Z, the matrix's busiest 40s), settling to a shallower slope during the Vulkan cells and the guard call. wh_per_task is null: this window covers throughput + a capability guard, not a scored task set (the tau2 window is a separate energy record, joined once that run completes). tariff at the standing 30.3 p/kWh.
eng-0146
- 2026-08-17 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the run-meta window) · window recorded
- wh_total 22.7285 · mean 54.69 W · baseline 10.1 W (box-idle-all-empty) · delta 44.59 W
- per task: 4.5457 Wh · 0.1377 p (tariff 30.3 p/kWh)
- covers: run-0372
- aihydra deep-thought-posttrain FULL 26-task tau2 airline run (run-0372), window 06:37:19Z-07:02:15Z (1496s) matching run-meta.jsonl's tau2-full-26task line exactly. Counter interpolated from raw ~10s-resolution HA history pulled fresh this session: window-start bracket 24.8130654543638276/24.8131226748228076 kWh (06:37:10.533Z/06:37:20.529Z) -> interpolated 24.8131139 kWh at 06:37:19Z; window-end bracket 24.8358238190412576/24.8358655422925976 kWh (07:02:10.545Z/07:02:20.551Z) -> interpolated 24.8358424 kWh at 07:02:15Z. Delta 0.0227285 kWh = 22.7285 Wh over 1496s = 54.69 W mean, 44.59 W over the 10.1 W box-idle baseline — a real, elevated draw (the raw samples show a sustained faster climb from ~06:39Z through ~06:56Z, tracking the busiest run of tasks in the log, settling toward the baseline slope in the final ~6 minutes as the remaining tasks resolved quickly or errored near-instantly). wh_per_task IS ARITHMETICALLY DEFINED HERE (5 tasks scored reward 1.0 out of 26 — not the zero-correct-answers case that would make this a literal division by zero) but is published with a hard caveat, not as a clean comparator: run-0372's own guard (run-0363) and this arm's own tool_call_messages count (0 across all 26 simulations, 233/233 assistant messages) place this run squarely inside the SMOKE validity gate (protocol.json comparison_rules / benchmark-protocol.md §11) — a mean computed where the agent never calls a tool is INVALID, because the airline domain hands out an untouched-database point and a vacuous COMMUNICATE point to any agent that never acts, which is mechanically how a model that only ever says "42" collected 5 nominal "passes". 4.5457 Wh / 0.1377p per NOMINALLY-correct answer is therefore reported for completeness and arithmetic honesty only, and is explicitly NOT comparable to any other model's Wh-per-correct-answer figure on this site (ornith-35b 6.4834, laguna-s-21 7.2588, deepseek-v4-flash 24.29, nemotron3-super 32.43, qwen38-27b 38.29) — those all cleared the SMOKE gate with a nonzero tool_call_messages count; this one did not, and its "correct answers" are not evidence of task completion. wh_total (22.7285 Wh, the undiluted figure) is the number this run actually supports.
eng-0147
- 2026-08-17 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the session window) · window reconstructed
- wh_total 108.382 · mean 173.64 W · baseline 10.1 W (box-idle-all-empty) · delta 163.54 W
- covers: run-0373 run-0374
- Deadline-critical batch join (protocol §11) for the 2026-08-17 depth-ladder backfill, qwen38-27b Q8_0 ROCm d204800 (cfg-0103), covering both the prefill row (run-0373) and decode row (run-0374) from the single llama-bench invocation. Session window reconstructed from ~/bench-results/depth-backfill-20260817/qwen38-27b/d204800.started (07:15:40Z) and the output JSON's file mtime (07:53:07Z) -- the backfill wrapper this session used did not call run_begin/run_end for every cell (only its first cell, qwen36-27b-mtp, has a run-meta.jsonl entry), so window_provenance is reconstructed, not recorded. Counter interpolated from raw ~10s-resolution HA history samples: start 24.838408..24.838411 kWh (07:15:40Z), end 24.946790 kWh (07:53:07Z).
eng-0148
- 2026-08-17 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the session window) · window reconstructed
- wh_total 159.289 · mean 176.77 W · baseline 10.1 W (box-idle-all-empty) · delta 166.67 W
- covers: run-0375 run-0376
- Deadline-critical batch join (protocol §11) for the 2026-08-17 depth-ladder backfill, qwen38-27b Q8_0 ROCm d262144 model-max (cfg-0103), covering both the prefill row (run-0375) and decode row (run-0376) from the single llama-bench invocation. Session window reconstructed from ~/bench-results/depth-backfill-20260817/qwen38-27b/d262144.started (08:37:58Z) and the output JSON's file mtime (09:32:02Z); same reconstructed-not-recorded provenance as eng-0147. Counter interpolated from raw ~10s-resolution HA history samples: start 24.955448 kWh (08:37:58Z), end 25.114737 kWh (09:32:02Z). The deepest cell in this backfill (~54 min wall time), consistent with the run-0375 comment's own wall-time estimate.
eng-0149
- 2026-08-17 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the session window) · window reconstructed
- wh_total 37.114 · mean 179.34 W · baseline 10.1 W (box-idle-all-empty) · delta 169.24 W
- covers: run-0377 run-0378
- Deadline-critical batch join (protocol §11) for the 2026-08-17 depth-ladder backfill, ornith-35b Q8_0 ROCm d204800 (cfg-0104), covering both the prefill row (run-0377) and decode row (run-0378) from the single llama-bench invocation. Session window taken from the launch/landing timestamps embedded in the run-0378 record's own comment (09:35:22Z launch, 09:47:35Z decode landing) cross-checked against the ~/bench-results/depth-backfill-20260817/ornith-35b/d204800.json file mtime (09:47:47Z, used as the end boundary); reconstructed, not runmeta-recorded (same gap as eng-0147/ eng-0148). Counter interpolated from raw ~10s-resolution HA history samples: start 25.115512 kWh (09:35:22Z), end 25.152626 kWh (09:47:47Z). Cheapest cell of the backfill by wall time (~12 min) -- sparse A3B-active MoE, per the run-0378 comment.
eng-0150
- 2026-08-17 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the session window) · window reconstructed
- wh_total 55.405 · mean 179.37 W · baseline 10.1 W (box-idle-all-empty) · delta 169.27 W
- covers: run-0379 run-0380
- Deadline-critical batch join (protocol §11) for the 2026-08-17 depth-ladder backfill, ornith-35b Q8_0 ROCm d262144 model-max (cfg-0104), covering both the prefill row (run-0379) and decode row (run-0380) from the single llama-bench invocation. Session window taken from the launch timestamp embedded in the run-0380 record's own comment (09:48:14Z launch) cross-checked against the ~/bench-results/depth-backfill-20260817/ornith-35b/d262144.json file mtime (10:06:46Z, used as the end boundary); reconstructed, not runmeta-recorded (same gap as eng-0147..eng-0149). Counter interpolated from raw ~10s-resolution HA history samples: start 25.152866 kWh (09:48:14Z), end 25.208271 kWh (10:06:46Z). Near-identical mean power to eng-0149 (179.37 W vs 179.34 W) -- same model/backend/build, only depth differs, consistent with a compute-bound GPU draw that doesn't move much with depth at this model's cheap-KV footprint.
eng-0151
- 2026-08-17 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the session window) · window reconstructed
- wh_total 36.621 · mean 171.44 W · baseline 10.1 W (box-idle-all-empty) · delta 161.34 W
- covers: run-0381 run-0382
- Deadline-critical batch join (protocol §11) for the 2026-08-17 depth-ladder backfill, ornith-35b Q8_0 Vulkan d204800 (cfg-0105), covering both the prefill row (run-0381) and decode row (run-0382) from the single llama-bench invocation. Session window taken from the launch timestamp embedded in the run-0382 record's own comment (10:07:10Z launch) cross-checked against the ~/bench-results/depth-backfill-20260817/ornith-35b/vk-d204800.json file mtime (10:19:59Z, used as the end boundary); reconstructed, not runmeta-recorded (same gap as eng-0147..eng-0150). Counter interpolated from raw ~10s-resolution HA history samples: start 25.208454 kWh (10:07:10Z), end 25.245075 kWh (10:19:59Z). Slightly lower mean draw than the ROCm arms (171.44 W vs ~179 W) despite Vulkan's higher decode throughput at this depth (26.53 vs 22.23 t/s, per run-0382's comment) -- a same-model cross-backend efficiency contrast worth surfacing per protocol.json's compute-surfaces comparison rule.
eng-0152
- 2026-08-17 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the run-meta window) · window recorded
- wh_total 59.668 · mean 157.94 W · baseline 10.1 W (box-idle-all-empty) · delta 147.84 W
- covers: run-0383 run-0384
- Live join (protocol §11, joined same session as the run per the job brief's "join Part 1/2's own arms live as you go") for ornith-35b Q8_0 Vulkan d262144 model-max (cfg-0106), covering both the prefill row (run-0383) and decode row (run-0384). Window is run-meta-recorded (run_begin/run_end used directly this time, unlike the Part 3 backfill this session also joined), so window_provenance is "recorded" not "reconstructed". Counter interpolated from raw ~10s-resolution HA history samples: start 25.269026 kWh (11:46:56Z), end 25.328694 kWh (12:09:36Z). Lower mean draw than the ROCm arms in this session's eng-0147..eng-0150 (157.94 W vs ~173-179 W), consistent with eng-0151's same cross-backend pattern at d204800.
eng-0153
- 2026-08-17 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the run-meta window) · window recorded
- wh_total 32.569 · mean 175.28 W · baseline 10.1 W (box-idle-all-empty) · delta 165.18 W
- covers: run-0385 run-0386
- Live join (protocol §11, "join Part 1/2's own arms live as you go") for nemotron3-super UD-Q4_K_M ROCm d131072 (cfg-0107), covering both the prefill row (run-0385) and decode row (run-0386). Window is run-meta-recorded. Counter interpolated from raw ~10s-resolution HA history samples: start 25.329403 kWh (12:12:36Z), end 25.361972 kWh (12:23:45Z).
eng-0154
- 2026-08-17 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the run-meta window) · window recorded
- wh_total 53.548 · mean 176.71 W · baseline 10.1 W (box-idle-all-empty) · delta 166.61 W
- covers: run-0387 run-0388
- Live join (protocol §11) for nemotron3-super UD-Q4_K_M ROCm d204800 (cfg-0107), covering both the prefill row (run-0387) and decode row (run-0388). Window is run-meta-recorded. Counter interpolated from raw ~10s-resolution HA history samples: start 25.362194 kWh (12:24:11Z), end 25.415742 kWh (12:42:22Z).
eng-0155
- 2026-08-17 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the run-meta window) · window recorded
- wh_total 39.369 · mean 172.42 W · baseline 10.1 W (box-idle-all-empty) · delta 162.32 W
- covers: run-0389 run-0390
- Live join (protocol §11) for nemotron3-super UD-Q4_K_M Vulkan d131072, attempt 1 of the house 2-attempt cap (cfg-0108), covering both the prefill row (run-0389) and decode row (run-0390). Window is run-meta-recorded. Counter interpolated from raw ~10s-resolution HA history samples: start 25.415905 kWh (12:42:48Z), end 25.455274 kWh (12:56:30Z).
eng-0156
- 2026-08-17 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the run-meta window) · window recorded
- wh_total 81.039 · mean 170.72 W · baseline 10.1 W (box-idle-all-empty) · delta 160.62 W
- covers: run-0391 run-0392
- Live join (protocol §11) for laguna-s-21 Q4_K_M ROCm d204800 (cfg-0109), covering both the prefill row (run-0391) and decode row (run-0392). Window is run-meta-recorded. Counter interpolated from raw ~10s-resolution HA history samples: start 25.457095 kWh (13:05:42Z), end 25.538134 kWh (13:34:11Z).
eng-0157
- 2026-08-17 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the run-meta window) · window recorded
- wh_total 2.1625 · mean 80.26 W · baseline 10.1 W (box-idle-all-empty) · delta 70.16 W
- covers: run-0396 run-0397 run-0398 run-0399 run-0400 run-0401
- aihydra qwen3-coder-next-80b-fullbench-matrix throughput cell, rocm d0 (3 rep(s), 6 runs: prefill+decode pairs), window 17:01:21-17:02:58Z (97s). Counter interpolated from raw ~10s-resolution HA history samples: start 25.925755 kWh, end 25.927918 kWh. Cell-window energy (server load + all reps together), not per-rep split.
eng-0158
- 2026-08-17 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the run-meta window) · window recorded
- wh_total 15.0605 · mean 156.7 W · baseline 10.1 W (box-idle-all-empty) · delta 146.6 W
- covers: run-0408 run-0409 run-0410 run-0411 run-0412 run-0413
- aihydra qwen3-coder-next-80b-fullbench-matrix throughput cell, rocm d32768 (3 rep(s), 6 runs: prefill+decode pairs), window 17:02:59-17:08:45Z (346s). Counter interpolated from raw ~10s-resolution HA history samples: start 25.927948 kWh, end 25.943008 kWh. Cell-window energy (server load + all reps together), not per-rep split.
eng-0159
- 2026-08-17 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the run-meta window) · window recorded
- wh_total 27.1375 · mean 178.28 W · baseline 10.1 W (box-idle-all-empty) · delta 168.18 W
- covers: run-0402 run-0403
- aihydra qwen3-coder-next-80b-fullbench-matrix throughput cell, rocm d131072 (1 rep(s), 2 runs: prefill+decode pairs), window 17:08:46-17:17:54Z (548s). Counter interpolated from raw ~10s-resolution HA history samples: start 25.943032 kWh, end 25.97017 kWh. Cell-window energy (server load + all reps together), not per-rep split.
eng-0160
- 2026-08-17 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the run-meta window) · window recorded
- wh_total 53.1701 · mean 178.06 W · baseline 10.1 W (box-idle-all-empty) · delta 167.96 W
- covers: run-0404 run-0405
- aihydra qwen3-coder-next-80b-fullbench-matrix throughput cell, rocm d204800 (1 rep(s), 2 runs: prefill+decode pairs), window 17:17:55-17:35:50Z (1075s). Counter interpolated from raw ~10s-resolution HA history samples: start 25.970193 kWh, end 26.023363 kWh. Cell-window energy (server load + all reps together), not per-rep split.
eng-0161
- 2026-08-17 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the run-meta window) · window recorded
- wh_total 77.4029 · mean 175.25 W · baseline 10.1 W (box-idle-all-empty) · delta 165.15 W
- covers: run-0406 run-0407
- aihydra qwen3-coder-next-80b-fullbench-matrix throughput cell, rocm d262144 (1 rep(s), 2 runs: prefill+decode pairs), window 17:35:51-18:02:21Z (1590s). Counter interpolated from raw ~10s-resolution HA history samples: start 26.023398 kWh, end 26.100801 kWh. Cell-window energy (server load + all reps together), not per-rep split.
eng-0162
- 2026-08-17 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the run-meta window) · window recorded
- wh_total 2.7852 · mean 53.05 W · baseline 10.1 W (box-idle-all-empty) · delta 42.95 W
- covers: run-0414 run-0415 run-0416 run-0417 run-0418 run-0419
- aihydra qwen3-coder-next-80b-fullbench-matrix throughput cell, vulkan d0 (3 rep(s), 6 runs: prefill+decode pairs), window 18:03:13-18:06:22Z (189s). Counter interpolated from raw ~10s-resolution HA history samples: start 26.101421 kWh, end 26.104206 kWh. Cell-window energy (server load + all reps together), not per-rep split.
eng-0163
- 2026-08-17 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the run-meta window) · window recorded
- wh_total 13.1749 · mean 128.89 W · baseline 10.1 W (box-idle-all-empty) · delta 118.79 W
- covers: run-0426 run-0427 run-0428 run-0429 run-0430 run-0431
- aihydra qwen3-coder-next-80b-fullbench-matrix throughput cell, vulkan d32768 (3 rep(s), 6 runs: prefill+decode pairs), window 18:06:22-18:12:30Z (368s). Counter interpolated from raw ~10s-resolution HA history samples: start 26.104206 kWh, end 26.117381 kWh. Cell-window energy (server load + all reps together), not per-rep split.
eng-0164
- 2026-08-17 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the run-meta window) · window recorded
- wh_total 22.0572 · mean 169.31 W · baseline 10.1 W (box-idle-all-empty) · delta 159.21 W
- covers: run-0420 run-0421
- aihydra qwen3-coder-next-80b-fullbench-matrix throughput cell, vulkan d131072 (1 rep(s), 2 runs: prefill+decode pairs), window 18:12:31-18:20:20Z (469s). Counter interpolated from raw ~10s-resolution HA history samples: start 26.117418 kWh, end 26.139475 kWh. Cell-window energy (server load + all reps together), not per-rep split.
eng-0165
- 2026-08-17 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the run-meta window) · window recorded
- wh_total 46.8883 · mean 169.82 W · baseline 10.1 W (box-idle-all-empty) · delta 159.72 W
- covers: run-0422 run-0423
- aihydra qwen3-coder-next-80b-fullbench-matrix throughput cell, vulkan d204800 (1 rep(s), 2 runs: prefill+decode pairs), window 18:20:20-18:36:54Z (994s). Counter interpolated from raw ~10s-resolution HA history samples: start 26.139475 kWh, end 26.186363 kWh. Cell-window energy (server load + all reps together), not per-rep split.
eng-0166
- 2026-08-17 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the run-meta window) · window recorded
- wh_total 77.4502 · mean 160.7 W · baseline 10.1 W (box-idle-all-empty) · delta 150.6 W
- covers: run-0424 run-0425
- aihydra qwen3-coder-next-80b-fullbench-matrix throughput cell, vulkan d262144 (1 rep(s), 2 runs: prefill+decode pairs), window 18:36:55-19:05:50Z (1735s). Counter interpolated from raw ~10s-resolution HA history samples: start 26.186384 kWh, end 26.263834 kWh. Cell-window energy (server load + all reps together), not per-rep split.
eng-0167
- 2026-08-17 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the run-meta window) · window recorded
- wh_total 9.682 · mean 119.77 W · baseline 10.1 W (box-idle-all-empty) · delta 109.67 W
- per task: 2.4205 Wh · 0.0733 p (tariff 30.3 p/kWh)
- covers: run-0432
- aihydra qwen3-coder-next-80b-screen SMOKE window (16:54:29Z-16:59:20Z, 291s), matching run-0432 exactly. Counter interpolated from raw ~10s-resolution HA history samples: start 25.915112 kWh, end 25.924794 kWh. wh_per_task is Wh-per-CORRECT-ANSWER: wh_total / tasks_passed (4 of 5, run-0432.metrics) = 2.4205 Wh, 0.0733 pence at the standing 30.3 p/kWh tariff. WHOLE-SESSION energy (server already warm from the FIT probe moments earlier; includes simulator wait time end to end). A 5-task SMOKE window, not a like-for-like comparator against any full standard-26-task tau2 Wh/correct figure on this board (see eng-0168 for that comparison).
eng-0168
- 2026-08-17 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the run-meta window) · window recorded
- wh_total 67.054 · mean 103.69 W · baseline 10.1 W (box-idle-all-empty) · delta 93.59 W
- per task: 4.7896 Wh · 0.1451 p (tariff 30.3 p/kWh)
- covers: run-0433
- aihydra qwen3-coder-next-80b-fullbench tau2 26-task run window (19:09:14Z-19:48:02Z, 2328s), matching run-0433 exactly (excludes the preceding 51s server-load window, per the deepseek-v4-flash/laguna-s-21/ ornith-35b convention). Counter interpolated from raw ~10s-resolution HA history samples pulled fresh this session: start 26.265702 kWh, end 26.332756 kWh. wh_per_task is Wh-per-CORRECT-ANSWER: wh_total / tasks_passed (14 of 26, run-0433.metrics) = 4.7896 Wh, 0.1451 pence at the standing 30.3 p/kWh tariff. WHOLE-SESSION energy -- includes simulator (user_llm) wait time end to end, not an agent-only or decode-only figure, directly comparable to ornith-35b's eng-0134 (6.4834 Wh), laguna-s-21's eng-0124 (7.2588 Wh), deepseek-v4-flash's eng-0107 (24.29 Wh), nemotron3-super's eng-0116 (32.43 Wh) and qwen38-27b's Wh figure (38.29 Wh) on the same denominator. THIS IS THE NEW LOWEST Wh-per-correct-answer among full standard-26-task tau2 comparisons on this box, ahead of ornith-35b's previous-leading 6.4834 Wh -- notable because this candidate has, at the SAME TIME, the LOWEST capability score of the same comparison set (0.5385 mean reward, below qwen38-27b's 0.577): the efficiency lead is real and measured on the identical protocol, but it does not come with a capability lead, and the verdict must not imply otherwise. Screened against every other benched model's Wh-per-correct-answer before being claimed (the qwen35-122b 5.71 Wh/correct figure remains a 5-task SMOKE window predating the standard-26-task convention and is not a like-for-like comparator, same caveat every prior leader's page already carries).
eng-0169
- 2026-08-17 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the run-meta window) · window recorded
- wh_total 0.3274 · mean 107.14 W · baseline 10.1 W (box-idle-all-empty) · delta 97.04 W
- covers: run-0395
- aihydra qwen3-coder-next-80b tau2-config Vulkan guard window (19:08:02Z-19:08:13Z, 11s), matching run-0395 exactly (the house 4-item guard against cfg-0111 immediately before the full tau2 run, run-0433). Counter interpolated from raw ~10s-resolution HA history samples: start 26.264707 kWh, end 26.265034 kWh. Trivially short window (four cheap HTTP probes); recorded for completeness per protocol §11, not itself a meaningful throughput/efficiency figure.
eng-0170
- 2026-08-18 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the run-meta window) · window recorded
- wh_total 0.3624 · mean 118.6 W · baseline 10.1 W (box-idle-all-empty) · delta 108.5 W
- covers: run-0434
- Ling-3.0-flash corrected full-bench guard window (11s), matching run-0434. Counter interpolated from raw ~10s-resolution HA history samples: start 26.505894624 kWh, end 26.506257024 kWh. Short guard window recorded for completeness per protocol §11, not itself a meaningful efficiency figure.
eng-0171
- 2026-08-18 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the run-meta window) · window recorded
- wh_total 1.6968 · mean 83.68 W · baseline 10.1 W (box-idle-all-empty) · delta 73.58 W
- covers: run-0435 run-0436
- Ling-3.0-flash ROCm f16 d0 llama-bench invocation, covering both the prefill row (run-0435) and decode row (run-0436). Counter interpolated from raw ~10s-resolution HA history samples: start 26.494758163 kWh, end 26.496454987 kWh. Mean draw is low because this short window includes model setup/teardown around a shallow synthetic throughput cell.
eng-0172
- 2026-08-18 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the run-meta window) · window recorded
- wh_total 8.6674 · mean 139.92 W · baseline 10.1 W (box-idle-all-empty) · delta 129.82 W
- covers: run-0437 run-0438
- Ling-3.0-flash ROCm f16 d32768 llama-bench invocation, covering both the prefill row (run-0437) and decode row (run-0438). Counter interpolated from raw ~10s-resolution HA history samples: start 26.496506417 kWh, end 26.505173845 kWh. Same build/backend/flags as the d0 cell, with no mmap.
eng-0173
- 2026-08-18 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the run-meta window) · window recorded
- wh_total 69.218 · mean 138.21 W · baseline 10.1 W (box-idle-all-empty) · delta 128.11 W
- per task: 5.3245 Wh · 0.1613 p (tariff 30.3 p/kWh)
- covers: run-0439
- Ling-3.0-flash corrected full 26-task tau2 airline run (run-0439), plain decode on ROCm, pinned Haiku-4.5 simulator, rc=0 with all 26 tasks completed. Counter interpolated from raw ~10s-resolution HA history samples: start 26.506257024 kWh, end 26.575475071 kWh, delta 0.069218047 kWh = 69.2180 Wh. wh_per_task is Wh per CORRECT answer: wh_total / tasks_passed (13 of 26). Whole-session energy includes simulator wait time end to end, matching the site's headline denominator.
eng-0174
- 2026-08-09 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the run-meta window) · window reconstructed
- wh_total 714.9505 · mean 156.4 W · baseline 10.1 W (box-idle-all-empty) · delta 146.3 W
- per task: 59.5792 Wh · 1.8053 p (tariff 30.3 p/kWh)
- covers: run-0100
- aihydra qwen35-122b tau2-bench-airline full 22-task arm window (2026-08-09T21:05:52Z..2026-08-10T01:40:06Z, 16454s), taken from run-0100's started_at/ended_at. run-0100 predates run-meta.jsonl instrumentation for this arm, so the window is RECONSTRUCTED from the transcript timestamps (first task start to last task end), not a recorded meter window — this is why window_provenance is reconstructed rather than recorded. Counter interpolated from raw ~10s-resolution HA history samples pulled fresh 2026-08-18: start 4.805120 kWh, end 5.520070 kWh. Note the run carries internal idle gaps (e.g. task 6 ends 21:18:05Z, task 7 begins 23:48:52Z) — those are included end-to-end as wall-clock part of the session, giving a whole-session figure that includes any simulator/dwell time, the same denominator as the other full-arm joins. wh_per_task is Wh-per-CORRECT-ANSWER: wh_total / tasks_passed (12 of 22, run-0100.metrics) = 59.58 Wh, 1.81 pence/correct at the standing 30.3 p/kWh tariff. NOTE: this arm ran an UNPINNED user simulator (self-play, user_llm = the model's own path, temperature 0) — it is a provisional arm per clm-0043 and does NOT change qwen35-122b's published headline energy (eng-0034, the 5-task smoke, remains the page's best.wh_correct_energy pointer).
eng-0175
- 2026-08-10 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the transcript-derived window) · window reconstructed
- wh_total 802.6684 · mean 153.8 W · baseline 10.1 W (box-idle-all-empty) · delta 143.7 W
- per task: 100.3336 Wh · 3.0401 p (tariff 30.3 p/kWh)
- covers: run-0101
- aihydra qwen35-122b tau2-bench-airline 9-scored-task arm window (2026-08-10T03:06:47Z..2026-08-10T08:19:49Z, 18782s), taken from run-0101's started_at/ended_at. run-0101 predates matching run-meta instrumentation for this arm, so the window is reconstructed from transcript timestamps rather than a recorded meter window. Counter interpolated from raw ~10s-resolution HA history samples pulled fresh 2026-08-18 as part of the run-0101/run-0102/run-0111 batch: start 5.740920207 kWh, end 6.543588647 kWh. The transcript contains a long internal idle gap between task 6 and the later task 7/8/9 block; the gap is included end-to-end as wall-clock session energy. wh_per_task is Wh-per-CORRECT-ANSWER: wh_total / tasks_passed (8 of 9 scored tasks, run-0101.metrics) = 100.3336 Wh, 3.0401 pence/correct at the standing 30.3 p/kWh tariff. This arm ran an unpinned self-play user simulator (user_llm is the same qwen35-122b path as the agent), so it remains provisional per clm-0043.
eng-0176
- 2026-08-10 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the transcript-derived window) · window reconstructed
- wh_total 876.4963 · mean 166.9 W · baseline 10.1 W (box-idle-all-empty) · delta 156.8 W
- per task: 87.6496 Wh · 2.6558 p (tariff 30.3 p/kWh)
- covers: run-0102
- aihydra nemotron3-super UD-IQ4_XS tau2-bench-airline 16-scored-task arm window (2026-08-10T09:07:30Z..2026-08-10T14:22:34Z, 18904s), taken from run-0102's started_at/ended_at. run-0102 predates matching run-meta instrumentation for this arm, so the window is reconstructed from transcript timestamps rather than a recorded meter window. Counter interpolated from raw ~10s-resolution HA history samples pulled fresh 2026-08-18 as part of the run-0101/run-0102/run-0111 batch: start 6.663011680 kWh, end 7.539507967 kWh. The transcript contains internal idle gaps between task groups and one unscored infrastructure-error simulation; the gaps are included end-to-end as wall-clock session energy. wh_per_task is Wh-per-CORRECT-ANSWER: wh_total / tasks_passed (10 of 16 scored tasks, run-0102.metrics) = 87.6496 Wh, 2.6558 pence/correct at the standing 30.3 p/kWh tariff. This arm ran an unpinned self-play user simulator (user_llm is the same nemotron3-super path as the agent), so it remains provisional per clm-0043.
eng-0177
- 2026-08-11 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the run-meta window) · window recorded
- wh_total 73.3464 · mean 154 W · baseline 10.1 W (box-idle-all-empty) · delta 143.9 W
- per task: 73.3464 Wh · 2.2224 p (tariff 30.3 p/kWh)
- covers: run-0111
- aihydra qwen35-122b task-9 regression arm window (2026-08-11T20:08:30Z..2026-08-11T20:37:05Z, 1715s), matching run-0111's recorded run-meta window. Counter interpolated from raw ~10s-resolution HA history samples pulled fresh 2026-08-18 as part of the run-0101/run-0102/run-0111 batch: start 11.444064340 kWh, end 11.517410772 kWh. wh_per_task is Wh-per-CORRECT-ANSWER: wh_total / tasks_passed (1 of 8 trials, run-0111.metrics) = 73.3464 Wh, 2.2224 pence/correct at the standing 30.3 p/kWh tariff. This is a narrow task-9 regression run with the pinned Haiku simulator, not a full field-comparison arm.
eng-0178
- 2026-08-18 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the recorded window) · window recorded
- wh_total 1.2185 · mean 146.23 W · baseline 10.1 W (box-idle-all-empty) · delta 136.13 W
- covers: run-0440
- DeepSeek-V4-Flash v0.6.4-fork Vulkan guard. Counter interpolated to the recorded 30-second window: 26.670128189 to 26.671346734 kWh. Short load-hot guard window recorded for completeness; it is not a throughput or Wh-per-correct measurement.
eng-0179
- 2026-08-18 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the recorded window) · window recorded
- wh_total 5.3741 · mean 58.63 W · baseline 10.1 W (box-idle-all-empty) · delta 48.53 W
- covers: run-0441 run-0442
- DeepSeek-V4-Flash v0.6.4-fork Vulkan d0 llama-bench cell, covering three fresh-process repetitions and both prefill/decode rows. Counter interpolated from 26.671430153 to 26.676804260 kWh. Wh total only: synthetic throughput has no correct-answer denominator. Mean draw includes setup/teardown around the shallow cell.
eng-0180
- 2026-08-18 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the recorded window) · window recorded
- wh_total 33.9372 · mean 139.95 W · baseline 10.1 W (box-idle-all-empty) · delta 129.85 W
- covers: run-0443 run-0444
- DeepSeek-V4-Flash v0.6.4-fork Vulkan d32768 llama-bench cell, covering three fresh-process repetitions and both prefill/decode rows. Counter interpolated from 26.676804260 to 26.710741461 kWh. Wh total only: synthetic throughput has no correct-answer denominator.
eng-0181
- 2026-08-18 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the recorded window) · window recorded
- wh_total 127.9076 · mean 166.72 W · baseline 10.1 W (box-idle-all-empty) · delta 156.62 W
- covers: run-0445 run-0446
- DeepSeek-V4-Flash v0.6.4-fork Vulkan d131072 llama-bench cell, covering three fresh-process repetitions and both prefill/decode rows. Counter interpolated from 26.717403614 to 26.845311182 kWh. Wh total only: synthetic throughput has no correct-answer denominator.
eng-0182
- 2026-08-18 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the recorded window) · window recorded
- wh_total 70.0781 · mean 173.51 W · baseline 10.1 W (box-idle-all-empty) · delta 163.41 W
- covers: run-0447 run-0448
- DeepSeek-V4-Flash v0.6.4-fork Vulkan d204800 llama-bench cell, covering its single repetition and both prefill/decode rows. Counter interpolated from 26.845311182 to 26.915389310 kWh. Wh total only: synthetic throughput has no correct-answer denominator.
eng-0183
- 2026-08-18 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the recorded window) · window recorded
- wh_total 94.2977 · mean 174.54 W · baseline 10.1 W (box-idle-all-empty) · delta 164.44 W
- covers: run-0449 run-0450
- DeepSeek-V4-Flash v0.6.4-fork Vulkan d262144 llama-bench cell, covering its single repetition and both prefill/decode rows. Counter interpolated from 26.915404177 to 27.009701854 kWh. Wh total only: synthetic throughput has no correct-answer denominator.
eng-0184
- 2026-08-19 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the recorded window) · window recorded
- wh_total 1.9689 · mean 78.76 W · baseline 10.1 W (box-idle-all-empty) · delta 68.66 W
- covers: run-0452 run-0453
- HO-002 plain-control whole window, including model startup, 15-prompt varied probe and 4/4 guard. Counter interpolated from 27.128520619 to 27.130489507 kWh. Wh total only: this activation/performance probe has no correct-answer denominator.
eng-0185
- 2026-08-19 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the recorded window) · window recorded
- wh_total 2.2954 · mean 64.06 W · baseline 10.1 W (box-idle-all-empty) · delta 53.96 W
- covers: run-0454 run-0455
- HO-002 MTP n_max=1 whole window (startup, 15-prompt probe and 4/4 guard). Counter interpolated from 27.130489507 to 27.132784865 kWh. Wh total only; differing startup/window lengths prevent treating this as token-level energy.
eng-0186
- 2026-08-19 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the recorded window) · window recorded
- wh_total 2.3259 · mean 94.08 W · baseline 10.1 W (box-idle-all-empty) · delta 83.98 W
- covers: run-0456 run-0457
- HO-002 MTP n_max=2 whole window (startup, 15-prompt probe and 4/4 guard). Counter interpolated from 27.132784865 to 27.135110808 kWh. Wh total only; differing startup/window lengths prevent treating this as token-level energy.
eng-0187
- 2026-08-19 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the recorded window) · window recorded
- wh_total 2.6385 · mean 104.38 W · baseline 10.1 W (box-idle-all-empty) · delta 94.28 W
- covers: run-0458 run-0459
- HO-002 MTP n_max=3 whole window (startup, 15-prompt probe and 4/4 guard). Counter interpolated from 27.135110808 to 27.137749330 kWh. Wh total only; differing startup/window lengths prevent treating this as token-level energy.
eng-0188
- 2026-08-16 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the recorded window) · window reconstructed
- wh_total 84.1797 · mean 99 W · baseline 10.1 W (box-idle-all-empty) · delta 88.9 W
- covers: run-0356
- Aggregate failure-window energy for both allowed Vulkan d131072 attempts. Exact attempt boundaries came from the preserved matrix timestamps and run-meta records: attempt 1 15:11:44Z..15:35:36Z and attempt 2 15:35:36Z..16:02:45Z, both rc=134 and contention=false. The combined counter endpoints were linearly interpolated from raw samples to 22.343875097 and 22.428054755 kWh. Both attempts ended in DeviceLostError and produced no throughput; this is failure-window Wh only, never throughput energy or Wh per correct answer.
eng-0189
- 2026-08-19 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, high-confidence raw ~10s-resolution samples linearly interpolated to the recorded window) · window recorded
- wh_total 1.2437 · mean 114.81 W · baseline 10.1 W (box-idle-all-empty) · delta 104.71 W
- covers: run-0460 run-0461
- HG-002 plain whole window. Counter interpolated from 27.221368793 to 27.222612542 kWh. Gross Wh total only; no correct-answer denominator.
eng-0190
- 2026-08-19 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, high-confidence raw ~10s-resolution samples linearly interpolated to the recorded window) · window recorded
- wh_total 1.3166 · mean 139.4 W · baseline 10.1 W (box-idle-all-empty) · delta 129.3 W
- covers: run-0462 run-0463
- HG-002 MTP n1 whole window. Counter interpolated from 27.222612542 to 27.223929099 kWh. Gross Wh total only; no correct-answer denominator.
eng-0191
- 2026-08-19 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, high-confidence raw ~10s-resolution samples linearly interpolated to the recorded window) · window recorded
- wh_total 1.2939 · mean 137 W · baseline 10.1 W (box-idle-all-empty) · delta 126.9 W
- covers: run-0464 run-0465
- HG-002 MTP n2 whole window. Counter interpolated from 27.223929099 to 27.225222996 kWh. Gross Wh total only; no correct-answer denominator.
eng-0192
- 2026-08-19 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, high-confidence raw ~10s-resolution samples linearly interpolated to the recorded window) · window recorded
- wh_total 1.5277 · mean 157.13 W · baseline 10.1 W (box-idle-all-empty) · delta 147.03 W
- covers: run-0466 run-0467
- HG-002 MTP n3 whole window. Counter interpolated from 27.225222996 to 27.226750695 kWh. Gross Wh total only; no correct-answer denominator.
eng-0193
- 2026-08-19 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s samples linearly interpolated to the recorded window) · window recorded
- wh_total 8.2682 · mean 106.69 W · baseline 10.1 W (box-idle-all-empty) · delta 96.59 W
- per task: 2.0671 Wh · 0.6263 p (tariff 30.3 p/kWh)
- covers: run-0468
- Counter interpolated from 27.238559043 to 27.246827271 kWh; four correct answers.
eng-0194
- 2026-08-19 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s samples linearly interpolated to the recorded window) · window recorded
- wh_total 9.6174 · mean 124.1 W · baseline 10.1 W (box-idle-all-empty) · delta 114 W
- per task: 2.4044 Wh · 0.7285 p (tariff 30.3 p/kWh)
- covers: run-0469
- Counter interpolated from 27.249618336 to 27.259235741 kWh; four correct answers.
eng-0195
- 2026-08-19 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s samples linearly interpolated to the recorded window) · window recorded
- wh_total 17.6954 · mean 140.01 W · baseline 10.1 W (box-idle-all-empty) · delta 129.91 W
- covers: run-0470
- Counter interpolated from 27.259235741 to 27.276931143 kWh. Gross window energy only: the arm is invalid and has no aggregate mean or correct-answer denominator.
eng-0196
- 2026-08-19 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s samples linearly interpolated to the recorded window) · window recorded
- wh_total 0.0197 · mean 35.49 W · baseline 10.1 W (box-idle-all-empty) · delta 25.39 W
- covers: run-0472
- Counter interpolated from 27.287945226 to 27.287964945 kWh. Gross energy for the failed two-second server/load/request-admission window only; the runner was invalid and produced no scientific retrieval result, so no per-correct metric exists.
eng-0197
- 2026-08-19 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s samples linearly interpolated to the recorded window) · window recorded
- wh_total 0.0463 · mean 41.63 W · baseline 10.1 W (box-idle-all-empty) · delta 31.53 W
- covers: run-0473
- Counter interpolated from 27.298206880 to 27.298253140 kWh across the exact four-second retrieval window. This is gross energy for a valid failed retrieval cell; no correct answer exists, so no per-correct metric is defined.
eng-0198
- 2026-08-19 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s samples linearly interpolated to the recorded window) · window recorded
- wh_total 69.13 · mean 171.75 W · baseline 10.1 W (box-idle-all-empty) · delta 161.65 W
- covers: run-0492 run-0493
- Counter interpolated from 27.597717078 to 27.666847088 kWh across the exact HO-006 admitted Nemotron-3-Super Vulkan d204800 throughput window. Gross energy only: this llama-bench window covers paired prefill/decode rows and has no correct-answer denominator.
eng-0199
- 2026-08-20 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s samples linearly interpolated to the recorded window) · window recorded
- wh_total 0.1788 · mean 80.44 W · baseline 10.1 W (box-idle-all-empty) · delta 70.34 W
- covers: run-0494
- Counter interpolated from 27.672228684 to 27.672407449 kWh across the exact HO-005 Ornith UD-Q4_K_XL r3 guard window. Gross guard energy only; no correct-answer denominator is defined.
eng-0200
- 2026-08-20 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s samples linearly interpolated to the recorded window) · window recorded
- wh_total 16.0646 · mean 117.55 W · baseline 10.1 W (box-idle-all-empty) · delta 107.45 W
- per task: 3.2129 Wh · 0.0974 p (tariff 30.3 p/kWh)
- covers: run-0495
- Counter interpolated from 27.672407449 to 27.688472081 kWh across the exact HO-005 Ornith UD-Q4_K_XL r3 5-task smoke window. The run passed 5/5, so wh_per_task and pence_per_correct_answer use five correct tasks as the bounded smoke denominator; this is not a full-26 capability energy claim.
eng-0201
- 2026-08-20 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s samples linearly interpolated to the recorded window) · window recorded
- wh_total 61.9776 · mean 161.68 W · baseline 10.1 W (box-idle-all-empty) · delta 151.58 W
- covers: run-0496 run-0497
- Counter interpolated from 27.688510932 to 27.750488504 kWh across the exact HO-005 Ornith UD-Q4_K_XL r3 d262144 llama-bench window. Gross throughput energy only: this paired prefill/decode window has no correct-answer denominator.
eng-0202
- 2026-08-20 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s samples linearly interpolated to the recorded window) · window recorded
- wh_total 0.1888 · mean 75.53 W · baseline 10.1 W (box-idle-all-empty) · delta 65.43 W
- covers: run-0498
- Counter interpolated from 27.780810308 to 27.780999128 kWh across the exact HO-005 Ornith UD-Q4_K_XL full-26 successor guard window immediately before tau2. Guard energy only: this 4/4 guard has no correct-answer denominator.
eng-0203
- 2026-08-20 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s samples linearly interpolated to the recorded window) · window recorded
- wh_total 146.5378 · mean 126.05 W · baseline 10.1 W (box-idle-all-empty) · delta 115.95 W
- per task: 6.6608 Wh · 2.0182 p (tariff 30.3 p/kWh)
- covers: run-0499
- Counter interpolated from 27.780999128 to 27.927536928 kWh across the exact HO-005 Ornith UD-Q4_K_XL full-26 tau2 successor window. The run passed 22/26, so wh_per_task and pence_per_correct_answer use 22 correct tasks as the measured Q4_K_XL full-26 denominator. This is not Q8 inheritance and not a Q4/Q8 equivalence or energy-ranking claim.
eng-0204
- 2026-08-20 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s samples linearly interpolated to the recorded window) · window recorded
- wh_total 0.1023 · mean 33.49 W · baseline 10.1 W (box-active-ho009-window) · delta 23.39 W
- covers: run-0523
- Counter interpolated from 29.409234720 to 29.409337046 kWh across HO-009 mtp-context-proof n1 window (8 HA samples in widened query). Wh total only; no Wh/correct because HO-009 is activation/safety/performance evidence, not a capability suite.
eng-0205
- 2026-08-20 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s samples linearly interpolated to the recorded window) · window recorded
- wh_total 0.8248 · mean 95.78 W · baseline 10.1 W (box-active-ho009-window) · delta 85.68 W
- covers: run-0524
- Counter interpolated from 29.409337046 to 29.410161840 kWh across HO-009 stageA plain window (10 HA samples in widened query). Wh total only; no Wh/correct because HO-009 is activation/safety/performance evidence, not a capability suite.
eng-0206
- 2026-08-20 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s samples linearly interpolated to the recorded window) · window recorded
- wh_total 0.829 · mean 119.37 W · baseline 10.1 W (box-active-ho009-window) · delta 109.27 W
- covers: run-0525
- Counter interpolated from 29.410161840 to 29.410990833 kWh across HO-009 stageA n1 window (10 HA samples in widened query). Wh total only; no Wh/correct because HO-009 is activation/safety/performance evidence, not a capability suite.
eng-0207
- 2026-08-20 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s samples linearly interpolated to the recorded window) · window recorded
- wh_total 0.7408 · mean 98.78 W · baseline 10.1 W (box-active-ho009-window) · delta 88.68 W
- covers: run-0526
- Counter interpolated from 29.410990833 to 29.411731664 kWh across HO-009 stageA n2 window (9 HA samples in widened query). Wh total only; no Wh/correct because HO-009 is activation/safety/performance evidence, not a capability suite.
eng-0208
- 2026-08-20 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s samples linearly interpolated to the recorded window) · window recorded
- wh_total 0.8241 · mean 109.87 W · baseline 10.1 W (box-active-ho009-window) · delta 99.77 W
- covers: run-0527
- Counter interpolated from 29.411731664 to 29.412555718 kWh across HO-009 stageA n3 window (10 HA samples in widened query). Wh total only; no Wh/correct because HO-009 is activation/safety/performance evidence, not a capability suite.
eng-0209
- 2026-08-20 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s samples linearly interpolated to the recorded window) · window recorded
- wh_total 0.4477 · mean 115.11 W · baseline 10.1 W (box-active-ho009-window) · delta 105.01 W
- covers: run-0528
- Counter interpolated from 29.412555718 to 29.413003373 kWh across HO-009 guard plain window (8 HA samples in widened query). Wh total only; no Wh/correct because HO-009 is activation/safety/performance evidence, not a capability suite.
eng-0210
- 2026-08-20 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s samples linearly interpolated to the recorded window) · window recorded
- wh_total 0.3656 · mean 77.42 W · baseline 10.1 W (box-active-ho009-window) · delta 67.32 W
- covers: run-0529
- Counter interpolated from 29.413003373 to 29.413368972 kWh across HO-009 guard n2 window (9 HA samples in widened query). Wh total only; no Wh/correct because HO-009 is activation/safety/performance evidence, not a capability suite.
eng-0211
- 2026-08-20 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s samples linearly interpolated to the recorded window) · window recorded
- wh_total 0.445 · mean 133.5 W · baseline 10.1 W (box-active-ho009-window) · delta 123.4 W
- covers: run-0530
- Counter interpolated from 29.413368972 to 29.413813968 kWh across HO-009 bench-plain d0 window (8 HA samples in widened query). Wh total only; no Wh/correct because HO-009 is activation/safety/performance evidence, not a capability suite.
eng-0212
- 2026-08-20 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s samples linearly interpolated to the recorded window) · window recorded
- wh_total 3.0244 · mean 187.72 W · baseline 10.1 W (box-active-ho009-window) · delta 177.62 W
- covers: run-0531
- Counter interpolated from 29.413813968 to 29.416838350 kWh across HO-009 bench-plain d32768 window (13 HA samples in widened query). Wh total only; no Wh/correct because HO-009 is activation/safety/performance evidence, not a capability suite.
eng-0213
- 2026-08-20 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s samples linearly interpolated to the recorded window) · window recorded
- wh_total 17.7607 · mean 185.87 W · baseline 10.1 W (box-active-ho009-window) · delta 175.77 W
- covers: run-0532
- Counter interpolated from 29.416838350 to 29.434599075 kWh across HO-009 bench-plain d131072 window (42 HA samples in widened query). Wh total only; no Wh/correct because HO-009 is activation/safety/performance evidence, not a capability suite.
eng-0214
- 2026-08-20 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s samples linearly interpolated to the recorded window) · window recorded
- wh_total 34.8998 · mean 180.26 W · baseline 10.1 W (box-active-ho009-window) · delta 170.16 W
- covers: run-0533
- Counter interpolated from 29.434599075 to 29.469498864 kWh across HO-009 bench-plain d204800 window (76 HA samples in widened query). Wh total only; no Wh/correct because HO-009 is activation/safety/performance evidence, not a capability suite.
eng-0215
- 2026-08-20 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s samples linearly interpolated to the recorded window) · window recorded
- wh_total 53.1943 · mean 180.15 W · baseline 10.1 W (box-active-ho009-window) · delta 170.05 W
- covers: run-0534
- Counter interpolated from 29.469498864 to 29.522693180 kWh across HO-009 bench-plain d262144 window (114 HA samples in widened query). Wh total only; no Wh/correct because HO-009 is activation/safety/performance evidence, not a capability suite.
eng-0216
- 2026-08-20 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s samples linearly interpolated to the recorded window) · window recorded
- wh_total 0.3653 · mean 93.94 W · baseline 10.1 W (box-active-ho009-window) · delta 83.84 W
- covers: run-0535
- Counter interpolated from 29.522693180 to 29.523058521 kWh across HO-009 server-perf-n2 d0 window (8 HA samples in widened query). Wh total only; no Wh/correct because HO-009 is activation/safety/performance evidence, not a capability suite.
eng-0217
- 2026-08-20 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s samples linearly interpolated to the recorded window) · window recorded
- wh_total 6.4222 · mean 182.05 W · baseline 10.1 W (box-active-ho009-window) · delta 171.95 W
- covers: run-0536
- Counter interpolated from 29.523058521 to 29.529480702 kWh across HO-009 server-perf-n2 d32768 window (20 HA samples in widened query). Wh total only; no Wh/correct because HO-009 is activation/safety/performance evidence, not a capability suite.
eng-0218
- 2026-08-20 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s samples linearly interpolated to the recorded window) · window recorded
- wh_total 46.9044 · mean 178.87 W · baseline 10.1 W (box-active-ho009-window) · delta 168.77 W
- covers: run-0537
- Counter interpolated from 29.529480702 to 29.576385063 kWh across HO-009 server-perf-n2 d131072 window (101 HA samples in widened query). Wh total only; no Wh/correct because HO-009 is activation/safety/performance evidence, not a capability suite.
eng-0219
- 2026-08-20 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s samples linearly interpolated to the recorded window) · window recorded
- wh_total 33.9028 · mean 178.7 W · baseline 10.1 W (box-active-ho009-window) · delta 168.6 W
- covers: run-0538
- Counter interpolated from 29.576385063 to 29.610287909 kWh across HO-009 server-perf-n2 d204800 window (75 HA samples in widened query). Wh total only; no Wh/correct because HO-009 is activation/safety/performance evidence, not a capability suite.
eng-0220
- 2026-08-20 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s samples linearly interpolated to the recorded window) · window recorded
- wh_total 0.1308 · mean 52.3 W · baseline 10.1 W (box-active-ho009-window) · delta 42.2 W
- covers: run-0539
- Counter interpolated from 29.610287909 to 29.610418664 kWh across HO-009 server-perf-n2 d262144 window (8 HA samples in widened query). Wh total only; no Wh/correct because HO-009 is activation/safety/performance evidence, not a capability suite.
eng-0221
- 2026-08-20 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the recorded window; power from sensor.hardware_ai_hydra_power where available) · window recorded
- wh_total 96.7822 · mean 175.81 W · baseline 10.1 W (box-idle-all-empty) · delta 165.71 W
- covers: run-0540
- HO-001 v0.6.6 DeepSeek quant-KV sparse boundary llama-bench d262144 cell for cfg-0134. Counter interpolated from 28.931674412 to 29.028456651 kWh across the exact run-meta window from /home/aihydra/bench-results/ho001-deepseek-v066-quantkv-sparse-r1/run-meta-anchors-final.jsonl. Wh total only: this performance/sparse-path boundary has no correct-answer denominator and no q8_0/q4_0 quality-equivalence claim.
eng-0222
- 2026-08-20 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples linearly interpolated to the recorded window; power from sensor.hardware_ai_hydra_power where available) · window recorded
- wh_total 16.3303 · mean 142.76 W · baseline 10.1 W (box-idle-all-empty) · delta 132.66 W
- covers: run-0541
- HO-004 negative Stage A plain/control window. Counter interpolated from 29.620174110 to 29.636504455 kWh across the exact reviewer-admitted stageA window; 66 energy samples in the widened query and 39 power samples inside the window. Wh total only: this was a rejected sentinel-gate diagnostic, not a capability or throughput-energy result.
eng-0223
- 2026-08-20 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples linearly interpolated to the recorded window; power from sensor.hardware_ai_hydra_power where available) · window recorded
- wh_total 9.8588 · mean 164.64 W · baseline 10.1 W (box-idle-all-empty) · delta 154.54 W
- covers: run-0542
- HO-004 negative Stage A stock MTP n_max=3 window. Counter interpolated from 29.636657087 to 29.646515865 kWh across the exact reviewer-admitted stageA window; 47 energy samples in the widened query and 21 power samples inside the window. Wh total only: this was a rejected sentinel-gate diagnostic, not a capability or throughput-energy result.
eng-0224
- 2026-08-20 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples linearly interpolated to the recorded window; power from sensor.hardware_ai_hydra_power where available) · window recorded
- wh_total 9.5089 · mean 186.81 W · baseline 10.1 W (box-idle-all-empty) · delta 176.71 W
- covers: run-0543
- HO-004 negative Stage A DFlash2 Q8_0 width-4 window. Counter interpolated from 29.646683073 to 29.656192004 kWh across the exact reviewer-admitted stageA window; 43 energy samples in the widened query and 18 power samples inside the window. Wh total only: this was a rejected sentinel-gate diagnostic, not a capability, throughput, or DFlash2 energy-efficiency result.
eng-0225
- 2026-08-20 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy; raw ~10s-resolution history samples queried from hborchestrator using retained HA history) · window recorded
- wh_total 5.3059 · mean 48.09 W · baseline 11.4 W (box-idle-all-empty) · delta 36.69 W
- covers: run-0544
- HG-004 Phase-1 GLM diagnostic window. Counter history query returned 43 energy samples from 29.691154021 kWh at 2026-08-20T23:08:00Z to 29.696459951 kWh at 2026-08-20T23:14:59.720295Z, within the runner window ending 23:15:04Z. Power history returned 39 samples with mean 48.09 W and peak 149.13 W. Wh total only: this was a narrow non-scoring response-capture diagnostic, not a tau2 score, throughput, Wh/correct, capability, or production-efficiency result.
eng-0226
- 2026-08-22 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-resolution history samples, linearly interpolated to the run window) · window recorded
- wh_total 62.1014 · mean 157.33 W · baseline 10.1 W (box-idle-all-empty) · delta 147.23 W
- per task: 12.4203 Wh · 0.3763 p (tariff 30.3 p/kWh)
- covers: run-0553
- LLaDA2.2-flash tau2 airline 5-task SCREEN (run-0553): block-diffusion on the headbouyJB/diffuse-cpp fork (GPU-resident KV cache + full Levenshtein editing), GPU-accelerated on aihydra (gfx1151), pinned Haiku-4.5 simulator, all 5 tasks passed. Counter interpolated from raw ~10s HA history: start 32.737790577 kWh, end 32.799891990 kWh, delta 0.062101413 kWh = 62.1014 Wh. wh_per_task is Wh per CORRECT answer: wh_total / tasks_passed (5 of 5). Whole-session energy includes simulator wait time end to end, matching the site's headline denominator. Retroactive join (the run was not harness-wrapped) within the ~10-day HA retention window; the window was verified contention-free — a dedicated diffuse-server was the only GPU consumer and the counter is smooth and monotonic. This is a 5-task screen, not the full 26-task arm: at 12.42 Wh per correct answer it is ~2.3x Ling-3.0-flash's 5.32 Wh per correct on its full run, an early cross-paradigm indicator (diffusion is compute-heavy per committed token), not a settled comparison (different N and task mix).
eng-0227
- 2026-08-22 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s samples queried by hborchestrator via Nabu Casa API, idle baseline 10.1 W) · window recorded
- wh_total 254.44 · mean 163.08 W · baseline 10.1 W (idle-ho003-patched-throughput) · delta 153 W
- covers: run-0562 run-0563 run-0564 run-0565 run-0566
- PATCHED q8_0 throughput ladder aggregate (run-0562..0566). Counter 31.682629 -> 31.952827 kWh = 270.20 Wh total, 15.76 Wh idle -> 254.44 Wh active. Build ce7689f, q8_0/q8_0 KV. Slightly MORE efficient than the f16 control ladder (254.44 vs 272.84 Wh). Wh-total only (energy join yaml patched-throughput).
eng-0228
- 2026-08-21 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s samples queried by hborchestrator via Nabu Casa API, idle baseline 10.1 W) · window recorded
- wh_total 420.58 · mean 153.17 W · baseline 10.1 W (idle) · delta 143.07 W
- per task: — Wh · 1.274 p (tariff 30.3 p/kWh)
- covers: run-0555
- CONTROL arm tau2 full-26 tau2 (run-0555) energy window. Counter 30.455198 -> 30.903512 kWh = 448.31 Wh total, 27.73 Wh idle -> 420.58 Wh active over 9886 s. 10 correct tasks (mean 0.40) => **42.06 Wh/correct**. Idle-subtracted Wh/correct is a joined capability-energy metric (energy join yaml control-tau2). Control Wh/correct is the quant-KV baseline the patched arm beats.
eng-0229
- 2026-08-22 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s samples queried by hborchestrator via Nabu Casa API, idle baseline 10.1 W) · window recorded
- wh_total 448.76 · mean 149.97 W · baseline 10.1 W (idle) · delta 139.87 W
- per task: — Wh · 0.906 p (tariff 30.3 p/kWh)
- covers: run-0556
- PATCHED tau2 full-26 (run-0556) energy window. Counter 31.203646 -> 31.682629 kWh = 478.98 Wh total, 30.23 Wh idle -> 448.76 Wh active over 10774 s. 15 correct (run-0556) => **29.92 Wh/correct**. Patched q8_0 recovers +5 correct tasks and is **29% better energy per correct answer** (29.92 vs 42.06 Wh/correct, eng-0228): more correct answers at only 7% more active energy than control. (energy join yaml patched-tau2).
eng-0230
- 2026-08-22 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s samples linearly interpolated to the recorded window) · window recorded
- wh_total 83.39 · mean 137.7 W · baseline 10.1 W (box-idle-all-empty) · delta 127.6 W
- per task: 3.2073 Wh · 0.4211 p (tariff 30.3 p/kWh)
- covers: run-0569
- HO-009-AB PLAIN arm tau2 full-26 (reasoning OFF). Counter interpolated from 33.100245 to 33.183670 kWh across the plain tau2 window 20:57:25-21:34:00Z on 2026-08-22. 6/26 tasks passed, so wh_per_task uses 26 and pence_per_correct_answer uses 6 correct tasks as the denominator. Power profile shows a single isolated active band (idle ~11.5W before 20:55Z and after 21:34Z) — no contention with the prior LLaDA external job, which had already released. Build 7077abb, ROCm0/gfx1151, f16/f16 KV, -rea off. Plain-vs-n2 capability comparison only; no production or throughput claim.
eng-0231
- 2026-08-22 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s samples linearly interpolated to the recorded window) · window recorded
- wh_total 49.65 · mean 161.03 W · baseline 10.1 W (box-idle-all-empty) · delta 150.93 W
- per task: 1.9096 Wh · — p (tariff 30.3 p/kWh)
- covers: run-0571
- HO-009-AB n2/MTP arm tau2 full-26 (reasoning OFF). Counter interpolated from 33.202848 to 33.252787 kWh across the n2 tau2 window 23:05:23-23:24:06Z on 2026-08-22 (first task start to last task end, from tau2-full26-summary.json). 0/26 tasks passed, so pence_per_correct_answer is null (no correct-task denominator; a zero-count capability cell). wh_per_task uses 26. Power profile shows a single isolated active band (idle ~11.5W before 23:05Z and after 23:24Z) under the HO-009-AB claim, no contention. Build 7077abb, ROCm0/gfx1151, f16/f16 KV, --spec-type draft-mtp --spec-draft-n-max 2, -rea off. Plain-vs-n2 capability comparison only; the higher mean_w_active reflects the same idle baseline subtracted and a slightly shorter wall window, not a throughput claim.
eng-0232
- 2026-08-23 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-history samples queried by hborchestrator via Nabu Casa API and linearly interpolated to the recorded window; power from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 16.9032 · mean 142.15 W · baseline 10.1 W (box-idle-all-empty) · delta 132.05 W
- covers: run-0572
- HO-004-DIAG Stage-A plain/control sentinel window (run-0572 / cfg-0146). Counter interpolated from 33.291573 to 33.308476 kWh across the recorded 428s arm window (02:27:15.904973Z-02:34:23.984284Z); 40 power samples inside the window, mean 142.15 W, peak 217.91 W. Wh-total only: sentinel/gate arm, not a capability or throughput result — no Wh/correct, no Wh/task anywhere for this null-update diagnostic (run-0572-0574 and clm-0109 already assert this). Build 2586f6edd, c32768 plain/control with no spec draft.
eng-0233
- 2026-08-23 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-history samples queried by hborchestrator via Nabu Casa API and linearly interpolated to the recorded window; power from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 11.5626 · mean 177.97 W · baseline 10.1 W (box-idle-all-empty) · delta 167.87 W
- covers: run-0573
- HO-004-DIAG Stage-A stock-MTP sentinel window (run-0573 / cfg-0147). Counter interpolated from 33.308641 to 33.320204 kWh across the recorded 234s arm window (02:34:29.993049Z-02:38:23.887888Z); 22 power samples inside the window, mean 177.97 W, peak 206.22 W. Wh-total only: sentinel/gate arm, not a capability or throughput result — no Wh/correct, no Wh/task anywhere for this null-update diagnostic (run-0572-0574 and clm-0109 already assert this). Build 2586f6edd, c32768, --spec-type draft-mtp --spec-draft-n-max 3.
eng-0234
- 2026-08-23 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-history samples queried by hborchestrator via Nabu Casa API and linearly interpolated to the recorded window; power from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 9.9923 · mean 182.34 W · baseline 10.1 W (box-idle-all-empty) · delta 172.24 W
- covers: run-0574
- HO-004-DIAG Stage-A DFlash2 sentinel window (run-0574 / cfg-0148). Counter interpolated from 33.320370 to 33.330363 kWh across the recorded 197s arm window (02:38:29.899941Z-02:41:47.186671Z); 20 power samples inside the window, mean 182.34 W, peak 202.5 W. Wh-total only: sentinel/gate arm, not a capability or throughput result — no Wh/correct, no Wh/task anywhere for this null-update diagnostic (run-0572-0574 and clm-0109 already assert this). Build 2586f6edd, c32768, --spec-type draft-dflash DFlash2-Q8_0 --spec-draft-n-max 3.
eng-0235
- 2026-08-23 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-history samples queried by hborchestrator via Nabu Casa API and linearly interpolated to the recorded window; power from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 76.9601 · mean 122.99 W · baseline 10.1 W (box-idle-all-empty) · delta 112.89 W
- per task: 2.96 Wh · 0.106 p (tariff 30.3 p/kWh)
- covers: run-0576
- HO-009-v0610 PLAIN arm tau2 full-26 (reasoning OFF) on the v0.6.10 fork/Vulkan stack. Counter interpolated from 33.364054 to 33.441014 kWh across the plain tau2 window 03:50:27.640-04:28:01.945Z on 2026-08-23. 22/26 tasks passed, so wh_per_task uses 26 and pence_per_correct uses 22 correct tasks. Single isolated active band (idle floor 10.1W, power samples n=225), no contention. cfg-0149 / build 2586f6edd Vulkan0 from cfg-0144 on v0.6.10. Capability comparison only; no production/throughput claim.
eng-0236
- 2026-08-23 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-history samples queried by hborchestrator via Nabu Casa API and linearly interpolated to the recorded window; power from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 53.0968 · mean 131.33 W · baseline 10.1 W (box-idle-all-empty) · delta 121.23 W
- per task: 2.0422 Wh · 0.0946 p (tariff 30.3 p/kWh)
- covers: run-0578
- HO-009-v0610 n2/MTP arm tau2 full-26 (reasoning OFF) on the v0.6.10 fork/Vulkan stack. Counter interpolated from 33.441373 to 33.494470 kWh across the n2 tau2 window 04:28:18.406-04:52:36.488Z on 2026-08-23. 17/26 tasks passed, so wh_per_task uses 26 and pence_per_correct uses 17 correct tasks. Single isolated active band (idle floor 10.1W, power samples n=146), no contention. cfg-0150 / build 2586f6edd Vulkan0 from cfg-0145 on v0.6.10. Capability comparison only; no production/throughput claim.
eng-0237
- 2026-08-23 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-history samples queried by hborchestrator via Nabu Casa API and linearly interpolated to the recorded window; power from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 2.9378 · mean 76.73 W · baseline 10.1 W (box-idle-all-empty) · delta 66.63 W
- covers: run-0579
- HO-012 Cell-1 baseline c1-b2048-ub512 (q8_0/q8_0 KV, build ce7689f ROCm0) latency arm. Counter interpolated 33.5040978 -> 33.5070357 kWh across the server+guard+probe window 05:23:51-05:26:08Z. 13 power samples, no contention. cfg-0151; latency result in run-0579. Production-optimisation matrix cell only, no capability/throughput claim.
eng-0238
- 2026-08-23 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-history samples queried by hborchestrator via Nabu Casa API and linearly interpolated to the window; power from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 2.0998 · mean 67.75 W · baseline 10.1 W (box-idle-all-empty) · delta 57.65 W
- covers: run-0580
- HO-012 Cell-1 arm b1024/ub512 (cfg-0152, q8_0/q8_0 KV, build ce7689f ROCm0). Counter 33.50709 -> 33.50919 kWh over 05:26:10-05:28:05Z. 12 power samples, no contention. latency run-0580. Matrix cell only.
eng-0239
- 2026-08-23 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-history samples queried by hborchestrator via Nabu Casa and interpolated to the window; power from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 2.0531 · mean 64.6 W · baseline 10.1 W (box-idle-all-empty) · delta 54.5 W
- covers: run-0581
- HO-012 Cell-1 arm b512/ub512 (cfg-0153, q8_0/q8_0 KV, build ce7689f ROCm0). Counter 33.509254 -> 33.511307 kWh over 05:28:08-05:29:48Z. 9 power samples, no contention. run-0581 latency arm.
eng-0240
- 2026-08-23 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, sampled and interpolated to the window; power from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 2.1011 · mean 77.25 W · baseline 10.1 W (box-idle-all-empty) · delta 67.15 W
- covers: run-0582
- HO-012 Cell-1 arm b512/ub256 (cfg-0154, q8_0/q8_0 KV, build ce7689f ROCm0). Counter 33.511364 -> 33.513465 kWh over the window. 10 power samples, no contention. run-0582 latency arm.
eng-0241
- 2026-08-23 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, sampled and interpolated to the window; power from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 2.3934 · mean 82.18 W · baseline 10.1 W (box-idle-all-empty) · delta 72.08 W
- covers: run-0583
- HO-012 Cell-1 worst-ttft arm b512/ub128 (cfg-0155, q8_0/q8_0 KV, build ce7689f ROCm0). Counter 33.513556 -> 33.515950 kWh over the window. 10 power samples, no contention. run-0583 latency arm.
eng-0242
- 2026-08-23 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, sampled and interpolated to the window; power from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 2.309 · mean 69.84 W · baseline 10.1 W (box-idle-all-empty) · delta 59.74 W
- covers: run-0584
- HO-012 Cell-1 arm b256/ub256 (cfg-0156, q8_0/q8_0 KV, build ce7689f ROCm0). Counter 33.516005 -> 33.518314 kWh over the window. 12 power samples, no contention. run-0584 latency arm.
eng-0243
- 2026-08-23 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, sampled and interpolated to the window; power from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 144.3385 · mean 179.95 W · baseline 10.1 W (box-idle-all-empty) · delta 169.85 W
- covers: run-0585
- HO-012 Cell-2 depth d262144 (cfg-0151 -> b2048/ub512 q8_0/q8_0 KV, build ce7689f ROCm0) llama-bench over the 48-min window 05:35:15-06:23:18Z. The largest energy cell in the matrix (144.3 Wh, sustained ~180W, peak 208W) because it is a full prefill+decode at 262144 depth. pp16384 62.25 tok/s, tg256 8.448 t/s. run-0585. No capability/throughput-per-correct claim.
eng-0244
- 2026-08-23 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, sampled and interpolated to the window; power from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 3.1818 · mean 102.56 W · baseline 10.1 W (box-idle-all-empty) · delta 92.46 W
- covers: run-0586
- HO-012 Cell-3 f16/f16 KV arm (cfg-0157, b2048/ub512, build ce7689f ROCm0). Counter 33.662714 -> 33.665896 kWh over 06:23:19-06:25:12Z. 12 power samples, no contention. run-0586 f16 comparison anchor.
eng-0245
- 2026-08-23 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, sampled and interpolated to the window; power from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 3.1337 · mean 85.57 W · baseline 10.1 W (box-idle-all-empty) · delta 75.47 W
- covers: run-0587
- HO-012 Cell-3 q4_0/q4_0 KV arm (cfg-0158, b2048/ub512, build ce7689f ROCm0). Counter 33.665920 -> 33.669054 kWh over the window. 13 power samples, no contention. run-0587 latency arm (q4_0 measured ~0.4% slower than the f16 anchor at tpot; recorded in run-0587/cll-0111, never as equivalence).
eng-0246
- 2026-08-23 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, sampled and interpolated to the window; power from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 0.0749 · mean 42.37 W · baseline 10.1 W (box-idle-all-empty) · delta 32.27 W
- covers: run-0591
- HO-005-REVISED guard window (run-0591, cfg-0159, build 2586f6ed Vulkan v0.6.10). Counter 33.674558 -> 33.674633 kWh over ~5s. 1 power sample. Guard windows are sub-minute warm-ups; carried for completeness, never as an energy ranking (no tau2).
eng-0247
- 2026-08-23 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, sampled and interpolated to the window; power from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 13.2142 · mean 181.16 W · baseline 10.1 W (box-idle-all-empty) · delta 171.06 W
- covers: run-0592
- HO-005-REVISED control arm b512/ub512 d131072 (cfg-0159, build 2586f6ed Vulkan). Counter 33.674757 -> 33.687971 kWh over 258s. 25 power samples, no contention. run-0592 = 34.357 tps. No Wh/correct (no tau2).
eng-0248
- 2026-08-23 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, sampled and interpolated to the window; power from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 12.9191 · mean 174.75 W · baseline 10.1 W (box-idle-all-empty) · delta 164.65 W
- covers: run-0593
- HO-005-REVISED b1024/ub512 d131072 (cfg-0159, build 2586f6ed Vulkan). Counter 33.688031 -> 33.700950 kWh over 263s. 26 power samples, no contention. run-0593 = 34.430 tps (numeric max, still inside the flat 0.26% matrix). No fastest-config claim (clm-0112). No Wh/correct (no tau2).
eng-0249
- 2026-08-23 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, sampled and interpolated to the window; power from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 12.4689 · mean 176.97 W · baseline 10.1 W (box-idle-all-empty) · delta 166.87 W
- covers: run-0594
- HO-005-REVISED b2048/ub1024 d131072 (cfg-0159, build 2586f6ed Vulkan). Counter 33.701001 -> 33.713469 kWh over 254s. 23 power samples, no contention. run-0594 = 34.363 tps. No Wh/correct (no tau2). Mid-range numeric within the flat spread (clm-0112).
eng-0250
- 2026-08-23 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, sampled and interpolated to the window; power from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 13.0222 · mean 177.73 W · baseline 10.1 W (box-idle-all-empty) · delta 167.63 W
- covers: run-0595
- HO-005-REVISED b4096/ub2048 d131072 (cfg-0159, build 2586f6ed Vulkan). Counter 33.713525 -> 33.726557 kWh over 266s. 24 power samples, no contention. run-0595 = 34.342 tps (numeric minimum, still flat). No Wh/correct (no tau2).
eng-0251
- 2026-08-23 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-history samples queried by hborchestrator via Nabu Casa API and linearly interpolated to the recorded window; power from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 57.1737 · mean 141.45 W · baseline 10.1 W (box-idle-all-empty) · delta 131.35 W
- per task: 2.199 Wh · 0.1019 p (tariff 30.3 p/kWh)
- covers: run-0599
- HO-013 Cell-1 n4-off arm tau2 full-26 (reasoning OFF, MTP n_max=4) on the v0.6.10 fork/Vulkan stack (2586f6edd). Counter interpolated from 33.772208 to 33.829381 kWh across the n4 tau2 window 07:29:04.077901-07:53:21.027156Z on 2026-08-23 (1456.9s, n_power=144). 17/26 tasks passed, so wh_per_task uses 26 and pence_per_correct uses 17 correct tasks. Single isolated active band (idle floor 10.1W), no contention. cfg-0160 / build 2586f6edd Vulkan/RADV on v0.6.10. Compare arms already wired: eng-0235 (plain tau2 76.96 Wh), eng-0236 (n2 tau2 53.10 Wh). Capability comparison only; no production/throughput claim. Window within HA retention (~10d; run ended 07:53Z, joined same day).
eng-0252
- 2026-08-23 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-history samples queried by hborchestrator via the Nabu Casa API and linearly interpolated to the window; power from sensor.hardware_ai_hydra_power) · window reconstructed
- wh_total 207.7 · mean 180.5 W · baseline 10.1 W (box-idle-all-empty) · delta 170.4 W
- covers: run-0602 run-0603 run-0604
- HK-001-FULLBENCH FIT + depth-ladder engine band. Counter 34.3666 -> 34.5739 kWh over 14:19-15:28Z (207.7 Wh), the continuous active band covering the FIT c32768 cell (run-0602, pp=279.6/tg=11.73) and the retained ENGINE-path depth-ladder rungs d0/d32768 (run-0603..0604, llama-bench with the PROMPTFORGE sidecar env). Window RECONSTRUCTED from the HA power/energy history (the runner's per-run timestamps live in on-box logs not reachable from this session; retention ~10d so still joinable): I inferred the band edges (active startup 14:19, idle drop 15:28) from the sustained ~180 W draw, and attributed the one continuous engine band to the retained FIT/depth set rather than splitting per-rung. Single isolated active band (idle floor 10.1 W), no contention. cfg-0164 / TheRock 7.14 build e1da26bb8. No Wh/correct (no tau2 in this retained FIT/depth band).
eng-0253
- 2026-08-23 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-history samples queried by hborchestrator via the Nabu Casa API and linearly interpolated to the window; power from sensor.hardware_ai_hydra_power) · window reconstructed
- wh_total 8 · mean 40 W · baseline 10.1 W (box-idle-all-empty) · delta 29.9 W
- HK-001-FULLBENCH aggregate window detached from run records. Counter 34.5739 -> 34.5819 kWh over 15:28-15:41Z (8.0 Wh), but the former window covered both the run-0601 guard and an unresolved served-sweep interlude. It cannot be allocated solely to the guard without inventing a per-run split, so no run association or Wh allocation is retained here. Window RECONSTRUCTED from HA history; no contention. cfg-0164 / TheRock 7.14 build e1da26bb8.
eng-0254
- 2026-08-23 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-history samples queried by hborchestrator via the Nabu Casa API and linearly interpolated to the window; power from sensor.hardware_ai_hydra_power) · window reconstructed
- wh_total 284.4 · mean 162.4 W · baseline 10.1 W (box-idle-all-empty) · delta 152.3 W
- Reconstructed aggregate hardware window retained without a run association: no local Kairic run ID resolves this former HK-001-FULLBENCH full-26 window. It is not a Kairic energy claim, metric, or allocation. No contention within the window. cfg-0164 / TheRock 7.14 build e1da26bb8.
eng-0255
- 2026-08-23 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-history samples queried by hborchestrator via the Nabu Casa API and linearly interpolated to the window; power from sensor.hardware_ai_hydra_power) · window reconstructed
- wh_total 208.3 · mean 140.7 W · baseline 10.1 W (box-idle-all-empty) · delta 130.6 W
- Reconstructed aggregate hardware window retained without a run association: no local Kairic run ID resolves this former HK-001-FULLBENCH smoke window. It is not a Kairic energy claim, metric, or allocation; no per-task value is retained. No contention. cfg-0164 / TheRock 7.14 build e1da26bb8.
eng-0256
- 2026-08-23 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-history samples queried by hborchestrator via Nabu Casa API and linearly interpolated to the recorded window; power from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 13.4443 · mean 183.61 W · baseline 10.1 W (box-idle-all-empty) · delta 173.51 W
- covers: run-0606
- HO-014 Cell-1 b2048/ub512 latency arm (cfg-0165, f16/f16 KV, build 2586f6ed Vulkan/RADV). Counter interpolated 33.845891 -> 33.859336 kWh over the window (13.44 Wh). 25 power samples, no contention. Perf-case, no capability claim. Energy join = explicit HA counter-diff, never a silent null.
eng-0257
- 2026-08-23 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-history samples queried by hborchestrator via Nabu Casa API and linearly interpolated to the window; power from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 13.2554 · mean 176.02 W · baseline 10.1 W (box-idle-all-empty) · delta 165.92 W
- covers: run-0607
- HO-014 Cell-1 b1024/ub512 latency arm (cfg-0166 BEST-TTFT / production config, f16/f16 KV, build 2586f6ed Vulkan/RADV). Counter 33.859392 -> 33.872648 kWh (13.26 Wh). 24 power samples, no contention. Perf-case, no capability claim. Energy join = explicit HA counter-diff, never a silent null.
eng-0258
- 2026-08-23 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-history samples queried by hborchestrator via Nabu Casa API and linearly interpolated to the window; power from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 13.3337 · mean 177.49 W · baseline 10.1 W (box-idle-all-empty) · delta 167.39 W
- covers: run-0608
- HO-014 Cell-1 b512/ub512 latency arm (cfg-0167, f16/f16 KV, build 2586f6ed Vulkan/RADV). Counter 33.872714 -> 33.886047 kWh (13.33 Wh). 25 power samples, no contention. Perf-case, no capability claim. Energy join = explicit HA counter-diff, never a silent null.
eng-0259
- 2026-08-23 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-history samples queried by hborchestrator via Nabu Casa API and linearly interpolated to the window; power from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 15.5186 · mean 178.42 W · baseline 10.1 W (box-idle-all-empty) · delta 168.32 W
- covers: run-0609
- HO-014 Cell-1 b512/ub256 latency arm (cfg-0168, f16/f16 KV, build 2586f6ed Vulkan/RADV). Counter 33.886113 -> 33.901632 kWh (15.52 Wh). 30 power samples, no contention. Perf-case, no capability claim. Energy join = explicit HA counter-diff, never a silent null.
eng-0260
- 2026-08-23 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-history samples queried by hborchestrator via Nabu Casa API and linearly interpolated to the window; power from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 19.5859 · mean 166.11 W · baseline 10.1 W (box-idle-all-empty) · delta 156.01 W
- covers: run-0610
- HO-014 Cell-1 b512/ub128 latency arm (cfg-0169, f16/f16 KV, build 2586f6ed Vulkan/RADV). Counter 33.901694 -> 33.921280 kWh (19.59 Wh). 42 power samples, no contention. Perf-case, no capability claim. Energy join = explicit HA counter-diff, never a silent null.
eng-0261
- 2026-08-23 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-history samples queried by hborchestrator via Nabu Casa API and linearly interpolated to the window; power from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 15.6145 · mean 179.14 W · baseline 10.1 W (box-idle-all-empty) · delta 169.04 W
- covers: run-0611
- HO-014 Cell-1 b256/ub256 latency arm (cfg-0170, f16/f16 KV, build 2586f6ed Vulkan/RADV). Counter 33.921340 -> 33.936954 kWh (15.61 Wh). 29 power samples, no contention. Perf-case, no capability claim. Energy join = explicit HA counter-diff, never a silent null.
eng-0262
- 2026-08-23 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-history samples queried by hborchestrator via Nabu Casa API and linearly interpolated to the window; power from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 19.827 · mean 181.66 W · baseline 10.1 W (box-idle-all-empty) · delta 171.56 W
- covers: run-0612
- HO-014 Cell-2 depth rung d131072 (cfg-0166 b1024/ub512, f16/f16 KV, build 2586f6ed Vulkan/RADV). Counter 33.937025 -> 33.956852 kWh (19.83 Wh). 34 power samples, no contention. Depth reachability probe, no capability claim. Energy join = explicit HA counter-diff, never a silent null.
eng-0263
- 2026-08-23 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-history samples queried by hborchestrator via Nabu Casa API and linearly interpolated to the window; power from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 35.1696 · mean 178.57 W · baseline 10.1 W (box-idle-all-empty) · delta 168.47 W
- covers: run-0613
- HO-014 Cell-2 depth rung d200000 (cfg-0166 b1024/ub512, f16/f16 KV, build 2586f6ed Vulkan/RADV) — the >=200k production-depth probe. Counter 33.956908 -> 33.992077 kWh (35.17 Wh). 62 power samples, no contention. Depth reachability probe, no capability claim. Energy join = explicit HA counter-diff, never a silent null.
eng-0264
- 2026-08-23 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-history samples queried by hborchestrator via Nabu Casa API and linearly interpolated to the window; power from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 54.7472 · mean 180.88 W · baseline 10.1 W (box-idle-all-empty) · delta 170.78 W
- covers: run-0614
- HO-014 Cell-2 depth rung d262144 (cfg-0166 b1024/ub512, f16/f16 KV, build 2586f6ed Vulkan/RADV) — no OOM, d262144 viable. Counter 33.992130 -> 34.046877 kWh (54.75 Wh). 98 power samples, no contention. NOTE the run's pp_tps 167.4 carries pp_cv 8.22% (coarse survival figure, not a clean mean); energy Wh-total is unaffected. Depth reachability probe, no capability claim. Energy join = explicit HA counter-diff, never a silent null.
eng-0265
- 2026-08-23 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-history samples queried by hborchestrator via Nabu Casa API and linearly interpolated to the window; power from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 13.638 · mean 180.43 W · baseline 10.1 W (box-idle-all-empty) · delta 170.33 W
- covers: run-0616
- HO-014-RUN Cell-3B f16/f16 KV bench arm at prod config b1024/ub512 d131072 (cfg-0166, build 2586f6ed Vulkan/RADV). Counter 34.870648 -> 34.884286 kWh (13.64 Wh). 25 power samples, no contention. Perf-case (q4_0 comparison baseline), no capability claim. Energy join = explicit HA counter-diff, never a silent null.
eng-0266
- 2026-08-23 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-history samples queried by hborchestrator via Nabu Casa API and linearly interpolated to the window; power from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 13.1539 · mean 182.4 W · baseline 10.1 W (box-idle-all-empty) · delta 172.3 W
- covers: run-0617
- HO-014-RUN Cell-3B q4_0/q4_0 KV bench arm at prod config b1024/ub512 d131072 (cfg-0171, build 2586f6ed Vulkan/RADV). Counter 34.884342 -> 34.897496 kWh (13.15 Wh). 25 power samples, no contention. Perf-case (KV-quant arm), no capability claim. Energy join = explicit HA counter-diff, never a silent null.
eng-0267
- 2026-08-23 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-history samples queried by hborchestrator via the HA API and linearly interpolated to the window; power from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 139.097 · mean 157.4 W · baseline 10.1 W (box-idle-all-empty) · delta 147.3 W
- per task: 5.35 Wh · — p (tariff 30.3 p/kWh)
- covers: run-0618
- HK-RERUN-REASONOFF tau2 full-26 (-rea off, run-0618). Counter interpolated 139.097 Wh over the exact runner window 18:26:25-19:19:26Z (all-windows.jsonl timestamps; well inside HA retention, joined same day). Mean 157.4 W active, peak 198.8 W. Wh/correct-answer basis: 11/26 passed = 5.35 Wh/task-passed (12.64 Wh per task attempted). No tariff attribution.
eng-0268
- 2026-08-23 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-history samples queried by hborchestrator via the HA API and linearly interpolated to the band edges; power from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 17.096 · mean 106.5 W · baseline 10.1 W (box-idle-all-empty) · delta 96.4 W
- covers: run-0619
- HK-RERUN-REASONOFF specsweep spec=ON band (run-0619): 17.096 Wh over 19:25:30-19:35:08Z covering all 10 cells (per-cell windows also joined and range 0.76-1.56 Wh each; shallow d0 cells are short so their per-cell Wh is dominated by fixed overhead). Band-level figure preferred over per-cell sums because cells are back-to-back with server restarts between them.
eng-0269
- 2026-08-23 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw ~10s-history samples queried by hborchestrator via the HA API and linearly interpolated to the band edges; power from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 22.206 · mean 125.5 W · baseline 10.1 W (box-idle-all-empty) · delta 115.4 W
- covers: run-0620
- HK-RERUN-REASONOFF specsweep spec=OFF band (run-0620): 22.206 Wh over 19:35:31-19:46:08Z covering all 10 matched cells (per-cell 0.72-2.20 Wh). Matched-band comparison vs eng-0268: spec=ON used 17.1 Wh vs spec=OFF 22.2 Wh for identical work at identical throughput (~1.0 ratio) — i.e. --kairic-edge drew ~23% LESS wall energy while producing the same tokens. Plausible mechanism: speculative verification shifts work to cheaper batched draft passes even at flat acceptance-driven throughput; treat as one observation, not an established effect. Per-cell windows joined from all-windows.jsonl timestamps; within HA retention.
eng-0270
- 2026-08-23 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw history samples queried by hborchestrator via the HA API and linearly interpolated to the window edges; power from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 43.243 · mean 172.8 W · baseline 10.1 W (box-idle-all-empty) · delta 162.7 W
- per task: 8.649 Wh · — p (tariff 30.3 p/kWh)
- covers: run-0621
- HK-001-REASONING-MATRIX cell R0-B4k (reasoning OFF, max_tokens=4096, wall 901 s): 43.243 Wh, mean active ~172.8 W — roughly HALF the per-cell energy of every reasoning-ON cell (83.995/88.72/91.892/86.561 Wh), matching the ~2x wall-clock cost of reasoning-ON at equal reward-rate-per-token budgets. Window edges from run.log timestamps; 90 samples in-window, no gaps — joined same-day within HA retention.
eng-0271
- 2026-08-23 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw history samples queried by hborchestrator via the HA API and linearly interpolated to the window edges; power from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 83.995 · mean 167.2 W · baseline 10.1 W (box-idle-all-empty) · delta 157.1 W
- per task: 16.799 Wh · — p (tariff 30.3 p/kWh)
- covers: run-0622
- HK-001-REASONING-MATRIX cell R1-B6k (reasoning ON unlimited budget, wall 1808 s): 83.995 Wh, mean active ~167.2 W — roughly DOUBLE R0's energy for +0.2 reward, consistent with reasoning-ON's ~2x wall clock. Window edges from run.log timestamps; 178 samples in-window, no gaps — joined same-day within HA retention.
eng-0272
- 2026-08-23 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw history samples queried by hborchestrator via the HA API and linearly interpolated to the window edges; power from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 88.72 · mean 168.5 W · baseline 10.1 W (box-idle-all-empty) · delta 158.4 W
- per task: 17.744 Wh · — p (tariff 30.3 p/kWh)
- covers: run-0623
- HK-001-REASONING-MATRIX cell R2-B8k (reasoning ON unlimited budget, wall 1896 s): 88.72 Wh, mean active ~168.5 W. This is the reviewer-recommended config for the full-26 confirmation run. Window edges from run.log timestamps; 185 samples in-window, no gaps — joined same-day within HA retention.
eng-0273
- 2026-08-23 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw history samples queried by hborchestrator via the HA API and linearly interpolated to the window edges; power from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 91.892 · mean 170.3 W · baseline 10.1 W (box-idle-all-empty) · delta 160.2 W
- per task: 18.378 Wh · — p (tariff 30.3 p/kWh)
- covers: run-0624
- HK-001-REASONING-MATRIX cell R3-B8k-L (reasoning ON, budget=1024, wall 1942 s): 91.892 Wh, mean active ~170.3 W — highest of the matrix despite matching R1/R2/R4 reward; budget capping did not reduce energy here. Window edges from run.log timestamps; 189 samples, no gaps — joined same-day within HA retention.
eng-0274
- 2026-08-23 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw history samples queried by hborchestrator via the HA API and linearly interpolated to the window edges; power from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 86.561 · mean 167.5 W · baseline 10.1 W (box-idle-all-empty) · delta 157.4 W
- per task: 17.312 Wh · — p (tariff 30.3 p/kWh)
- covers: run-0625
- HK-001-REASONING-MATRIX cell R4-B8k-M (reasoning ON, budget=2048, wall 1847 s): 86.561 Wh, mean active ~167.5 W. Matrix total across all five cells: 394.4 Wh. Window edges from run.log; 184 samples, no gaps — joined same-day within HA retention.
eng-0275
- 2026-08-27 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw history queried via ha_get_history and linearly interpolated to the window edges; power cross-checked from sensor.hardware_ai_hydra_power 5-minute statistics) · window recorded
- wh_total 256.08 · mean 139.9 W · baseline 10.1 W (box-idle-all-empty) · delta 129.8 W
- per task: 9.849 Wh · 0.3 p (tariff 30.3 p/kWh)
- covers: run-0632
- tau2 airline full-26 (run-0632, wall 6605 s). Energy counter 37.975973 -> 38.232057 kWh = 256.08 Wh total; idle floor 10.1 W * 6605 s / 3600 = 18.53 Wh; active 237.55 Wh. Headline Wh-per-correct = active_wh / 24 correct = 9.898 Wh; pence-per-correct = 0.300 p @ 30.3 p/kWh. Whole-run cost 7.76 p. Power sensor mean (~140 W) independently matches the counter-derived 139.5 W. Joined same-day, well inside HA ~10-day raw retention. Context: ho003 stock-f16 tau2 control (eng-batch) recorded 42.06 Wh per correct answer at 10 correct — Flash-Next is ~4x more energy-efficient per correct answer, solving 24 vs 10 in less wall time.
eng-0276
- 2026-08-29 · wall-meter · AI Hydra Home Assistant cumulative kWh counter sensor.hardware_ai_hydra_energy, queried by the authorized hborchestrator reader and linearly interpolated to retained exact UTC edges. · window recorded
- wh_total 419.1478 · mean 146.4982 W · baseline 10.1 W (box-idle-all-empty) · delta 136.3982 W
- per task: 20.5395 Wh · — p (tariff p/kWh)
- covers: run-0639
- POST-HOC RECOVERED JOIN, not campaign-time metering compliance. The immutable campaign package remains unchanged. Exact 10,300-second Tau2 window from retained receipts; 1,151 monotonic counter samples interpolate 39.97000636723048 to 40.389154126317216 kWh, yielding 419.14775908673363 Wh total and 146.49824589439234 W mean. The dated 10.1 W empty-box baseline contributes 28.897222222222222 Wh, leaving 390.2505368645114 active Wh and 20.539501940237443 active Wh per 19 correct. Independent power history has 1,003 in-window samples, a 146.37548016313303 W unweighted sample mean, 146.38710568774053 W linearly time-weighted mean and 192.460464477539 W sampled peak. Counter difference is the admitted method; sampled power is only a consistency check. Raw/derived receipts SHA-256: 7ca023539b92c7d8e4547d15e78f1122297b7358bd2417df2d2cc78980d217d9 / 1c48258aec1b975860df1fd471315a0ecd597084464b8c982d5c600acdfe31cd.
eng-0277
- 2026-09-02 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy, raw history queried via ha_get_history at the window edges and a midpoint; power cross-checked from sensor.hardware_ai_hydra_power hourly statistics) · window recorded
- wh_total 286.3 · mean 153.8 W · baseline 10.1 W (box-idle-all-empty) · delta 143.7 W
- per task: 12.74 Wh · 0.386 p (tariff 30.3 p/kWh)
- covers: run-0653
- tau2 airline full-26 (run-0653, wall 6701 s). Energy counter 44.38310 -> 44.66940 kWh = 286.3 Wh total; steady ~154 W throughout (midpoint 22:04Z = 44.52170 kWh cross-checks: 150.9 W first half, 156.6 W second half). Idle floor 10.1 W * 6701 s / 3600 = 18.80 Wh; active 267.5 Wh. Headline Wh-per-correct = active_wh / 21 correct = 12.74 Wh; pence-per-correct = 0.386 p @ 30.3 p/kWh; whole-run cost 8.68 p. baseline_w 10.1 W is the same box-idle-all-empty floor used for eng-0275 (identical aihydra hardware). Power-sensor hourly statistic for the 22:00-23:00 UTC bucket read anomalously low (~17 W mean) against the counter-derived ~154 W — a sensor-reporting gap, not real idle; the cumulative energy counter is authoritative and its steady progression (verified at the midpoint) is used. NOT reasoning-matched to eng-0275: this run used the model's DEFAULT reasoning effort (heavier generation) in the single-slot adaptive-MTP serving config, whereas eng-0275 (UD-Q4_K_XL) used reasoning_effort=low — so the 12.74 vs 9.90 Wh-per-correct gap conflates quant, reasoning depth, and a lower correct count (21 vs 24); a reasoning=low re-run is the matched comparison.
eng-0278
- 2026-09-05 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy) · window recorded
- wh_total 333.87 · mean 189.6 W · baseline 10.1 W (box-idle-all-empty) · delta 179.5 W
- per task: 15.176 Wh · 0.46 p (tariff 30.3 p/kWh)
- covers: run-0658 run-0659
- The τ² 0-25 worker-config window (cfg-0183, run-0658 / run-0659). Counter differenced on the HA energy sensor: 47.96667902 kWh at 13:41:58Z -> 48.30054615 kWh at 15:27:37Z = 333.87 Wh over 1.761 hours, mean 189.6 W. wh_per_task is Wh-per-CORRECT-ANSWER: wh_total / tasks_passed (22 of 24 scored, run-0658), matching the method eng-0081 used for the earlier Q8_0 run so the two are directly comparable: 15.18 Wh per correct answer, 0.46 pence at the 30.3 p/kWh tariff. This is WHOLE-SESSION energy — the window spans simulator (user_llm) wait time and inter-task gaps, not agent-only inference — so it should not be read against a decode-only bench cell. It is far below the earlier single-slot Q8_0 run's 38.29 Wh per correct (eng-0081), for two reasons that are not a clean A/B: this run served --parallel 4, so the ~190 W draw is amortised across ~4 concurrent tasks, and it got more answers right (22 vs 15). Quant, build, reasoning and parallelism all differ at once (clm-0131). BASELINE CAVEAT: baseline_w is the canonical aihydra box-idle-all-empty floor of 10.1 W (idle-aihydra-2026-08), used so delta_w stays comparable across the whole energy series. A spot idle reading taken this session was higher (~20 W), most likely because the GPU power state had been forced high and the strix-halo fan-control daemon is now resident; a fresh idle-baseline for aihydra's post-fan-control era is an open item and would refine delta_w (it does not touch wh_per_task, which uses wh_total).
eng-0279
- 2026-09-05 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy) · window recorded
- wh_total 130.18 · mean 162.2 W · baseline 10.1 W (box-idle-all-empty) · delta 152.1 W
- covers: run-0661 run-0662 run-0663 run-0664
- Single-stream depth-sweep window (cfg-0184, run-0661..run-0664, four depths). Counter differenced on the HA energy sensor: 47.81016026 kWh at 12:39:09Z -> 47.94033955 kWh at 13:27:18Z = 130.18 Wh over 0.80 hours, mean 162.2 W. This is a throughput sweep with no task/answer concept, so wh_per_task is null. Baseline is the canonical 10.1 W floor; see eng-0278's baseline caveat.
eng-0280
- 2026-09-06 · wall-meter · aihydra energy sensor via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy) · window recorded
- wh_total 205.16 · mean 186.56 W · baseline 10.1 W (box-idle-all-empty) · delta 176.46 W
- per task: 10.258 Wh · 0.31 p (tariff 30.3 p/kWh)
- covers: run-0665
- The pwilkin candidate's τ² 0-25 window (cfg-0185, run-0665). Counter differenced on the HA energy sensor: 48.72781343 kWh at 09:27:14Z -> 48.93297700 kWh at 10:33:13Z = 205.16 Wh over 1.10 hours, mean 186.6 W. wh_per_task is Wh-per-CORRECT-ANSWER: wh_total / tasks_passed (20 of 24 scored, run-0665), the same method as eng-0081 / eng-0278 so the three compare directly: 10.26 Wh per correct answer, 0.31 pence at 30.3 p/kWh. Well below the worker's 15.18 (eng-0278) for two reasons, neither a clean A/B: the pwilkin stack decodes faster (66-min window vs the worker's 106 min -> less total energy) and got fewer answers right (20 vs 22). Whole-session energy (includes user-sim wait). Baseline is the canonical 10.1 W aihydra floor (see eng-0278's baseline caveat).
Contributions (19)
con-0001 — ggml-org/llama.cpp (carrying)
- fork · opened 2026-07-21 · link
- carrying reproducibility hazard — in use but not upstream
- Disk slot save/restore silently loses all prompt reuse on hybrid/recurrent models because context checkpoints are never persisted. Fixed with a .ckpt sidecar and pushed to headbouyJB/llama.cpp@fix-25913. Independently confirmed working by two community testers. A third-party PR (#26004) fixes the same bug by appending checkpoints inside the save file instead; production carries our fork until one of the two lands upstream.
con-0002 — ggml-org/llama.cpp (carrying)
- fork · opened 2026-08-18 · link
- carrying reproducibility hazard — in use but not upstream
- Nathanw1014's Strix Halo Vulkan fork, staged first as the v0.6.4 portable payload at build baf6360be and then as the reviewed v0.6.6 portable payload at build 7b6c6133/source 7b6c61330edf370659f531932e0b91aca67ba055. The payloads carry RADV wave32/cooperative-matrix tuning, dense and sparse prefill work, DeepSeek-V4 Lightning Indexer support and upstream merges. Results produced with them are fork-specific and must not be attributed to stock llama.cpp or to any one bundled patch.
con-0003 — ggml-org/llama.cpp (open)
- pr · opened 2026-08-18 · link
- Open llama.cpp DFlash2 support PR for Qwen3.8-27B-style drafters. Community comments currently report hardware-dependent results: modest single-slot Strix Halo Vulkan gains over MTP on code but near-parity on prose, a severe Intel B70 multi-agent collapse, a V100 multimodal/M-RoPE draft-context failure with a proposed fix, RTX 3090 near-parity to modest gains, and Blackwell scaling better at higher parallelism. This is upstream/community context only; no HaloBench measured-here result depends on it yet.
con-0004 — ggml-org/llama.cpp (open)
- report · opened 2026-08-20 · link
- Independent Strix Halo Vulkan follow-up on llama.cpp #25618. The report reproduced the earlier F16-V byte-exact PASS on the original prose prompt using matching Qwen3.8-27B Q6_K target and Q8_0 MTP draft artifacts, llama.cpp 9d57ce456c94d241dde672b2db9cf18879766568, f16 K/V cache and the reporter's greedy request contract. Under the same runtime, server flags, template path and completion path, five additional prompts had 0/5 exact-token parity, and a separate fixed-sampler replay of those same prompts also had 0/5 exact parity with unchanged first mismatch locations. Treat this as an upstream diagnostic report, not a HaloBench measured-here run: it establishes prompt sensitivity of the exactness check for that target/draft/runtime combination and argues that a single exact MTP trajectory is not evidence of general baseline/MTP invariance.
con-0005 — Akicou/diffuse-cpp (carrying)
- fork · opened 2026-08-22 · link
- carrying reproducibility hazard — in use but not upstream
- Makes inclusionAI LLaDA2.2-flash (100B-A13B block-diffusion MoE) run GPU-accelerated on AMD Strix Halo (gfx1151) and benchmarkable as an agent. Over the Akicou base (b799157): GPU offload of the MoE forward to the HIP/ROCm backend; OpenAI tool-calling in diffuse-server; a manual soft_max attention path that works around a gfx1151 flash_attn_ext mask leak [clm-0104]; a resurrected + GPU-resident inter-step KV cache with in-place flash decode and cross-turn prompt reuse (the key speedup, ~1 -> ~5-12 tok/s at long context) [clm-0103]; and a faithful port of the model's Levenshtein editing (M2T + T2T + DELETE/SPLIT + anti-loop + post-steps) restoring its self-correction. Fork to be published at headbouyJB/diffuse-cpp; patch currently held on aihydra. Base runtime AND the diffuse.* GGUF are a matched pair, so this carries the Akicou base rather than rebasing.
con-0006 — ggml-org/llama.cpp (open)
- report · opened 2026-08-22 · link
- New llama.cpp #25618 follow-up by snick525 (2026-08-22T02:11:18Z) on RDNA4 (AMD Radeon AI PRO R9700 / gfx1201, 32 GB, Vulkan/radv), target Qwen3.8-27B Q6_K_XL (MTP head intact), greedy (temperature 0, top_k 1, top_p 1, seed 42), f16 K/V. The reporter additionally built open PR #27342 (DFlash2, an external 2B draft model) so the built-in MTP head is no longer the only speculation type in the comparison. Three findings against a --spec-type none baseline on the same binary: 1. DIVERGENCE IS DRAFTER-INDEPENDENT. At -c 65536 / f16 KV the code prompt diverges at the same first-difference byte 204 for the built-in MTP head AND for DFlash2 at three drafter quants (Q4_K_M, Q8_0, BF16); all four configs are lossless on a reasoning prompt. Same first-diff byte holds across two context sizes (16384 and 65536) and two builds. Whatever drifts is on the target/verify side; the drafter is not a variable in it. 2. DRAFTER QUANTIZATION HAS NO EFFECT. DFlash2 Q4_K_M vs Q8_0 vs BF16 produce byte-identical outputs on both prompts and both context sizes, and identical acceptance counters (draft 347, accepted 283). Consistent with the verify path emitting only the target's own token: a drafter can only change output by proposing different tokens, and here even a Q4 drafter proposes identically, so drafter precision can be ruled out when bisecting. 3. KV-CACHE QUANTIZATION IS ITSELF OUTPUT-MOVING AND COMPOUNDS. Swapping f16 for q8_0 KV (everything else fixed) turns the reasoning prompt from LOSSLESS under f16 to DIVERGES at byte 155 under q8_0 for both MTP and DFlash2, and moves the code prompt's first-diff from byte 204 to byte 277. Separately, KV quantization alone moves greedy output with no speculation on either side: --spec-type none with f16 KV versus q8_0 KV differs on the code prompt at byte 561. The reporter notes q8_0 KV is accordingly never output-preserving and compounds with whatever this issue turns out to be. Throughput context (mentioned only because it came up in #27342, and reported with an active caveat): -c 65536 / f16 KV on Qwen3.8-27B Q6_K_XL on a 32 GB card gave 23.8 tok/s no-spec, 40.4 tok/s draft-mtp n_max 1, 61.9 tok/s DFlash2 Q4_K_M, and 29.4 tok/s for DFlash2 Q8_0 (the Q8 and BF16 drafter rows are NOT drafter-cost evidence: at 31.2-31.6 GB they cross the card's usable ceiling and silently fall back to host memory per #26432, with acceptance unchanged). Below the ceiling every drafter quant performs identically, so the Q8 spill is a memory-footprint artefact, not a drafter-precision finding. Upstream diagnostic record only. Surface is RDNA4 R9700/gfx1201 Vulkan, not our ROCm/Vulkan Strix Halo gfx1151 aihydra box, so no HaloBench benchmark claim is made from these figures.
con-0007 — Nathanw1014/strix-halo-llamacpp (open)
- report · opened 2026-08-22 · link
- Nathanw1014/strix-halo-llamacpp (the Strix Halo Vulkan fork tracked in con-0002) shipped v0.6.8 and v0.6.9, two stable releases that carry a deliberate MTP rollback EXACTNESS tradeoff on hybrid GDN targets. v0.6.8 (2026-08-22T02:40:00Z) added DFlash2 speculative decoding support for mmproj/vision (community-reported vision-prompt HTTP 500 fixed by injecting draft rows at dense per-token positions and triming the draft cache in token space) and introduced a one-line change, "common: use full checkpoints for MTP rollback", that forced MTP rollback through full sequence-state checkpoints. v0.6.9 (2026-08-22T04:09:43Z) REVERTS that one line (revert commit a17e843, payload Nathanw1014/llama.cpp@a17e8432b, branch strix-halo-vulkan): on Vulkan with a hybrid GDN target (qwen35moe, e.g. Qwen3.6-35B-A3B) the full-checkpoint save/restore path deadlocked a few hundred tokens into a long response - every server thread parked in futex_do_wait with the GPU idle and GTT flat, reproduced deterministically and bisected on hardware. v0.6.7, CPU, and DFlash2 (incl. the vision path) were unaffected. The vendor states the tradeoff plainly: with the revert, MTP rollback returns to v0.6.4-v0.6.7 behaviour - fast snapshot-plane restore that CAN DIVERGE SLIGHTLY from a no-draft run after a rejected draft (called "a subtle distribution drift after rejected drafts, not garbled output"). Re-land of the full-checkpoint restore is deferred until the state-save path is fixed on Vulkan. Speed figures in the notes are single runs, explicitly NOT the BENCHMARKS.md protocol. Treat as an upstream/report record only: single-payload fork-vendor verification on gfx1151 Vulkan, not a HaloBench benchmark result, and not inherited into any HaloBench number. SUPERSEDED as an operative statement by fork v0.6.10 (see con-0009/clm-0107, 2026-08-22T12:03:16Z), which restored token-exact MTP rollback (root cause: the server re-verifying replayed draft tokens after a checkpoint restore) — this record stands as the historical v0.6.8/v0.6.9 report.
con-0008 — ggml-org/llama.cpp (open)
- report · opened 2026-08-22 · link
- llama.cpp #25618 comment #16 by F-Mangini (2026-08-22T20:32:01Z) corroborates the quantized-target divergence on a NEW model/OS/GPU family: the official LiquidAI LFM2.5 DSpark pair also diverges from vanilla on a quantized target while an F16 target preserves greedy parity. Environment: llama.cpp b10566 (bb4caa754), Vulkan build, Windows 11, NVIDIA GTX 1660 SUPER 6 GB (driver 551.76). Target: LiquidAI/LFM2.5-2.6B-GGUF Q8_0 (SHA-256 36587fdf27bdfc69caf2637273679a0870ec155162161bde6fd16e8c70bdb757); draft: LiquidAI/LFM2.5-2.6B-DSpark-GGUF Q8_0 (SHA-256 85a98fafd9af1328b6876fd1360d7ed69e74c6cefc14dd07fb6306e1940386c87). Minimal raw /completion reproduction without a chat template, server `-c 4096 -np 1 -fa on -ngl 99 -ctk f16 -ctv f16`, requests temperature=0 seed=42 n_predict=256, on the LRU-cache Python prompt. Quantized-target finding: vanilla Q8_0 target outputs SHA-256 2eeb8174..., same target + DSpark --spec-draft-n-max 1 outputs a8d38c69... (draft accepted 107/147) -- DIVERGES. Differential controls: (1) Q8_0 target + DSpark loaded but --spec-draft-p-min 1 (zero draft tokens) exactly matches vanilla Q8_0; (2) the active Q8_0 DSpark run is deterministic but consistently differs from vanilla; (3) the SAME Q8_0 draft against an F16 target produces byte-identical greedy output to the F16 vanilla target while accepting 40/64 proposed tokens over 64 generated tokens -- i.e. F16 target preserves greedy parity. The Q8_0 mismatch also persists with --spec-draft-n-max 1 and target KV in F16, so it is not caused by a larger speculative block or by quantized target KV. Notably the n_max=1 divergence differs from the Qwen3 boundary reported earlier in the same issue, where n_max=1 was said to remain lossless. Reporter concludes this is consistent with path-dependent target numerics between sequential single-token decoding and the speculative verification path for quantized weights, and confirms the issue affects LFM2 DSpark, NVIDIA Vulkan, Windows and Q8_0 -- not only the Q4/MTP configs in the original report. Upstream diagnostic record only: surface is NVIDIA Vulkan/Windows, not our ROCm/Vulkan Strix Halo gfx1151 aihydra box, so no HaloBench benchmark claim is made from these figures.
con-0009 — Nathanw1014/strix-halo-llamacpp (open)
- report · opened 2026-08-22 · link
- Nathanw1014/strix-halo-llamacpp (the Strix Halo Vulkan fork tracked in con-0002) shipped STABLE v0.6.10 (2026-08-22T12:03:16Z) that REVERSES the v0.6.8/v0.6.9 MTP rollback-EXACTNESS tradeoff recorded in con-0007/clm-0105. Token-exact MTP rollback returns on hybrid GDN (qwen35moe) targets (e.g. Qwen3.6-35B-A3B). Root cause of the v0.6.9 deadlock was NOT the state-save path: the server was RE-VERIFYING replayed draft tokens after a checkpoint restore, which livelocked the slot. Fix lands in two commits — "server: do not re-verify replayed draft tokens after a checkpoint restore" (9c5d899) and re-apply "common: use full checkpoints for MTP rollback" (f25eefe). With the re-verify eliminated, MTP rollback through full sequence-state checkpoints is token-exact again and the long-run stall is gone (smoke: Qwen3.6-35B-A3B MTP Q6_K 800-token repro 13.9 s no stall; 3000-token run 46.4 s / ~65 tok/s past the former stall horizon). v0.6.10 ALSO adds DSpark speculative-decode support for bailingmoe3 (Ling 3.0), cherry-picked from upstream llama.cpp PR #27508 (merged 2026-08-22T09:19:49Z, btw616). Runtime side is complete; the reworked bailingmoe3 forward pass was A/B verified byte-identical to the prior tip on Ling-3.0-tiny (greedy, 96 tokens), but the DSpark path is NOT exercisable end-to-end until Ling-3.0-flash draft GGUFs are published (inclusionAI/Ling-3.0-flash-dspark has model.safetensors up, GGUFs not out as of 2026-08-22T16:19Z). Also gates the RADV coopmat LDS pad of 2 to RADV >= 25.3 (older drivers violated VUID-08986). Payload built from Nathanw1014/llama.cpp@2586f6ed (branch strix-halo-vulkan). Speed figures are single runs, explicitly NOT the BENCHMARKS.md protocol; a pre-existing CPU-only divergence in one 800-token exactness matrix cell (token 776) predates v0.6.8 and is unrelated to the rollback path. Treat as an upstream/fork release-note record only: single-payload fork-vendor verification on gfx1151 Vulkan, not a HaloBench benchmark result, and not inherited into any HaloBench number.
con-0010 — ggml-org/llama.cpp (open)
- report · opened 2026-08-22 · link
- llama.cpp #27122 comment by mazinist (2026-08-22T23:51:13Z) independently confirms the MTP/CUDA multi-GPU lockup originally reported by tripletto, and validates a workaround on a completely different platform from the earlier zyxyunxin note (#issuecomment-5310571689, 2026-08-17). Hardware: Threadripper Pro 3945WX (WRX80, Gigabyte MC62-G40), 4x RTX A4000 16GB (PCIe 4.0 x16, NO NVLink / no P2P — nvidia-smi topo shows NODE), Ubuntu 24.04, kernel 6.14, closed driver 595.84, llama.cpp build 749f688fc (Aug 21, CUDA 13.2). Repro: Qwen3.8-27B-UD-Q6_K, --split-mode tensor -ts 25,25,25,25, --spec-type draft-mtp --spec-draft-n-max 3, 131072-token deep-context prefill via llama-benchy. 6/6 hard crashes, dying 2-4 min into the prefill; symptom on this platform is harsher than a lockup — the entire machine hard-powers-off (BMC logs "Power off/down", no Xid/AER/panic). KEY DISCRIMINATOR: the same workload on 2 GPUs (which uses the internal 2-device AllReduce path, not the meta-backend) is completely stable with MTP enabled, so the defect is specific to the multi-GPU tensor-split MTP path. WORKAROUND CONFIRMED: LLAMA_GRAPH_REUSE_DISABLE=1 lets the exact crash config survive the full 131072 prefill (610.82 t/s prefill, 27.53 t/s decode @131K, 38.43 t/s @4K), consistent with PR #24549's mechanism (graph reuse leaves dangling per-device tensor references when MTP/target contexts share memory under SPLIT_MODE_TENSOR). Also matches the reporter's notes: no-MTP tensor is stable, layer split never affected, and lockup frequency scales with --spec-draft-n-max. Upstream diagnostic record only: surface is CUDA multi-GPU, not our single-GPU ROCm/Vulkan Strix Halo gfx1151 aihydra box, so no HaloBench benchmark claim is made from these figures.
con-0011 — ggml-org/llama.cpp (open)
- pr · opened 2026-08-17 · link
- llama.cpp PR #27210 by stew675 "spec : add adaptive MTP draft depth (draft-mtp-adaptive)" — adds a new --spec-type draft-mtp-adaptive with a counting state machine (climb counter + weighted drop-pressure accumulator) so draft depth adjusts per segment instead of staying at a fixed n_max. Suggested config --spec-draft-n-max 12; floor/cold-start default 3. Author's own table (Qwen3.8-27B Q8_0, 2x Radeon AI PRO R9700/gfx1201 ROCm, temp 0.6, ctx 8192): on coding, adaptive (n_max 12, floor 2) reaches 86.4 t/s vs fixed depth 3 at 78.8 t/s, driven by long mean draft length (8.0) with 58.4% accept; on prose and hard prose the adaptive rows sit roughly AT or slightly BELOW fixed depth 3 (54.4 vs 56.4 and 51.0 vs 52.6), which is the small reasoning/prose penalty trace. KEY FINDING (comment 5382662863, 2026-08-22T21:21:33Z): stew675 modified the algorithm to stay at FIXED depth 3 but kept max depth 10, and found that merely HAVING max depth 10 incurred a fixed ~2.6% performance penalty completely independent of the adaptive logic — i.e. the penalty is the cost of a high fixed n_max CEILING, not the adaptive switching itself. NEW FIX pushed 2026-08-23T01:32:47Z (comment 5383599943): limits the full MTP buffer scan on truncated/short drafts, giving a ~2-3% speed boost for --spec-draft-n-max 10 --spec-draft-p-min >0.5 configs. INDEPENDENT CORROBORATION (comments 5375448677 at 2026-08-21T21:09:53Z and 5377578075 at 2026-08-22T03:18:20Z, marcusds on a single RTX 5090 32GB, CUDA 13.3, same e1a5754 commit and script): GGML_CUDA_DISABLE_GRAPHS=1 vs default deltas are small (baseline C0 -3.0% / -1.9% / -1.8% / -1.7%), i.e. depth CHANGING does not slow CUDA graph perf; the adaptive 3..10 + p-min config (C3) was the only row with a POSITIVE recall delta (+2.6%) — corroborating that a fixed high ceiling, not adaptive switching, is the cost. Upstream PR record only: surface is CUDA/gfx1201, not our single-GPU ROCm/Vulkan Strix Halo gfx1151 aihydra box, so no HaloBench benchmark claim is made from these figures.
con-0012 — ciru-ai/ROCmFPX (carrying)
- fork · opened 2026-08-23 · link
- carrying reproducibility hazard — in use but not upstream
- The Kairic Edge Qwen3.8-27B-IU4 build (branch kairic-edge-qwen38-27b-v1.1, commit e1da26bb8) from ciru-ai/ROCmFPX — a vendor-authored fork that binds the ROCmFPX binary (called "TheRock" stack here because the carrier vendor certifies it against TheRock 7.14 / AMD clang 23.0.0) to a custom 4-bit IU4 sidecar layout: a Q4-ish quant GGUF (Qwen3.8-27B-IU4-Kairic-Edge.gguf) + three PROMPTFORGE_* .pfs sidecars (FFN, GDN, GDN-Output) that accelerate the FFN and GDN projector paths on gfx1151, plus a --kairc-edge server flag and a compat mode (KAIRIC_EDGE_COMPATIBILITY_MODE) that enables tool-calling where fast-greedy forbids grammar/tool calls. HK-001 on aihydra rebuilt this at commit e1da26bb8 (binary e121a4d3, model sha 360caf73, sidecar hashes adcbb90a/82f93129/3b07e7b1, gfx1151 / AMD clang 23.0.0). It is a carried fork — results produced with it are fork-specific and must not be attributed to stock llama.cpp, and not to any one bundled patch. Backwards the fixed GDN output path is NOT accelerated (gdn_output_fallback 48) on this aihydra build and the toolchain differs from the vendor-certified one (ROCm 7.1 vs 7.14, LLVM 21 vs AMD clang 23), so the vendor's throughput claims do not transfer; every number must be treated as measured-here on this exact build. [2026-08-24 v1.2 RUNTIME MANDATE — HK-V12-ANNOTATE, per t_b05473c9] The vendor released runtime tag kairic-edge-qwen38-27b-v1.2 @ 205a3e5f (HF repo updated 2026-08-23T14:28:32Z): the v1.1 native IU4 M65 verifier could flip a greedy target token on a low-margin reproduced case; v1.2 defaults to the strict-compact authoritative verify path (unsafe native only via KAIRIC_UNSAFE_NATIVE_M65_VERIFY=1). GGUF + all three .pfs sidecars are byte-unchanged (gguf sha 360caf73…), so throughput records stand. Vendor validation: 6/6 exact-match to frozen no-spec target, draft acceptance 99.73%. ALL FUTURE Kairic Edge box runs MUST build from the v1.2 tag; every HK run recorded to date used v1.1 and carries a comparability annotation on its claim (clm-0116, clm-0118, clm-0119).
con-0013 — LaurentZuijdwijk/llama.cpp (carrying)
- fork · opened 2026-08-24
- carrying reproducibility hazard — in use but not upstream
- Vulkan fork used by the Qwen3.8-27B ROCmFPX/DFlash2 autonomous tuning handoff. The handoff records commit 16f0799a6 and Mesa RADV 26.0.3. Its production-performance evidence remains screening-only until the required floor, guard, and artifact-identity records are captured.
con-0014 — Nathanw1014/strix-halo-llamacpp (open)
- report · opened 2026-08-25 · link
- Primary-source runtime-watch record: Nathanw1014/strix-halo-llamacpp v0.6.11, published 2026-08-24, fixes the fork-only image-request regression introduced by the fork's v0.6.8 DFlash M-RoPE change. For a position-sharing draft-mtp context, the old token-space trim left the draft behind the target after an image span and led to inconsistent M-RoPE positions, llama_decode failure, and HTTP 500. Commit 0eb528051a56f34567312ce63ab4e14a3fc71d89 records whether a draft cache is dense-row indexed: DFlash/DSpark retain token-space trimming, while draft-mtp and eagle3 use target position space. The release's gfx1151 Vulkan/RADV smoke on Qwen3.8-27B-UD-Q6_K_XL plus mmproj-F16 reports that draft-mtp n_max=3 image requests now return HTTP 200 with speculation remaining active; its DFlash image and speculation-off controls are unchanged. The release asset digest is sha256:a4306edefdefb2eff925cbc43cbf32cd42cf07c00e06011e906957217bf1ddad and MANIFEST.txt identifies the portable payload source as 0eb528051. This is a fork-vendor functional-correctness report, not a HaloBench benchmark: the release explicitly says its smoke checks are not BENCHMARKS.md throughput and makes no throughput claim. It has no text-only agent outcome, guard result, capability score, energy join, or served-path/floor pair. The regression never applied to upstream llama.cpp or to fork configurations with speculation off, so it cannot rebaseline existing text-only Qwen3.8 throughput/capability records or replace their recorded v0.6.10/stock runtime provenance. A future vision/MTP prep card is justified only as a separately scoped, on-box validation of this exact payload with multimodal request controls and the house guard; it is not a reason to dispatch a benchmark or change the production pick now.
con-0015 — ggml-org/llama.cpp (open)
- issue · opened 2026-08-24 · link
- Raw upstream diagnostic comment by frizikk on open issue #25618, retrieved directly from GitHub on 2026-08-25. Surface: pinned llama.cpp sources, Vulkan, Qwen3.8-27B Q6_K target plus Q8_0 native MTP, flash attention on, f16 K/V, greedy sampling, and --spec-draft-n-max 1. Its first confirmed target-verify boundary is the natural generation-1 layer-0 Q8_0 x F32 MUL_MAT: sequential N=1 uses F32 MUL_MAT_VEC, while speculative N=2 stages activations as Q8_1 for MMVQ; the direct predecessor is exact but every one of the 5120 output values changes. Replaying the N=2 fixture through the non-MMVQ F32 path made row 0 bit-exact, but later full generation logits still differ, so the narrow diagnostic gate is not a complete fix. The same comment identifies a second N-dependent ADD/RMS-reduction boundary and explicitly confines its flash-attention observation to one captured fixture. This is upstream operator-level evidence, not a HaloBench benchmark or a production fix. It corroborates the standing varied-prompt/raw-token MTP-invariance requirement, but does not identify any local Qwen3.8 result as affected: the closest recorded local Vulkan arms use Q8_0 targets and n_max=3 (cfg-0068/run-0269 and cfg-0147/run-0573), already retain a negative varied-prompt invariance result and are not promoted as baseline/MTP equivalent. It has no direct transfer to the local ROCm Qwen3.8 records.
con-0016 — ggml-org/llama.cpp (open)
- pr · opened 2026-08-24 · link
- Raw upstream comment by alexpooley on open PR #25863, retrieved directly from GitHub on 2026-08-25. The report is HIP/ROCm on the integrated gfx1151 surface, using Qwen3-4B-Instruct-2507 Q8_0, llama-perplexity, c4096, b4096, 12 chunks, and variable ubatch. Its master-versus-proposed-fix PPL table is stable at ub4096 (9.1542 versus 9.1542), but degrades on master as prompt chunking increases: ub2048 11.8957 versus 9.1527, ub1024 2259.7396 versus 9.1557, and ub512 49142.0371 versus 9.1594. The proposed explanation is that the scheduler does not protect caller-writable input during chunking; the author disclosed AI assistance for identifying and implementing the patch, while stating that the final change was personally tested and reviewed. This is a source-bound upstream report, not a HaloBench result or a validated fix. It shares HIP/gfx1151 and low-ubatch shape with the local Qwen3.5-122B production-optimisation matrix (cfg-0151, b2048/ub512), but differs in model, weight quant, tool, corpus, batch, context and runtime. The local four-item guard (run-0588) passed but is not a chunked-input forward correctness test; it cannot prove this source does or does not apply. Therefore no recorded Qwen3.5 number is reinterpreted or invalidated, but future or modified HIP/gfx1151 low-ubatch performance work needs the new exact-fingerprint chunked-input control before a claim is admitted.
con-0017 — kingjones777/Qwen3.8-Flash-Next-ROCmFP4-STRIX-GGUF (carrying)
- report · opened 2026-08-27 · link
- carrying reproducibility hazard — in use but not upstream
- Immutable publisher model-page revision 069dddb53bab04218d734fa9a771f8a0242ab059 describes the full-STRIX artifact and the required kingjones30/ROCmFPX runtime. Its long-context table was measured on gfx1151 with ROCm 7.2.4, different target prompt counts and context allocations, and no HaloBench 512-checkpoint/ 2048-token-cadence proof. The page reports full-STRIX rows through c131072, while its c262144 rows are explicitly STRIX_LEAN. HaloBench used the full-STRIX artifact on HIP 7.1.52801-9999, exact cache-false prompt_n targets, c262144, and the corrected checkpoint flags. The 32k-131k rates are close in magnitude, but neither source transfers: runtime version, prompt bytes/counts, allocation, repetition method and checkpoint boundary differ. The publisher page is a comparison source and attribution caveat, not evidence for a HaloBench number.
con-0018 — ggml-org/llama.cpp (open)
- pr · opened 2026-08-26 · link
- PR #27311 "Scheduler UMA ring buffer (+ sanitizer and fixes)" adds an input ring buffer on UMA devices so the host cannot clobber in-flight graph inputs. OPEN / unmerged as of the 2026-08-26 head commit. Carried on aihydra as branch pr27311 (commit c530ea79c) and is the load-bearing dependency of the Qwen3.8-27B worker serving config (cfg-0183): 0 EMPTY across a 106-minute 4-slot run vs the master build's ~4-minute failure (run-0660), and it is why the worker can serve --parallel 4 at all where the earlier Q8_0 config had to pin --parallel 1. A reproducibility hazard until merged: any rebuild of this worker must fetch the PR branch, not stock master. See cfg-0183, clm-0130.
con-0019 — ROCm/rocm-systems (carrying)
- fork · opened 2026-08-26 · link
- carrying reproducibility hazard — in use but not upstream
- "Retained PM4 dispatch" — a custom AMD HIP/ROCm runtime (rocm-systems ilintar-experiments @ 78d1160, the CLR/hipamd graph path) that keeps a graph's low-level PM4 command buffer resident and replays it, cutting per-token launch overhead. Activated by ENABLE_RETAINED_PM4=1 -> DEBUG_HIP_GRAPH_PM4=1 with HIP graphs on. This is a RUNTIME change (libamdhip64), not a llama.cpp change. Carried on aihydra as the custom rocr+clr runtime behind the pwilkin 27B build (cfg-0185/cfg-0186). It is the ONE piece of the pwilkin stack with NO upstream PR: the author states "all of this is in upstream PRs except for the PM4 stuff which I don't expect to be accepted (but I might retire for proper HRX support when it's there)." A novel-fork reproducibility hazard: any rebuild must build this custom runtime from source; there is no merge path. Measured effect on our worker model is small (~+4%, clm-0135) — it is not the source of the pwilkin speed advantage. The pwilkin llama.cpp changes, by contrast, ARE upstream PRs (e.g. #27311, con-0018).
Claims (135)
clm-0001
- measured-here low ●○○ volatility medium · verified 2026-08-03
- Static pool utilisation did not predict the 2026-07-21 OOM. The box crashed at 83.8% estimated utilisation, and cfg-0002 sits at that same 83.8% today.
- total_gb is weights + KV + a fixed overhead — a static estimate. What actually exhausted memory was runtime growth: context checkpoints defaulting to 32, cache-ram, and fragmentation. None of it appears in the record. The mitigations worked precisely because they targeted dynamic growth (checkpoints capped at 4, desktop stack removed, swap raised to 16G) — which is why they moved the ratio not at all. Until observed_peak_gb is collected from telemetry, the 0.8 warn line is a threshold on a number that has never predicted an OOM. Falsifiable: collect peak RSS/GTT under load and compare against the estimate.
- evidence: inc-0004
clm-0002
- community med ●●○ volatility medium · verified 2026-08-03
- Multi-token prediction gains MORE at heavier quantisation, not less. On this silicon Q8_0 saw 2.44x (7.7 -> 18.1 t/s) against Q4_K_M's 1.81x (12.1 -> 21.2).
- Counterintuitive, and it matters for quant selection: baseline decode here is bandwidth-saturated, so MTP's only lever — fewer memory passes per token — pays more the heavier the weights. Partially offsets the cost of a higher quant. Not yet reproduced on our own hardware; the quant x MTP matrix is queued for the lab box. Our own measured figure is the 122B at UD-Q4_K_M: 20.96 t/s single-slot baseline to 31.17 with MTP n=3 (+49%), and 36.5 t/s after tuning to n=6.
clm-0003
- measured-here med ●●○ volatility low · verified 2026-07-20
- MTP's speedup tracks how predictable the text is: +81% on code, +21% on freeform, with the real agent workload mix landing around +29-40%.
- The headline "+50%" is the favourable end of a range, not a constant. Draft acceptance measured at 75%, averaging ~5.0 tokens per speculative call. Prefill paid a ~10% tax, discounted in practice by a ~90% cache-hit rate — so the tax lands only on cold turns. Single-slot cost zero decode, which is what made the architecture viable.
- evidence: run-0002
clm-0004
- measured-here med ●●○ volatility high · verified 2026-06-28
- gpt-oss-120b placed last on our blind conceptual set — 3/8, mean 1.38 — against Qwen3.6-35B-A3B's 7/8, mean 2.12.
- FINGERPRINT, without which this is a slight rather than a finding. Config cfg-0003: Q4_K_M, ROCm, -c 131072, q8_0 KV, reasoning OFF, `-ngl 999 --no-mmap --parallel 1 -fa 1 --jinja`, on a ~82 GiB budget. Tasks C1-C8, one response per model per task. Judging: a separate, fresh Claude Opus sub-agent per response, given only the task prompt, the rubric and one anonymised answer — no model name, no comparison set, no project context. Every rationale preserved and re-checkable. SCOPE: eight tasks, N=1 per cell, one quant, one date, one llama.cpp build, and a rubric aimed at proactive/generative assistant work. It is not a general capability verdict. The same model was the FASTEST reactive performer in the field (~5 s). Volatility high: this was true of that build on that date and has not been re-tested.
- evidence: run-0003
clm-0005
- measured-here high ●●● volatility low · verified 2026-06-28
- Nemotron 3 Super's benchmark result is NOT quant-equivalent to its peers and must not be read as a like-for-like verdict on the model.
- Every model was tested at the best quant that fits 128K in ~82 GiB. For the field that was a Q4_K-class quant; for Nemotron it was IQ4_XS, because its Q4_K needs 77 GiB at 4K context and OOMs by 16K. It was additionally the only entry with no speculative decoding available — llama.cpp does not support the Mamba-hybrid MTP path — while peers could use it. So it was handicapped twice, and came last on speed (~28 s). The comparison is honest about what THIS BOX can run; it is not honest about the model. This caveat travels with the result rather than sitting in a footnote, which is the whole reason it is a record and not a sentence.
- evidence: run-0003
clm-0006
- community med ●●○ volatility high · verified 2026-08-03
- Quantized KV cache is mis-implemented for Strix Halo in stock llama.cpp: the code dequantizes to full precision repeatedly during inference, which a discrete GPU hides in cache and this box cannot. Community fixes report ROCm text generation +75% to +203% depending on depth, and q8_0 KV generating 23-53% FASTER than f16.
- Source: independent 128GB verification (Nathanw1014) of a fork by another user, four controlled builds — patched and stock, Vulkan and ROCm — flash attention pinned, build hash and flags carried per row. Vulkan prefill +45%/+71%/+87% at 32k/64k/128k on Qwen3-Coder-30B. Full 262,144-token native context runs on both backends; at 262k the compressed cache generates 65% faster for a 2.6% prefill cost. One reported regression: MoE loses 0.3-2.4% ROCm prompt speed with the fix. DIRECTLY LOAD-BEARING FOR US: cfg-0002 runs q8_0/q8_0 KV on ROCm at 200K context, which is the deep end where the reported gain is largest — potentially a bigger win than MTP's +50%. NOT REPRODUCED HERE, and not upstream: no PR exists in ggml-org/llama.cpp, so adopting it means carrying a SECOND fork alongside con-0001. The author states plainly that none of it measures output quality.
clm-0007
- inferred low ●○○ volatility medium · verified 2026-08-03
- A GPU cache wall around 32-40 MB, past which read speed drops roughly 4x, is a candidate explanation for our own prefill degradation from ~354 tok/s at small context to ~144 tok/s at 128K.
- The wall was measured on another 128GB Strix Halo box with a purpose-built benchmark; the originating theory about where the collapse should begin did NOT match the measured curves, so the mechanism is real but not yet modelled. Our ~2.5x prefill drop is the right order for a ~4x bandwidth cliff partially amortised. TWO REASONS TO BE CAUTIOUS: this is a GPU-cache effect, entirely distinct from the pool-utilisation metric in clm-0001 — conflating them would be a mistake; and the reporter's dense model collapsed while his MoE shrugged the same conditions off, which is unexplained. We run MoE, so the q4_0 mitigation that helped his dense model may not transfer. Falsifiable: run the same cache benchmark on aibeast and compare the knee against our measured prefill curve.
- evidence: run-0003
clm-0008
- community low ●○○ volatility high · verified 2026-08-05
- Identity-grounded persistent sessions with a shared memory and peer-to-peer messaging can self-organise into useful working groups without an orchestrator. Evaluated 2026-08-05 and DEFERRED: the prerequisite is many warm concurrent lanes, which our single-KV-slot architecture cannot provide.
- Source: a white paper describing 8 autonomous OpenClaw sessions ("Council of Minions") self-organising into 3 teams over ~3 hours, coordinated via a shared MongoDB brain and sessions_send, at ~$0.10 of tokens. THE INSIGHT WORTH KEEPING: "an orchestrator coordinates tasks; a council coordinates attention." Their strongest evidence is that nobody had been ASSIGNED to notice auth was added to a new engine and never backfilled to the old one — a task decomposition can only cover work you can name in advance. That limitation is real and matches our own experience: the R1 scorer artefact was caught by reading transcripts, not by any metric we had designed. SCEPTICISM: the headline ROI ($0.10 vs EUR 1,500-4,000 of consultancy) prices output volume as professional deliverable. "79 of 98 routes unauthenticated" is a grep. The cross-verification is agents agreeing with agents, with no independent validation. And the emergence is partly pre-seeded — an identity file that says "Auth, Secrets and Identity Security" makes a credential audit declared rather than emergent. Their gateway was also at ~50% RAM across 555 sessions, which is not a healthy system. WHY WE CANNOT DO IT TODAY, AND IT IS NOT ABOUT POWER: their architecture trades model capability for concurrency; ours does the opposite. MTP forces -np 1, so we have ONE KV slot behind one very large model. Eight concurrent sessions would each cold-prefill and destroy the warm lane. The prerequisite is many CHEAP WARM LANES, not more compute. THE OPERATOR'S EXTENSION (2026-08-05), which is the stronger form: use several DIFFERENT small specialist models rather than many instances of one. The paper ran 8 copies of a single model, so all 8 shared its blind spots — diversity of attention, not of capability. Different models have genuinely different failure modes, which is the same principle underpinning our blind-judging methodology. Our own field showed real profile differences: gpt-oss-120b terse to the point of under-delivery, 35B-A3B best on conceptual quality, 27B strongest on code. FEASIBLE SHAPE once aihydra lands: 122B deep reasoner (aibeast) + 35B-A3B fast generalist and a coder model (aihydra) + a small classifier on the M4 mini. Four genuinely different skills across three machines — heterogeneous by hardware as well as by prompt. ⚑ CORRECTED 2026-08-05 (operator catch). I filed this as a design conflict between self-cycling and Warden's restraint. That was the wrong layer. Restraint lives at the DELIVERY GATE (loop3: daily_cap, quiet hours, validity check, teaser-first), not in deliberation. Deliberation is already separated from delivery by design — lots may happen in the background; the operator is involved only when something of value needs delivering AND it is the right moment. So a self-cycling council in loop2 costs COMPUTE, not attention, and the conflict dissolves. THE RESULTING SHAPE — self-cycling as the first escalation tier: sentinel (cheap, continuous, high-recall, deliberately low-precision) flags a change -> council self-cycles on that one thing, bounded -> exits, and only then does loop3 decide if and when the operator hears about it. THREE EXITS, not two. The operator named "nothing is wrong" and "value delivered"; the third is BUDGET EXHAUSTED, INCONCLUSIVE — and it is the one that bites, because both clean exits are conclusions and real investigations often reach neither. Without it a council does not fail loudly, it quietly eats the box. BOUNDARY PRECISION: the council does not deliver. It produces a CANDIDATE WORTH QUEUEING. Loop3 remains the only path to the operator. A council that believes it can deliver has escaped the gate. FAILURE MODES PER TIER: sentinel too sensitive -> constant convening, cost; too conservative -> the 6h blind spot returns; council unbounded -> runs forever; council concluding "nothing wrong" too eagerly -> misses the thing. OUR OWN INCIDENT IS THE SPECIFICATION: inc-0004 saw the box crash-looping for 46 minutes with three llama-server kills before the kernel panicked, and nothing surfaced any of it — with the risk flagged in that morning's brief. Three process kills in 46 minutes is exactly what a change-detector flags, and "why does this keep restarting" is exactly a bounded investigation with a clear conclusive exit. TRANSFERABLE ENGINEERING (the most useful part): their inter-session sends timed out at a ~10s gateway default against 30-60s cold starts, and the fix was retry-with-backoff plus spill-to-disk staging rather than a larger timeout. That is directly applicable to our own scheduler-preemption work.
clm-0009
- community med ●●○ volatility high · verified 2026-08-05
- A ROCmFP4 iMatrix quant of our exact model (Qwen3.5-122B-A10B) reports 60.70 GiB and 28.505 tok/s decode with MTP OFF — against our measured 20.96 tok/s MTP-off at UD-Q4_K_M. That is roughly +36% decode for ~5-6 GB less memory, if it reproduces.
- Source: vmlinux/Qwen3.5-122B-A10B-ROCmFP4-iMatrix-GGUF on HuggingFace, updated 2026-07-26. Built with ROCmFPX's Q4_0_ROCMFP4_STRIX_LEAN recipe and tested ONLY on gfx1151 — our exact silicon. Published figures: 60.70 GiB, greedy decode 28.505 tok/s, 4,277-token prefill 356.9 tok/s, and BF16 mean KLD 0.041366 +/- 0.002531. NOTABLE ON METHOD: they publish KL divergence WITH an uncertainty interval — the same instrument we selected for KV-quant quality, applied to weight quantisation. That gives us a comparable reference point for what a small quality delta looks like on this model, which we did not previously have. WHY IT MATTERS FOR PLACEMENT: 60.70 GiB against our ~66 GB frees ~5-6 GB. Config 2 in model-phases.md has the 122B and 35B co-resident at ~108 of ~120 GiB, which we flagged as uncomfortably tight given our OOM history. This quant would meaningfully relieve that — possibly the difference between viable and reckless. ⚠ THE COST: it uses custom ROCmFP4 tensor types and **stock llama.cpp cannot load it**. It requires the ROCmFPX runtime. That is a THIRD carried fork alongside con-0001 (our checkpoint sidecar) and the community quantized-KV fix — and our reproducibility debt is already a stated concern. Worse for measurement: a different runtime is a different config fingerprint, so ROCmFP4 numbers CANNOT sit in the same series as our existing ones (protocol section 8). It is a new experiment, not a faster version of the old one. Our Tier 1 pipeline also assumes stock llama-bench, which would need the ROCmFPX build. ALSO: the MTP heads ship as a SEPARATE 2.14 GiB companion file, where our current GGUF embeds them. Flags: "experimental", 1,929 downloads, 6 likes — low adoption, so we would be early rather than following. TEST IT AS A LAB-BOX CANDIDATE, not a production swap: same suite, same context, against our own UD-Q4_K_M baseline, with the runtime difference recorded as the config change it is. If +36% MTP-off holds and MTP scales on top, it plausibly lands near 42 tok/s against our current 36.5.
clm-0010
- community med ●●○ volatility high · verified 2026-08-05
- Draft-free ngram speculation (ngram-mod) reportedly beats MTP by a wide margin on repetitive/agentic work on Strix Halo — 71 t/s unspeculated to 216 solo on a code-edit probe, with a shared hash pool letting concurrent streams feed each other's drafts (247 pooled across 4 streams, later 302 end-to-end on HIP).
- Source: a detailed Strix Halo write-up (HP ZBook, Ryzen AI MAX+ PRO 395, 128 GB, Windows, Vulkan then HIP) on a Qwen3.6-35B-A3B finetune in ROCmFP4 via the ciru-ai/ROCmFPX fork. MECHANISM: the speculator matches n-grams against text already in the context window, so the draft costs nothing to produce and only verify rows are paid. Reported 86% acceptance on 48-64 token copied drafts, and on fresh content the gate simply stays closed with no penalty. That asymmetry is what makes it different from a draft model. WHY IT MATTERS TO US MORE THAN MOST: our heaviest repetitive workload is exactly their best case — Claude Code doing code edits, re-emitting large chunks of a file. And the shared-pool effect (acceptance RISING with stream count, 93% to 97%) is a genuine argument for the council pattern in clm-0008, where several agents work adjacent problems. ⚠ MTP IS THROUGHPUT-NEGATIVE AT BATCH per the same source: roughly breakeven at 2 streams, -13% at 4, because every accepted draft still costs a verify row through the MoE and the expert-union already saturates the bus. We run -np 1, so we are in MTP's favourable regime — but this is a further argument against multi-slot, on top of llama.cpp #25992 and the memory cost. CAVEATS: single reporter; the 4-stream figures used IDENTICAL prompts across streams, which is maximum pool sharing and flattering (they say so). Requires the ROCmFPX fork, which would be a third carried runtime (see clm-0009). Their own updates report ngram fighting flash attention at larger contexts, and crashing when combined with MTP. TEST: ngram-mod vs MTP vs neither, on OUR code-edit workload, single stream. It is the cheapest large win available if it reproduces.
clm-0011
- community med ●●○ volatility medium · verified 2026-08-05
- Greedy decode is NOT run-to-run deterministic on the Vulkan backend — same config, same prompt, temperature 0, batch 1, no speculation, three different outputs.
- Reported on Strix Halo/gfx1151 with outputs diverging mid-generation. Deep MTP (n>=4) amplifies it to first-token divergence via draft-length to batch-shape variance. The reporter is explicit that outputs stay coherent — this is variance, not corruption. DIRECT CORRECTION TO OUR PROTOCOL: benchmark-protocol.md adopted "temperature 0, fixed seed" as a determinism control, borrowed from homebench. On this backend that control does not hold, which has two consequences. First, N=1 is not defensible for anything output-dependent even where we assumed determinism. Second, and worse: **speculative losslessness cannot be verified by diffing outputs on Vulkan**, which is exactly how one would naturally check that MTP or ngram speculation is not changing results. UNVERIFIED HERE, and worth checking on ROCm rather than assuming it transfers — the report is Vulkan-specific and our production backend is ROCm. If it holds on ROCm too, quality comparisons need a distributional instrument (KL divergence) rather than output equality, which is the approach we had already chosen for KV quant.
clm-0012
- community low ●○○ volatility high · verified 2026-08-05
- 128 GB Strix Halo systems appear to be repricing upward materially — from a roughly $2,500-4,000 enthusiast tier toward $4,500-6,000 — with 128 GB SKUs scarce while 64 GB configurations remain available.
- Source: a community PSA aggregating six vendors (Bosgame, GMKtec, ACEMAGIC, Framework, Corsair, MSI). Direction is well-evidenced across independent listings; MAGNITUDE is not — the post itself flags ACEMAGIC's $4,999 as possibly a placeholder, and Corsair's jump came with reopened preorders, which could be a reset rather than a trend. ROOT CAUSE MATCHES WHAT WE FOUND ELSEWHERE: this is the same memory squeeze that delayed NVIDIA's RTX 50 Super refresh to CES 2027, where 3 GB GDDR7 modules reportedly cost $60-70 against $20 for 2 GB parts. Here it lands on high-density LPDDR5X for the 256-bit bus. The memory being SOLDERED is what gives manufacturers pricing power specifically on the high-capacity SKUs — a buyer cannot take a cheaper configuration and upgrade later — and it explains 64 GB staying in stock while 128 GB does not. WHAT IT CHANGES FOR US, which is why it is recorded at all: · A THIRD box becomes unlikely, so the two-box architecture is a durable design rather than a transitional one. · It WEAKENS the earlier "wait for Medusa Halo (LPDDR6, ~460 GB/s, 2027)" argument. A part needing even denser memory will be priced in the new tier, not the old one. · It STRENGTHENS the cross-host overlap case in the integration plan. That was justified on "most people replace rather than overlap, so this comparison is rare". If replacement is now expensive, keeping both boxes long-term is likely — and aibeast's 96 GB appreciates in usefulness rather than depreciating.
clm-0013
- community high ●●● volatility medium · verified 2026-08-05
- Qwen3.6-35B-A3B — our chosen reflex model — is published in FastFlowLM's .q4nx NPU format (23.2 GB, plus a 1.0 GB vision encoder), making a genuinely capable MoE, not a 1-2B classifier, runnable on the XDNA2 NPU.
- FastFlowLM/Qwen3.6-35B-A3B-NPU2, created 2026-07-08, ~2,000 downloads. WHY A 35B IS SUDDENLY PLAUSIBLE ON AN NPU: decode bandwidth is proportional to ACTIVE parameters, not total. At A3B the model reads roughly 2 GB per token at 4-bit, where a dense 35B would read ~20 GB. Low-active-parameter MoE is the ideal shape for a bandwidth-poor accelerator — which is exactly why the NPU catalogue was previously 1-8B dense models. WHAT IT CHANGES: the second-lane argument in npu-fastflowlm-2026-08-02 assumed the NPU could host a small classifier. If it can host our actual reflex model, the sentinel tier in clm-0008 could be a capable council member rather than a change-detector. It also ships a vision encoder, which could relieve the Mac mini. ⚑ CORRECTED 2026-08-05 with measured Strix Halo data (Framework Desktop, gfx1151, LMDE 7 / xanmod 7.1.3). I had framed the NPU as "slower at decode but wins prefill, so the value is concurrency". **On a 35B, the NPU loses BOTH phases decisively.** | ctx | NPU decode | NPU prefill | NPU TTFT | |-----|-----------|-------------|----------| | 1k | 12.11 t/s | 90.7 t/s | 10.8 s | | 8k | 10.94 | 209.9 | 36.9 s | | 32k | 8.11 | 245.2 | **126.3 s** | Same box, same model at a HIGHER quant (Q8_0, 35.21 GiB) on the iGPU: **pp4096 709-840 t/s on ROCm, 964-1039 on Vulkan; tg128 43 t/s ROCm, 53 t/s Vulkan.** So the iGPU is roughly **4x faster at decode AND 4-6x faster at prefill**, while running a heavier quantisation. My earlier "NPU wins prefill 2.3x" came from an ~8-9B dense comparison; that advantage does not survive at 35B on this silicon. And TTFT of **126 seconds at 32k** rules the NPU out of anything interactive or escalation-shaped outright. WHAT SURVIVES: the NPU lane is only defensible as genuinely CONCURRENT background work where latency is irrelevant — Phase A/D in model-phases.md — and even then it costs the iGPU about 14% of its own decode. The "capable second council member" framing is much weaker than it looked an hour ago: a member that answers in two minutes is not a peer. ALSO STILL TRUE: mutually exclusive with amd_iommu=off (+5-12%), and aibeast's NPU logged SMU init errors on the 21 July wedged boot, so NPU health is unverified.
clm-0014
- community med ●●○ volatility high · verified 2026-08-05
- On gfx1151, Vulkan measured ~22-24% faster than ROCm on the same 35B-A3B model and build — pp4096 1039 vs 840 t/s, tg128 53.1 vs 43.4 — but the Vulkan build produced GARBAGE OUTPUT for that model, so the numbers describe a broken configuration.
- Measured on a Framework Desktop (Strix Halo, Q8_0, 35.21 GiB) with llama-bench. The reporter states plainly that the current llama.cpp Vulkan build returns a single garbage Chinese character when actually asked anything with Qwen3.6. ⚠ THE METHODOLOGICAL POINT MATTERS MORE THAN THE NUMBERS. **llama-bench measures throughput and never checks that the output is valid.** A backend can be 24% "faster" while emitting nothing usable, and the benchmark cannot tell. Our Tier 1 is built on llama-bench, so we inherit that blind spot exactly. Fix adopted: an OUTPUT SANITY GATE before any throughput number is recorded — one real generation per (model, backend, build), checked for coherent text, and the throughput discarded if it fails. Cheap, and it is the difference between "Vulkan is faster" and "Vulkan is broken but quick about it". SECONDARY, AND DIRECTLY USEFUL: the same runs sweep ubatch. Prefill peaks around **ub 1024-2048** on both backends (ROCm 715 -> 833 t/s from 512 to 2048; Vulkan 964 -> 1039 at 1024) and DEGRADES at 4096. Decode is flat across ubatch, as expected. That narrows our own planned sweep to a sensible range rather than guessing.
clm-0015
- community high ●●● volatility high · verified 2026-08-06
- Ling-3.0-flash (Ant Group, 26 Jul 2026) is a 124B/5.1B-active hybrid MoE that is a near-ideal controlled comparison for our Qwen3.5-122B-A10B — same footprint, half the active parameters — but it cannot be benchmarked on a stock llama.cpp, and its MTP head ships INACTIVE, which would silently rig any decode comparison in Qwen's favour.
- STRUCTURAL FACTS (from the GGUF model card and the upstream PR; high confidence). 124B total, **5.1B active** per token — note 5.1, not 5.2. 42 layers: **35 KDA (Kimi Delta Attention) + 7 gated MLA** at layers 5/11/17/23/29/35/41, plus an MTP/nextn head at layer 42. Native 256K context; the GGUFs expose 262,144 tokens, double the 131,072 in config.json. Q4_K_M = **69.70 GB**, Q5_K_M = 84.18 GB, Q8_0 = 126.32 GB. NO PERFORMANCE CLAIM HERE IS VERIFIED. The vendor line — matching their 1T flagship on most benchmarks at 1/8 the total and 1/12 the active parameters — is marketing, and every number in it is exactly what we would be measuring. Recorded as motivation for a benchmark, not as a result. WHY IT IS WORTH THE SLOT: against Qwen3.5-122B-A10B this is almost a controlled experiment. 124B vs 122B total (near-identical memory footprint), 5.1B vs 10B active (half), both MoE, both hybrid, both carrying MTP heads. It isolates ACTIVE PARAMETER COUNT more cleanly than any other pairing available to us. ⚠ THE TRAP — THE MTP HEAD IS PRESENT BUT "UNUSED BY THE GRAPH". It loads without consuming VRAM and can be enabled later without requantisation, but today it is inert. Our production Qwen runs MTP and we measured +50% decode from it (mtp-single-slot decision, 2026-07-20). So a naive Ling-vs-Qwen decode comparison is **MTP-off versus MTP-on**, and would flatter Qwen by roughly the size of the effect we are trying to measure. This promotes the already-open "MTP-off Qwen build for the ratio" question from nice-to-have to a PRECONDITION for this comparison being meaningful at all. THE HYBRID CONFOUND, NOW QUANTIFIED. bench/kv-quality.sh warns that -ctk/-ctv only touch full-attention layers, so a near-zero KL divergence would be an ARCHITECTURE result rather than a quantisation result. Here the number is **7 of 42 layers**. KV-quant should barely register, and we can state that up front instead of discovering it. The corollary is the attractive part: with 35 of 42 layers holding constant-size recurrent state, KV growth at 256K should be far below what our current model demands. BUILD STATUS — NOT UPSTREAM. Files declare `general.architecture = bailing-hybrid`, a provisional name; upstream PR #26608 (opened **2026-08-05**, unmerged) proposes `bailingmoe3`, continuing bailingmoe → bailingmoe2. Stock builds refuse the model with "unknown model architecture". A patch against llama.cpp commit 6ea215d17 ships with the GGUFs; a fork exists at aetherbird/llama.cpp:bailingmoe3-support. Because the arch name is provisional, **today's GGUFs may need re-downloading** once upstream lands — 70 GB of reason not to rush. PROTOCOL CONSEQUENCE 1 — fingerprint. `build_commit` is a CONFIG_KEY. A patched-build Ling result and our patched-build Qwen result come from two DIFFERENT forks, so they are not the same experiment and must not be presented as one row set without saying so. PROTOCOL CONSEQUENCE 2 — the sanity gate is not strong enough here. The GGUF notes `rope_interleave: true` resolving to NORM rope rather than the NEOX that DeepSeek-style MLA usually uses. A subtly wrong rope on a fork build produces output that passes sweep.sh's 24-token coherence check and then degrades at depth. **For any unsupported-architecture build, the precondition should be a long-context retrieval check (RULER-lite), not a short coherence check.** This is a gap in our own protocol that this model exposed. FIT, AND AN ARGUMENT FOR THE 128 GB BOX. Q4_K_M at 69.7 GB fits aibeast's 96 GB with room for KV and compute — helped by only 7 layers holding a real cache. Q5_K_M at 84.18 GB does not fit aibeast's ~84 GiB GTT but is comfortable on aihydra's 128 GB. That is the first concrete case where the larger box buys a quantisation level rather than just headroom. UNVERIFIED RISK: the published quants use "SM_120-safe" types chosen around NVIDIA Blackwell (no iq1_s/iq2_s/iq3_s). Irrelevant to gfx1151 in intent, but it means these files were built and tested against CUDA, and their behaviour on ROCm is untested. A GAP WE COULD ACTUALLY FILL: the model card states reasoning quality beyond 128K is **unmeasured**. We have the KL-divergence and RULER-lite tooling and a box that fits 256K. That is a contribution-shaped hole, not just a benchmark. RECOMMENDATION: add to the bench queue, do NOT chase the fork. aibeast is dead — the 2026-08-07 hands-on attempt found a board that will not POST and a warranty claim is open (inc-0005), so the benchmark host is now aihydra. A second fork alongside our #25913 patch is real maintenance cost. Wait for #26608 to merge and fold it into the llama.cpp rebuild already queued for return (Step 1b) — that gets the stable arch name, the final GGUFs and the support in one move.
clm-0016
- community med ●●○ volatility medium · verified 2026-08-06
- Measured on Strix Halo: Qwen3.6-35B-A3B under ROCmFP4 + HIP + ngram-mod at parallel 4 sustains 121 tok/s across 500 varied IFEval prompts against a 64.8 tok/s no-speculation floor — a 1.87x production speedup with IFEval-strict at 78.6% — while prefill collapses from 1,211 tok/s cold to 136 tok/s at ~243k depth.
- Single community source (the ACE-SABER follow-up post), self-reported, on someone else's Strix Halo. Confidence is medium DESPITE the detail, because it is one operator on one box and none of it is independently reproduced. The internal arithmetic does check out: 1.32M tokens / 3.0 hours = 122 tok/s, consistent with the stated 121. CONFIGURATION (all five are fingerprint fields, so this is ONE config point, not a decomposition): ROCmFP4 quant, HIP backend, ngram-mod speculation with shared hash pool, --parallel 4, f16 KV. THE NUMBER THAT MATTERS: **121 tok/s sustained vs a 64.8 tok/s no-speculation floor = 1.87x**. Both figures come from the same operator under the same conditions, which makes this a genuine isolation of the speculation effect rather than a headline. The author explicitly labels 121 as "the production number". WHAT THE AUTHOR DISCARDS, AND WHY IT MATTERS MORE THAN WHAT THEY KEEP: 430 tok/s (same-prompt repeat) is flagged as an artifact — only one session doing real work while three ride the ngram pool. 302-305 tok/s (4 streams, identical content) is flagged as "inflated by max pool sharing". Only 380 tok/s (single stream, real agentic session) and 121 tok/s (sustained, varied) are claimed as trustworthy. This is exactly the discipline our protocol demands and it is why this post is usable at all. ⚠ IT ALSO CONFIRMS A HAZARD WE ALREADY FLAGGED. model-phases §3a notes that ngram-mod's shared pool only works within one llama-server, so isolation and speed conflict. This data shows the pool ALSO corrupts measurement: identical content across slots inflates throughput because the slots feed each other's n-gram cache. **Any multi-slot benchmark of ngram-mod must use VARIED prompts or it measures the pool, not the model.** Our sweep.sh does not currently guard against this. PREFILL DECAY AT DEPTH — THE FINDING WITH THE MOST CONSEQUENCE FOR US: | prompt | prefill | |---|---| | 8.1k cold | 1,211 tok/s | | 2.2k blob at ~3k depth | 1,186 tok/s | | 2.2k blob at ~13k depth | 899 tok/s | | 2.2k blob at ~243k depth | **136 tok/s** | Read carefully: that last row is the MARGINAL cost of appending 2.2k tokens when 243k are already resident — roughly **16 seconds of prefill before a single token is generated**, every turn, at depth. It is not the cost of prefilling 243k from scratch. This is the strongest external validation yet of the warm-lane and disk-restore work: our measured restore was 105 ms for 44k tokens, so restore-versus-reprefill is the difference between milliseconds and tens of seconds per turn. TTFT corroborates it — 6.7 s for a cold 8.1k prompt versus **71 ms** for turn 2 on a cached prefix. QUALITY — AND WHAT THE IFEval NUMBER ACTUALLY TESTS. IFEval strict 78.6% / loose 79.8%, reported as no measurable drop. Worth being precise about what that establishes: correctly implemented speculative decoding is mathematically exact — the verification step reproduces the base model's distribution — so an unchanged score is the EXPECTED result and mainly evidences that the ngram implementation is not lossy. The more interesting half is that it also passes under **ROCmFP4**, which is a genuine quantisation change and could have cost accuracy. Note too that IFEval measures instruction-following only: it says nothing about long-context retrieval at 243k, nor about tool calling, which are the two axes we care most about. The author says more benchmarks are coming. SCOPE LIMIT — 27B GETS NO SIMILAR BOOST. Reported plainly, and the mechanism is plausible: speculation pays in proportion to how memory-bound decode is, because batch verification of drafted tokens is nearly free when you are bandwidth-starved rather than compute-starved. A3B activates 3B; a dense 27B activates all 27B and is far more compute-bound, so there is less idle bandwidth for speculation to exploit. **Prediction worth testing rather than assuming:** the benefit should scale inversely with active parameters — largest at Ling-3.0-flash (5.1B active, clm-0015), smaller but real on our Qwen3.5-122B-A10B (10B active), negligible on dense models. VULKAN, SECOND INDEPENDENT REPORT: "Vulkan? Nope, has a bug that's kicking in, HIP is the fix." This is now the second unrelated account of Vulkan misbehaving on gfx1151 with a Qwen3.6 model, alongside our own note about a backend being faster while emitting garbage. Two sources is not proof, but it is enough to make HIP the default for our sweeps and to treat any Vulkan throughput advantage on this silicon as suspect until output is verified. WHAT WE SHOULD TAKE FROM IT: a quality anchor (IFEval strict 78.6%) and a throughput target (121 tok/s sustained, 64.8 floor) that our own runs can be compared against, on the same silicon we own — provided we state the five-way config difference rather than presenting it as a like-for-like row.
clm-0017
- community high ●●● volatility high · verified 2026-08-06
- On Strix Halo Vulkan, a single unmerged patch — contiguizing strided f16 KV data before the flash-attention prefill — removes a dense-model prefill collapse worth 2.5x at 32k and 6.6x at 65k, changes decode not at all, and collapses run-to-run scatter from 5-10% to 0.3%. The variance change is the more useful finding: high scatter at depth is a SYMPTOM of a broken code path, not noise to be averaged away.
- Single community source, but the confidence is high because of the METHOD rather than the reputation: five binaries built from the same upstream base, changing only which patches were present, same cells on each. That is proper attribution — the variable is isolated rather than argued about. It is the standard our own §9 asks for and rarely sees in the wild. | dense 27B, f16 KV, FA on, Vulkan | pp512 @ 32k | pp512 @ 65k | |---|---|---| | stock upstream | 94.5 t/s | 29.8 t/s | | nine patches, contiguize REMOVED | 134.5 | 45.1 | | stock + contiguize and prerequisites | 237.4 | 180.4 | | all nine patches | 252.0 | 198.2 | Removing one patch from the set reinstates the collapse; adding it to an otherwise plain build removes it. Everything else in the set is worth a few percent. Baseline for scale: ~350 t/s at empty context, so stock loses >90% of prefill purely for holding a long conversation. ⚑ THE FINDING THAT CHANGES HOW WE MEASURE, not just what we build. On the broken path identical repetitions of the same deep cell varied by **5-10%**; on the fixed path they reproduce within **0.3%**. The author reports this retroactively explained a 20% standard deviation they had flagged months earlier and never understood. We currently treat `--reps 3` as a way to average noise out. That is exactly wrong when the variance is the signal. **A cell whose repetitions disagree by more than a few percent at depth should be treated as evidence of a broken code path and investigated, not reported as a clean mean with an error bar.** Actioned: sweep.sh now computes and reports per-cell coefficient of variation and flags high-scatter cells; protocol §0 carries the rule. DECODE IS UNAFFECTED — within 0.09% across all five builds. So this is purely a prefill path issue, which also means it would be invisible to any benchmark that only reports tokens/sec of generation. HOW MUCH OF THIS TRANSFERS TO US — genuinely uncertain, and the differences are large: · **Backend.** Measured on VULKAN. The author dropped ROCm entirely because Fedora 44 version locks blocked the upgrade, so they CANNOT test the HIP path. Whether an analogous collapse exists on ROCm is open. · **Architecture.** Measured on a DENSE 27B. Our models are MoE and mostly hybrid. For Ling-3.0-flash only 7 of 42 layers hold a real KV cache (clm-0015), so a flash-attention prefill defect would touch a minority of layers and the effect size should be much smaller — but "smaller" is not "zero", and we run `-fa 1`. · **Our own unexplained decay.** clm-0016 records prefill falling from 1,211 t/s to 136 t/s at ~243k depth on HIP with a 35B-A3B MoE. Different backend, different architecture, but the same SHAPE. Worth testing as a hypothesis rather than assumed related: is there an analogous contiguity problem on the HIP prefill path? IT ALSO COMPLICATES "VULKAN IS BROKEN". We hold two reports of Vulkan misbehaving on gfx1151 (clm-0014 garbage output, clm-0016's author reporting a bug that HIP avoids). This author runs Vulkan exclusively and gets good results once patched. The honest reading is that these are DIFFERENT defects — a model-specific correctness bug and a general FA prefill inefficiency — and we should stop collapsing them into "prefer HIP". FLASH ATTENTION IS NOW A LEVER, NOT A SETTING. The author's practical advice: on a patched build FA wins every cell and can be left on; on stock the old caution still applies. Our sweep.sh hardcoded `FA=1` with no way to vary it, which on a stock build means silently measuring the collapse. Actioned: `--fa` is now a sweepable argument. FORK COST. The patch is unmerged and under discussion upstream. Adopting it would make a THIRD fork alongside our #25913 .ckpt sidecar and a possible bailingmoe3 build — and these would have to be combined, not merely chosen between. Each fork is a permanent reproducibility tax: `build_commit` is a fingerprint field, so a combined tree is a configuration nobody else can reproduce. A CONTRIBUTION WE ARE UNUSUALLY WELL PLACED TO MAKE. The author explicitly asks for replication on stock builds — depth 32k+, FA explicitly on, three runs, because a single run tells you little on the broken path. We will shortly have two Strix Halo boxes and, critically, **ROCm — which the author no longer has**. The unanswered question is not "does this reproduce on Vulkan" but "does the collapse exist on the HIP path at all", and we are positioned to answer exactly that. Qwen3.6-27B is dense and available, so we can replicate the architecture rather than only approximating it with an MoE. Verified 2026-08-06 that the branch `strix-halo-fa-fixes` exists on the Nathanw1014/llama.cpp fork. Individual commit contents were NOT verified — the patch is described here as the author describes it, not as we have read it.
clm-0018
- community med ●●○ volatility high · verified 2026-08-06
- DeepGrove Maple-Preview (20.2B-A1.49B ternary, 5.31 GB, MIT) is a credible second-lane candidate for the Mac mini, but not a Warden candidate — its own model card concedes underperformance on agentic benchmarks. The "on-device weight adaptation (dreaming)" feature is undocumented at source, and even if it works it is architecturally opposed to how Warden's state is built.
- PROVENANCE MATTERS HERE MORE THAN USUAL — the claims arrive at three different levels of evidence and the most interesting one is the weakest: · **Model card (primary).** 20B-A1B, 24 layers, 256 experts / 8 active, ternary weights 2-bit packed as {-a, 0, +a} with one a per row, 5.31 GB checkpoint, 131,072 context, 3:1 SWA-512:global attention, MIT licence. Quantisations exist for llama.cpp, Ollama, LM Studio, Jan. Apple Silicon runs via the authors' own MLX fork (deepgrove-ai/mlx-lm-deepgrove). · **Model card, self-reported, NO NUMBERS.** Evaluated on LCBv6, AIME 2026, HMMT 2026, GPQA-D — but published as a *visual comparison chart* with no scores. Claims a new point on the Pareto frontier for memory-to-performance. · **Launch demo on X + aggregator blog only.** The "dreaming" weight adaptation, the ~5.9 GB adaptation peak, and the 200+ tok/s Mac mini figure. The model card says **nothing** about weight adaptation. The company describes adaptive behaviour partly in FUTURE TENSE ("plans to enhance"), so this is a demo and a roadmap, not a released, reproducible capability. 419 downloads in the last month — very low adoption. For scale, the NPU build in clm-0013 had ~2,000. Nothing here has been independently reproduced. ⛔ DISQUALIFYING FOR WARDEN'S PRIMARY ROLE, on the authors' own evidence. The model card acknowledges **underperformance on agentic benchmarks** and notes minimal post-training for agentic tasks. Warden is an agentic system whose stated failure mode is tool selection. A model that is weak at exactly that is not a candidate for the main lane, however good its reasoning scores turn out to be. This is the rare case where the vendor tells you the disqualifying fact themselves. ✅ WHERE IT IS GENUINELY INTERESTING — THE MAC MINI. wardenmac is an M4 Mac mini with 16 GB currently carrying only vision (MLX) and STT; no LLM lane. A 5.31 GB checkpoint at a claimed 200+ tok/s would fit with enormous headroom, and MLX is a path already proven on that box. That makes it a candidate SECOND LANE that does not depend on aibeast at all — which is the point, because the single-slot bottleneck is an aibeast problem. It is a more plausible route than the XDNA2 NPU option (clm-0013), where measured data showed the iGPU winning both phases by ~4x. ON "DREAMING" — THE NAME COLLIDES WITH OURS AND THE MECHANISM IS OPPOSITE. Warden's dreaming consolidates recalls into MEMORY.md: text, dated, diffable, revertible. Maple's dreaming runs a local fine-tune that embeds a preference **into the weights**. The architectural appeal is real — persistent preference that costs no context is a direct attack on our dominant cost, since the entire warm-lane programme exists because prefill of a large persistent prompt is expensive. But it is opposed to the principle the whole system is built on: **legible, auditable, revertible state.** `anticipatory-flags.md` declares the operator as its owner. Every recovery this month depended on being able to read state and diff it — the 2026-08-06 audit found a SESSION-STATE that had asserted a wrong host address for six weeks, and it was fixable precisely because it was text. A preference baked into weights cannot be inspected, diffed, dated, or switched off with a flag; you would discover a wrong one only by noticing bad behaviour, and then have no way to attribute it. You cannot `git diff` a weight delta into an explanation. If we ever pursue this, the defensible split is: weights may hold **stable, low-stakes, stylistic** adaptation (tone, formatting); anything with consequences stays in text. And even then the audit story is poor enough that it should be a deliberate experiment on a lab box, never on the production agent. TERNARY CAVEAT: replacing multiplies with additions is real, but the speedup needs kernels that exploit it. Community llama.cpp quants exist (e.g. stamsam/maple-preview-gguf), but ternary support in llama.cpp is narrower than standard K-quants and the Apple Silicon numbers come from the authors' own MLX fork rather than a neutral runtime. Treat the 200+ tok/s figure as vendor-path until measured on a stock runtime. IF WE TEST IT, the honest first question is not speed but capability: does a 1.49B-active model actually hold up? That is precisely what the capability tier exists for — and per protocol §1a it must be measured, not inherited, since this differs from everything we run by architecture, quant class and runtime simultaneously.
clm-0019
- measured-here high ●●● volatility medium · verified 2026-08-08
- On aihydra (gfx1151, ROCm 7.1, llama.cpp 3653e6d), running llama-server with `--parallel 4` destroys long-context needle retrieval — 0/8 across controlled trials — while `--parallel 1` on the same build, model and prompt succeeds 8/8. Throughput is unaffected and reports nothing wrong, so no performance benchmark would ever see it.
- MEASURED HERE, controlled, same session, same model file, same server binary, only the slot count varying: | server flags | needle before concurrency | after | |---|---|---| | `-c 65536 --parallel 4` (cache-idle-slots default) | 0/3 | 0/5 | | `-c 65536 --parallel 4 --no-cache-idle-slots` | 0/3 | 0/5 | | `-c 16384 --parallel 1` | **3/3** | **5/5** | Probe: a needle at position 0 ("The maintenance codeword is chartreuse-viper-88."), ~5,368 prompt tokens of whole-sentence filler, question at the end, `temperature 0`, thinking disabled. Failure mode is not garbage — the model answers coherently *from the filler* ("The maintenance codeword is: **routine**") or echoes the instruction ("codeword"). It behaves exactly as if the first part of its context is absent. WHY IT MATTERS MORE THAN IT LOOKS: · **Throughput is unaffected and silent.** decode/prefill numbers at `--parallel 4` look entirely normal. A sweep-only benchmark would have published them. · **It is a `hazard`-class lever behaving exactly as protocol §1a predicts** — a performance setting silently corrupting correctness. This is the first time our own guard framework has caught a real one, and it is the whole argument for the guard tier existing. · **It threatens the multi-slot plans directly** — the second-lane idea, the council architecture, and any use of parallel slots for concurrent agents. ⚠ DISTINCT FROM llama.cpp #25992. That issue is cross-request *leakage* between concurrent requests. We tested for it explicitly — four concurrent requests with disjoint markers — and found **no contamination**, repeatedly. This is a different failure: single sequential requests on a multi-slot server lose their own early context. HONEST ABOUT THE MESSY PATH: earlier ad-hoc observations in the same session were inconsistent (the same prompt passed 10/10 on one server instance and failed 6/6 on another with the same flags), and I twice attributed failures to my own harness before running a controlled comparison. The three-way matrix above is the only evidence that should be relied on; the earlier anecdotes are recorded here only so nobody re-derives them and thinks they contradict this. NOT YET ESTABLISHED — the open questions that would make this reportable upstream: · Does it depend on slot COUNT (2? 8?) or merely on >1? · Does it depend on `n_ctx_slot`, or on depth relative to it? Our failures were at ~5.4k tokens against a 16,384-token slot, so it is not a simple overflow. · Does it reproduce on CPU or Vulkan, or is it HIP/gfx1151-specific? · Does it reproduce on a stock upstream build on other hardware? If yes, this is an upstream bug worth filing; if no, it is a gfx1151 backend issue. IMMEDIATE OPERATIONAL CONSEQUENCE: **run `--parallel 1` for anything that depends on long-context recall** until the above is answered. That is what production already did (`cfg-0002` used `--parallel 1`), so nothing shipped is affected — but the plan to use multiple slots for a second lane is on hold pending this.
clm-0020
- measured-here high ●●● volatility medium · verified 2026-08-08
- On ROCm/gfx1151 with a stock llama.cpp, flash attention is unambiguously BETTER at depth — at 32k it is worth 1.22x prefill and 1.84x decode on the 122B MoE — which is the opposite of the Vulkan cliff reported in clm-0017. Quantised KV on a STOCK build costs 16% decode at 32k versus f16 and cannot create a context at all without flash attention — which is what clm-0006 predicts for stock, and makes testing its fix the highest-value remaining experiment.
- Qwen3.5-122B-A10B UD-Q4_K_M, aihydra, ROCm 7.1.0, llama.cpp 3653e6d (STOCK — no fork, no patches), `--load-mode none`, 3 reps. Full rows ingested as run-0008..run-0023. | depth | fa | KV | pp512 | tg128 | |---|---|---|---|---| | 0 | 0 | f16 | 312.04 | 21.68 | | 0 | 1 | f16 | 321.72 | 21.91 | | 4096 | 0 | f16 | 283.47 | 19.26 | | 4096 | 1 | f16 | 300.67 | 21.39 | | 32768 | 0 | f16 | 170.81 | **9.87** | | 32768 | 1 | f16 | 208.77 | **18.17** | | 32768 | 1 | q8_0 | 204.49 | 15.26 | **FLASH ATTENTION AT 32k: 1.22x prefill, 1.84x decode.** Without it, decode collapses from 21.68 at empty context to 9.87 at 32k — it loses 54% of its speed just for holding a long conversation. With it, the same span costs only 17% (21.91 → 18.17). ⚑ THIS IS THE OPPOSITE OF clm-0017, AND THAT IS THE POINT. That report — dense 27B, **Vulkan**, stock build — found flash attention CAUSING a prefill collapse at depth, fixed only by an unmerged contiguize patch. Its author explicitly could not test ROCm, having lost it to a Fedora 44 upgrade, and asked for exactly this replication. On the HIP path, on a stock build, **there is no cliff — FA is the thing preventing one.** Scope honestly: this is a hybrid MoE, not the dense 27B they used, so architecture is uncontrolled. A dense-model run is in flight and will settle whether the difference is the backend or the architecture. Either answer is worth reporting back. PRACTICAL: leave `-fa on`. It wins every cell measured here and wins hugely at depth. QUANTISED KV — TWO FINDINGS, BOTH NEGATIVE: · **q8_0 costs 16% decode at 32k** (15.26 vs 18.17). At empty context the two are identical (21.63 vs 21.91), so the penalty is depth-dependent and would be invisible to a shallow benchmark. · **q8_0 with flash attention OFF cannot create a context at all** — llama-bench exits with "failed to create context". Quantised KV *requires* FA on this build. Recorded as a failed cell against its fingerprint rather than retried into a pass (protocol §9). ⚑ CORRECTED 2026-08-08 (operator catch). I first wrote that these results "do not support clm-0006". That is backwards. clm-0006's claim is that **stock llama.cpp dequantizes the KV cache to full precision repeatedly during inference on this silicon**, and that a community FIX then makes q8_0 run 23-53% FASTER than f16. **We are on a stock build.** So measuring q8_0 as *slower* is exactly what clm-0006 predicts for stock — it CONFIRMS the diagnosis and says nothing at all about the fixed build, which we have not tested. That inverts what this result means. Rather than deflating clm-0006, it raises the value of testing the patch: stock costs 16% decode at 32k, the fix claims +23-53%, so the available swing at depth is roughly 40-70%. **That is potentially larger than MTP's 1.45x**, and it compounds with it rather than competing. AND IT IMPLICATES PRODUCTION. `cfg-0002` ran **q8_0/q8_0 KV at 200,000 context** on a fork carrying only the `.ckpt` sidecar fix (`con-0001`) — nothing touching KV. So production sat on the slow dequantisation path, at the deep end where clm-0006 says the penalty is largest. Together with the `--spec-draft-n-max 6` finding above, that is two independent production settings measurably below optimum. q8_0 remains a memory-saving lever on a stock build, not a speed one — but on a patched build that may reverse entirely, and it is now the highest-value untested experiment. SCATTER: every surviving cell reproduced within **1.3% CV**, most under 1%. Per clm-0017 that signature indicates a healthy code path; the broken Vulkan path scattered 5-10%. Our sweep now reports CV per cell precisely so this is checked rather than assumed.
clm-0021
- measured-here high ●●● volatility medium · verified 2026-08-08
- The dense-model flash-attention prefill cliff reported in clm-0017 does NOT exist on ROCm. A stock ROCm build of a dense 27B reaches 214.66 t/s prefill at 32k and 153.63 at 65k — close to their PATCHED Vulkan numbers (237.4 / 180.4) and 5.2x their broken stock Vulkan at 65k (29.8). Flash attention is not the problem on HIP; it is what prevents one, worth 3.4x decode at 65k.
- THIS IS THE REPLICATION clm-0017's AUTHOR ASKED FOR AND COULD NOT RUN. They dropped ROCm because Fedora 44 version locks blocked their release upgrade, leaving them Vulkan-only and unable to answer whether the collapse exists on the HIP path. It does not. Qwen3.6-27B Q4_K_M (dense), aihydra, ROCm 7.1.0, **stock** llama.cpp 3653e6d, f16 KV, `--load-mode none`, 3 reps: | depth | fa | pp512 | tg128 | |---|---|---|---| | 0 | 1 | 355.52 | 12.04 | | 0 | 0 | 353.93 | 11.95 | | 32768 | 1 | **214.66** | 10.86 | | 32768 | 0 | 175.87 | **4.76** | | 65536 | 1 | **153.63** | 9.90 | | 65536 | 0 | 119.08 | **2.92** | SIDE BY SIDE WITH THE COMMUNITY REPORT (their dense 27B, f16 KV, FA on, pp512): | | @32k | @65k | |---|---|---| | Vulkan STOCK (theirs, broken) | 94.5 | 29.8 | | Vulkan PATCHED (theirs) | 237.4 | 180.4 | | **ROCm STOCK (ours)** | **214.66** | **153.63** | Our unpatched ROCm lands near their patched Vulkan and **5.2x their unpatched Vulkan at 65k**. Whatever the contiguize patch repairs, the HIP path does not suffer from it. FLASH ATTENTION IS LOAD-BEARING FOR DECODE, NOT PREFILL. At empty context FA is irrelevant (355.52 vs 353.93 prefill, 12.04 vs 11.95 decode — both within noise). The divergence is entirely depth-driven, and decode is hit far harder than prefill: · prefill at 65k: 1.29x with FA · **decode at 65k: 3.4x with FA** (9.90 vs 2.92) Without FA, decode falls 76% from empty context to 65k (11.95 → 2.92). With it, 18%. A benchmark that only reported prefill would have understated this badly. SCATTER: every cell reproduced within **0.6% CV**. Per clm-0017 the broken path scatters 5-10% while a healthy one holds ~0.3% — our numbers carry the healthy signature, which is independent corroboration that we are not on the defective code path. UNCONTROLLED: clm-0017 says "dense 27B" without naming the model; we used Qwen3.6-27B. If theirs was a different 27B, architecture is not perfectly matched — but the effect size (5.2x) is far larger than any plausible model-to-model difference. CONSEQUENCE, COMBINED WITH clm-0020: leave `-fa on` everywhere on this hardware. It wins at every depth measured, on both a dense 27B and a 122B hybrid MoE, and its absence is catastrophic for decode at depth. The clm-0017 caution applies to Vulkan only. WORTH REPORTING BACK. The author explicitly requested this test. Answering it costs us nothing and it resolves whether their patch is a general fix or a Vulkan-backend repair — which changes whether it belongs upstream as a backend fix or a cross-cutting one.
clm-0022
- measured-here high ●●● volatility medium · verified 2026-08-08
- The community KV-dequantisation fix is real, large, and scales monotonically with depth: one cherry-picked commit recovers +18.3% at 32k, +55.6% at 131k and **+70.3% at 204,800 — production's own context** — while leaving f16 unchanged at every depth. Warden's local model was therefore running at **59% of its achievable decode speed** at the context it actually used.
- EVIDENCE TRAIL: 32k stock/kvfix cells have run records for the visible decode arms; 131,072 and 204,800 source-JSON cells are partly represented by energy windows whose notes carry the source filename and avg_ts match, because several deep-context f16/stock q8_0 rows predate matching run records (`run: []`). The headline +70.3% production-depth q8_0 delta is anchored by `eng-0017` (stock q8_0 avg_ts 5.686) and `run-0037`/`eng-0018` (kvfix q8_0 avg_ts 9.6866), with f16 controls in `eng-0019` and `eng-0020`. METHOD — single-variable attribution. Both binaries built from the SAME upstream commit (3653e6d); the only difference is one cherry-picked patch, `ce7689f` = Nathanw1014's `2a24abc` "CUDA: dequantize KV on load in the tile FA kernel, use it for quantized decode" (6 files, +201/-45). Building the fork wholesale was DELIBERATELY avoided: it sits on upstream #25xxx while our tree is 2026-08-07, which would have confounded the patch with the base version — the attribution error clm-0017's author avoided with five same-base builds. Qwen3.5-122B-A10B UD-Q4_K_M, ROCm 7.1.0, `-fa 1`, `--load-mode none`, `--parallel 1`, 3 reps, page cache dropped between arms. | build | KV | depth | pp512 | tg128 | |---|---|---|---|---| | stock | f16 | 0 | 326.36 | 21.90 | | kvfix | f16 | 0 | 319.24 | 21.88 | | stock | f16 | 32768 | 210.21 | **18.16** | | kvfix | f16 | 32768 | 204.48 | **18.16** | | stock | q8_0 | 32768 | 200.26 | **15.29** | | kvfix | q8_0 | 32768 | 211.35 | **18.09** | **THE CONTROL IS THE POINT.** f16 decode at 32k is 18.16 on BOTH builds — identical to four significant figures. The patch touches the tile flash-attention kernel, so without an f16 arm a general FA speedup could have been misread as a quantised-KV win. It moved q8_0 by 18.3% and f16 by nothing, which is exactly what a correct KV-specific fix looks like. q8_0 prefill also gains (+5.5%); at depth 0 nothing moves, so the effect is depth-dependent as the mechanism predicts. WHAT IT CHANGES OPERATIONALLY: on stock, q8_0 KV costs 16% decode at 32k and **35% at 131k** — you pay steeply increasing speed for memory, exactly where long context is the point. **With the patch that cost disappears entirely** and slightly reverses. So quantised KV becomes free: roughly half the KV footprint at no speed penalty at any depth measured. On `cfg-0002`'s 200,000-token configuration that is ~12 GB of headroom recovered for nothing, which on a fixed-memory box buys context or a second resident model — and it is the concrete enabler for the 35B + 122B co-residency in clm-0023. ⚑ DEPTH TEST RUN — AND IT CHANGES THE CONCLUSION. My first write-up said the fix reached "parity, not superiority" and that clm-0006's "23-53% faster than f16" did not reproduce. That was an artefact of testing at only 32k. Decode at **131,072**: | build | KV | tg128 | |---|---|---| | stock | f16 | 12.03 | | stock | q8_0 | **7.81** | | kvfix | f16 | 12.04 (control — unchanged) | | kvfix | q8_0 | **12.15** | **+55.6% from the patch at 131k**, and q8_0 now sits marginally ABOVE f16 — which is clm-0006's claim, reproduced. The gain scales monotonically with depth exactly as the mechanism implies (dequantize once on load versus repeatedly during inference: the more cache you touch per token, the more you save): | depth | patch gain on q8_0 | |---|---| | 0 | none | | 32,768 | +18.3% | | 131,072 | +55.6% | | **204,800** | **+70.3%** | AT PRODUCTION DEPTH (204,800 — `cfg-0002` ran n_ctx 200,000): | build | KV | tg128 | |---|---|---| | stock | q8_0 | **5.69** | | kvfix | q8_0 | **9.69** | | stock | f16 | 9.60 | | kvfix | f16 | 9.54 (control — unchanged, -0.6% is noise) | The control holds at every depth tested: the patch moves q8_0 and leaves f16 alone. At 204,800 the patched q8_0 (9.69) now **exceeds** f16 (9.54-9.60), which is clm-0006's claim in full. Their +75% to +203% was measured at 262k, deeper than anything we ran, so our +70.3% sits just below their range on a shallower test — consistent, not conflicting. ALSO ESTABLISHED, AND WE DID NOT HAVE IT BEFORE: **f16 KV FITS at 204,800.** 9.60 t/s with memory stable at 79 GiB and no swap growth. The f16 arm was expected to be the one that might OOM against the 104 GiB ceiling; it did not. That is a real fit data point of the kind clm-0001 says we lack. 🔴 **MEASURED AT PRODUCTION'S OWN CONTEXT, AND IT IS WORSE THAN THE EXTRAPOLATION.** `cfg-0002` ran **q8_0/q8_0 KV at n_ctx 200,000** on a fork carrying only the `.ckpt` sidecar — nothing touching KV. At 204,800 that configuration produces **5.69 tok/s where 9.69 was available**: Warden's local model was running at **59% of its achievable decode speed** at the context it actually used. On a 500-token reply that is 88 seconds instead of 52. Two independent, compounding mistunings are now measured in the shipped configuration: this, and `--spec-draft-n-max 6` at ~18% below the optimum of 2-3. Neither was visible without an A/B, and neither would have shown up in any throughput number taken alone — the config simply looked like the speed of the hardware. FORK COST, NOW CONCRETE: adopting this means a SECOND divergence alongside `con-0001` (the `.ckpt` sidecar). They touch different files and both would have to be carried together. `build_commit` is a fingerprint field, so a combined tree is a configuration nobody else can reproduce — the tax is real. But a **55.6% decode recovery at 131k**, rising with depth, plus ~12 GB of headroom, on a box whose entire purpose is long context, makes this the strongest case for carrying a fork that we have. SCATTER: every cell ≤0.8% CV. Healthy code paths on both builds.
- evidence: run-0015 run-0019 run-0025 run-0027 run-0036 run-0037 eng-0012 eng-0013 eng-0014 eng-0015 eng-0017 eng-0018 eng-0019 eng-0020 cfg-0009 cfg-0010 cfg-0011 cfg-0013 cfg-0014
clm-0023
- measured-here high ●●● volatility medium · verified 2026-08-08
- Qwen3.6-35B-A3B is 2.3x the 122B's decode on identical hardware (51.01 vs 21.90 tok/s at empty context) and holds 42.45 at 32k, with prefill above 1000 tok/s. It passes the capability guard 4/4.
- QWEN3.6-35B-A3B (UD-Q4_K_XL, 21 GB), aihydra, ROCm 7.1.0, stock llama.cpp 3653e6d, `-fa 1`, f16 KV, `--parallel 1`, 3 reps. Ingested as cfg-0012 / run-0030..0032. | depth | pp512 | tg128 | |---|---|---| | 0 | **1079.38** | **51.01** | | 4096 | 965.37 | 49.80 | | 32768 | 598.68 | 42.45 | All cells ≤0.6% CV. Guard 4/4 at depth 8000. THE ACTIVE-PARAMETER PREDICTION, TESTED. Decode is bandwidth-bound, so throughput should scale roughly inversely with ACTIVE parameters: A3B (3B) against A10B (10B) predicts ~3.3x. We measure **2.3x** (51.01 vs 21.90). The direction and rough magnitude hold, and the shortfall is expected — per-token overhead does not scale down with active parameters. Good enough to keep using active-parameter count as a first-order predictor, not good enough to quote as a rule. It also retains speed at depth better than the 122B: 51.01 → 42.45 is a 17% loss to 32k, against the 122B's 21.90 → 18.16, also 17%. Proportionally identical, which suggests the depth penalty is an attention-cost property rather than a model-size one. REFLEX-TIER VERDICT: at 42-51 tok/s with sub-second prefill on 21 GB, this is a credible fast lane. Combined with clm-0022's finding that quantised KV is free on a patched build, a 35B and the 122B could plausibly be co-resident within the 105 GiB ceiling — which is the council architecture's first concrete opportunity on this hardware. ⚑ GRAMMAR CEILING: see clm-0024, which re-measures this properly and finds the ceiling does not reproduce — the constraint is context size, not the grammar builder. Kept below as the process record of how the wrong number nearly got published. NOT MEASURED HERE, AND THE FIRST ANSWER WAS WRONG. The probe reported "highest OK: 30 tools (127,415 B)" against the real 353 KB Home Assistant schemas. That is **not** the grammar ceiling. The 400 carried `{"type":"exceed_context_size_error", "n_prompt_tokens":16651, "n_ctx":16384}` — the server was started at `-c 16384` and 31 real tools need ~16.6k tokens. We measured the CONTEXT limit and nearly published it as a grammar limit, which would have understated the real ceiling badly and sent us hunting a regression that does not exist. The script's own `grammar-related error text: no` line caught it — and then printed a ceiling anyway. **A caveat nobody acts on is not a safeguard.** Fixed 2026-08-08: the probe now exits 2 and refuses to report any number when the failure is a context overflow, naming the token count it needed. Re-running at `-c 131072`. The real ceiling is somewhere above 30 tools and currently unknown. For scale: at ~4 KB per real MCP tool, probing 120 tools needs roughly 60k tokens of context — which is itself a useful fact, because production runs nowhere near that much context devoted to tool schemas. ALSO PASSED TONIGHT: tau2-bench smoke test. LiteLLM reaches our endpoint and gets a correct native tool call (`get_booking {"reference":"ABC123"}`), so the adopted agentic-benchmark path works end to end and is no longer an untested assumption.
clm-0024
- measured-here med ●●○ volatility high · verified 2026-08-08
- The tool-grammar ceiling does not reproduce on llama.cpp 3653e6d. Given adequate context, 200 real MCP tools totalling 752 KB were accepted with HTTP 200. The binding constraint is CONTEXT SIZE — tool schemas consume prompt tokens — not the GBNF grammar builder. This supersedes the ~55-59 tool figure that has shaped our MCP tool budget since 2026-07-05.
- Qwen3.5-122B-A10B UD-Q4_K_M, aihydra, ROCm 7.1.0, stock llama.cpp 3653e6d, `--jinja`, `--parallel 1`, real Home Assistant MCP schemas (78 tools, 353 KB, cycled to reach higher counts). Fresh server, `-c 131072`: | tools | request bytes | result | |---|---|---| | 60 | 226,121 | HTTP 200 | | 80 | 304,532 | HTTP 200 | | 100 | 367,660 | HTTP 200 | | 120 | 464,680 | HTTP 200 | | 160 | 593,137 | HTTP 200 | | **200** | **752,383** | **HTTP 200** | No failure was found at any count tested. WHAT THE OLD CEILING ACTUALLY WAS. Tonight's first probe reported "30 tools" — that was the server at `-c 16384` returning `exceed_context_size_error`, pure context exhaustion. A second probe at `-c 131072` reported 59/60, and a fresh server then accepted 60 AND 61. Only the third run — fresh server, ladder rather than binary search — established that nothing fails up to 200. The consistent story across all three is **context**: 200 tools is ~94k tokens of schema, which fits in 131k and cannot fit in 16k. So the mental model changes. It was "llama.cpp cannot build a grammar above ~55-59 tools by schema SIZE". It should be "tool schemas are prompt tokens, and you need context to hold them". That is a far more tractable constraint — it trades against context budget rather than being a hard wall. ⚠ CONFIDENCE IS medium, NOT high, AND VOLATILITY IS high. Three reasons to hold this loosely: · **One unexplained failure.** n=60 genuinely returned 400 during the binary search at `-c 131072`, then passed twice on fresh servers. Something is state-dependent and we have not characterised it — the same flavour of instability as clm-0019. · **We only proved acceptance, not correctness.** HTTP 200 means the grammar compiled and inference started. It does NOT mean the model selects correctly among 200 tools. Tool-selection accuracy at high tool counts is a CAPABILITY question and completely untested here. · **Build-specific by construction.** The ceiling is a property of the build; ours is one day old. It may differ on the version production runs. WHAT THIS WOULD CHANGE IF IT HOLDS (do not act before re-verifying on the production build): the ~55-59 figure has directly shaped Warden's configuration — the Home Assistant `toolFilter` cut 42 tools to 9, apple-mail was curated, and the live count sits around 51 deliberately close to the limit. If the wall is really context, that curation buys tokens rather than avoiding a cliff, and the trade can be reasoned about instead of feared. The **silent cloud fallback** risk (HTTP 400 → OpenRouter, looking healthy) is the thing that made the ceiling frightening; a context error is at least loud. NEXT: re-run on the production llama.cpp build, and — more importantly — measure whether tool-SELECTION accuracy degrades as tool count rises. Acceptance was never the interesting question; it was just the one that used to fail first.
clm-0025
- measured-here high ●●● volatility medium · verified 2026-08-08
- Four models measured on identical hardware give decode from 17.51 to 55.45 tok/s, and the bandwidth model predicts the ORDER but not the magnitude — realised efficiency ranges from 34% to 62% of the theoretical ceiling. Separately, gpt-oss-120b returns STALE ANSWERS FROM PREVIOUS REQUESTS at --parallel 1, which no other model tested does.
- All on aihydra, ROCm 7.1.0, stock llama.cpp 3653e6d, `-fa 1`, f16 KV, `--parallel 1`, `--load-mode none`, 3 reps, page cache dropped between models. | model | active | quant | pp512 @0 | tg128 @0 | tg128 @32k | |---|---|---|---|---|---| | gpt-oss-120b | ~5.1B | UD-Q4_K_XL | 450.83 | **55.45** | 41.70 | | Qwen3.6-35B-A3B | 3B | UD-Q4_K_XL | 1079.38 | 51.01 | 42.45 | | Qwen3.5-122B-A10B | 10B | UD-Q4_K_M | 326.36 | 21.90 | 18.16 | | Nemotron-3-Super-120B-A12B | 12B | UD-IQ4_XS | 225.36 | **17.51** | 17.16 | THE BANDWIDTH MODEL, TESTED PROPERLY FOR THE FIRST TIME. Decode should scale inversely with ACTIVE parameters at ~256 GB/s. Taking ~0.56 bytes/param at 4-bit: | model | active | predicted ceiling | measured | realised | |---|---|---|---|---| | Qwen3.6-35B-A3B | 3B | ~152 t/s | 51.01 | **34%** | | gpt-oss-120b | 5.1B | ~90 t/s | 55.45 | **62%** | | Qwen3.5-122B-A10B | 10B | ~46 t/s | 21.90 | **48%** | | Nemotron-120B-A12B | 12B | ~38 t/s | 17.51 | **46%** | **The ORDER is right — more active parameters, slower decode, monotonically. The MAGNITUDE is not.** Realised efficiency varies nearly two-fold, and the smallest model is the least efficient: the 35B-A3B converts only 34% of its theoretical bandwidth into tokens against gpt-oss's 62%. Per-token overhead does not shrink with active parameters, so a very sparse model spends proportionally more time on everything that is not weight streaming. This refines clm-0023, which took 2.3x on a single pair as broad support for the model. With four points the honest statement is: **active-parameter count predicts ranking reliably and throughput poorly.** Useful for choosing which model to try; useless for predicting what it will do. ⚠ gpt-oss-120b RETURNS STALE ANSWERS — AND ONLY gpt-oss DOES. Its guard failed retrieval in a way none of the others did: · On a **virgin server**, needle at 8000 tokens returns `'chartreuse'` — the first word of `chartreuse-viper-88`. Retrieval WORKS; the model simply answers partially. That is a probe-strictness issue on our side, not a model failure. · **After any prior request**, the same probe returns `'the quick brown fox'` — verbatim the answer to the COHERENCE check that ran earlier. Reproduced 6/6 across two sequences. The server hands back a previous response. · Control questions are unaffected: "2+2" → `4`, "capital of France" → `Paris`. This is at `--parallel 1`, so it is NOT clm-0019 (multi-slot) and NOT llama.cpp #25992 (concurrent leakage). ⚑ **PROMPT CACHING RULED OUT, 2026-08-08.** I predicted `cache_prompt` prefix reuse. Tested both ways with a priming request in between: | cache_prompt | priming request | needle | needle | |---|---|---|---| | true | 'the quick brown fox' | 'the quick brown fox' | 'the quick brown fox' | | **false** | 'the quick brown fox' | **'the quick brown fox'** | **'the quick brown fox'** | Disabling the prompt cache changes nothing. The model returns the previous answer regardless. Combined with the virgin-server result (first request after start answers correctly with 'chartreuse'), the pattern is: **the FIRST request on a fresh server is correct and every subsequent request returns the first one's answer.** That is slot state not being reset between requests, not a caching optimisation misfiring. Mechanism still unidentified. gpt-oss uses the harmony chat format, which llama.cpp handles through a separate code path, and that remains the most likely locus — but it is a hypothesis, not a finding. Reproducible in three lines against a fresh server, so it is cheap for anyone to confirm and worth an upstream report once characterised. ADDITIONAL, from the 2026-08-08 matrix run: · **gpt-oss REQUIRES flash attention.** Both `fa=0` arms failed outright (f16 and q8_0), where the 122B only failed the q8_0/fa=0 combination. · **q8_0 costs 56% decode at 131k** (10.70 vs f16's 24.53) — substantially worse than the 122B's 35% at the same depth, on the same stock build. If the KV-dequant patch (clm-0022) helps proportionally, gpt-oss has more to gain from it than anything else measured. · f16 at 131k flagged **3.5% CV**, above our 3% scatter threshold. Per clm-0017 that is a defective-path signature and warrants a look rather than a shrug. CONSEQUENCE: **gpt-oss-120b's throughput rows are recorded but NOT guard-cleared.** They are honest numbers for how fast it generates and say nothing about whether it generates the right thing. On this evidence it is not a candidate for any Warden role until the staleness is understood — a model that occasionally serves a previous answer is far worse than a slow one. Nemotron-3-Super passed its guard 4/4 and is the slowest of the four at 17.51 tok/s, which its 12B active parameters predict. It is also the flattest with depth (17.51 → 17.16, a 2% loss to 32k, against 17-20% for the others) — worth a second look if long-context stability ever matters more than raw speed.
clm-0026
- measured-here high ●●● volatility medium · verified 2026-08-08
- Speculation is close to worthless on Qwen3.6-35B-A3B with varied prompts — ngram-mod gives 1.11x with a 29% coefficient of variation, ngram-cache gives nothing, and MTP is unavailable because the model carries no NextN layers. The community's 121 tok/s is not a single-stream figure we can chase with configuration: their no-speculation FLOOR alone is 27% above ours, which points at the quant format, not at tuning.
- Qwen3.6-35B-A3B UD-Q4_K_XL, aihydra, ROCm 7.1.0, stock llama.cpp 3653e6d, `-fa on`, f16 KV, `--parallel 1`, ctx 16384, 3 reps x 5 varied prompts (n=15 per arm). | spec type | tok/s | stddev | ratio | |---|---|---|---| | none | **50.91** | 0.10 | 1.00x | | ngram-mod | 56.42 | **16.53** | 1.11x | | ngram-cache | 50.22 | 1.36 | 0.99x | | draft-mtp | FAILED | | | The floor agrees with llama-bench's independent 51.01 to 0.2%, so both harnesses are measuring the same thing. **draft-mtp failed cleanly and for a good reason:** `context type MTP requested but model doesn't contain MTP layers`. Unlike Qwen3.5-122B-A10B-MTP, this build of the 35B carries no NextN heads, so the 1.45x that MTP delivers on the 122B is simply unavailable here. A recorded failure against its fingerprint, not a retry candidate. **ngram-mod's 29% CV is the finding, not its 1.11x mean.** On the 122B it was 1.17x with 38% CV; here 1.11x with 29%. Draft-free speculation only pays when the output repeats text the n-gram pool has already seen, so on five deliberately dissimilar prompts some runs gain substantially and others gain nothing. The mean is not a number to plan with. This QUALIFIES clm-0016 rather than contradicting it: their 1.87x came from 500 IFEval prompts, which are short, structured and highly repetitive — the best case for the technique. On open-ended generation it largely evaporates. WHY WE WILL NOT REACH THEIR 121 tok/s BY CONFIGURATION. Decomposing the gap honestly: · **Their no-speculation floor is 64.8; ours is 50.91 — 27% apart.** A floor difference cannot be caused by speculation, parallelism or tuning. The remaining structural difference is the quant: they ran **ROCmFP4**, we ran UD-Q4_K_XL. Quant format sets bytes-per-weight and therefore the bandwidth ceiling itself, which is the one term that moves a floor. · **Their 121 is FOUR STREAMS POOLED, not single-stream.** Their own post labels the sustained figure as 4 slots across 500 prompts. Our 50.91 is one stream. Comparing the two directly overstates the gap; aggregate throughput and single-stream latency are different quantities and should never be set side by side. · Their trustworthy single-stream peak was 380 tok/s, but peak in a real agentic session is not a sustained rate either. So the honest reading: **ROCmFP4 is the lever worth chasing on this model, and it is worth roughly 27%.** Speculation is not — on varied prompts it buys ~11% with scatter three times larger than the gain. ⚠ AND WE CANNOT SIMPLY COPY THEIR PARALLEL-4 SETUP. clm-0019 measured that `--parallel 4` destroys long-context retrieval on this build (needle 0/8 against 8/8 at `--parallel 1`). Their 121 tok/s was measured at parallel 4. Either their build does not carry that defect, or their IFEval prompts were short enough never to expose it. Chasing their number by raising slot count would trade a capability we have verified for throughput we have not. NEXT, IF THE 35B MATTERS AS A REFLEX TIER: build the ROCmFPX fork and requantise. That is a third divergence to carry, on top of `con-0001` and the KV-dequant patch (clm-0022) — but unlike speculation it moves the floor, and the floor is what a reflex tier is for. ⚑ **CLM-0010 IS UNRECONCILED WITH THIS.** That report claims draft-free ngram alone takes a single stream from 71 to 216 tok/s — a 3x ratio, against our measured 1.11x on the same technique here. Neither their ROCmFPX-fork ROCmFP4 finetune nor their prompt regime (repetitive code-edit content, ngram-mod's best case) matches ours, and either could account for most of the gap — but it has not been tested, so the discrepancy stands rather than being averaged away.
clm-0027
- measured-here high ●●● volatility low · verified 2026-08-08
- Prefix cache reuse is worth 9.8x on this box — an 8,000-token prefix costs 25.99 s cold and 2.65 s warm — and it is strictly PREFIX-ANCHORED: prepending three characters to an otherwise identical prompt returns it to full cold cost (26.26 s), zero reuse despite 99.9% identical content. Separately, both f16 and q8_0 KV load at every context up to 204,800, and the f16/q8_0 footprint difference is only 2 GiB — far below the ~12 GiB cfg-0002 assumed.
- PROVENANCE CAVEAT: the cache-probe wall times and fit table below are claim-recorded measurements from the operator's `capability-probe.sh` run, not yet split into separate run/probe/config content records. Until those source artifacts are ingested, cite this claim as the record of the measurement rather than treating the Ref chip as a raw-run pointer. Qwen3.5-122B-A10B UD-Q4_K_M, aihydra, ROCm 7.1.0, stock llama.cpp 3653e6d, `-fa on`, f16 KV, `-c 32768`, `--parallel 1`, `--load-mode none`. ## Cache behaviour — the ladder nobody publishes | probe | wall | prefill | meaning | |---|---|---|---| | cold | **25.99 s** | 317.5 t/s | first sight of this prefix | | warm | **2.65 s** | 269.2 t/s | same prefix, different question | | diverged | **26.26 s** | 315.9 t/s | three characters prepended | **9.8x from prefix reuse.** That is the whole warm-lane thesis, measured on this hardware rather than inferred: the difference between a turn that feels instant and one that does not is almost entirely whether the prefix was seen before. **And reuse is prefix-anchored, not similarity-based.** The `diverged` probe is 99.9% identical to `cold` — the same 8,000 tokens, with `ZZZ ` prepended. It gets **no reuse whatsoever** and costs 101% of cold. A changed head invalidates everything after it. That is the measured justification for design decisions already taken on faith: · why the prefix-relocation fix (`cache-maxing-prefix-fix`, upstreamed as openclaw #98267) mattered — moving two sections below the cache boundary took shared tokens from 1.46K to 15.5K, and this shows what each shared token is worth; · why heartbeats and crons that carry their own preamble **cannot** share a warm lane with iMessage traffic, and why slot pinning was the right answer rather than a bigger cache; · why anything that mutates the head of a prompt — a timestamp, a rotating greeting, a changing tool list — is far more expensive than its size suggests. GAP: this build does not populate `prompt_n_cached` in `timings`, so the ratio is measured by wall clock rather than reported token counts. The effect is large enough (9.8x) that this does not threaten the conclusion, but a cached-token figure would let us measure PARTIAL reuse rather than just its presence or absence. ## Fit — both KV arms load everywhere | KV | ctx | loaded | load_s | peak GiB | |---|---|---|---|---| | f16 | 32,768 | yes | 46 | 75 | | f16 | 131,072 | yes | 78 | 77 | | f16 | 204,800 | yes | 48 | **79** | | q8_0 | 32,768 | yes | 95 | 75 | | q8_0 | 131,072 | yes | 62 | 76 | | q8_0 | 204,800 | yes | 47 | **77** | **No fit cliff anywhere.** f16 KV reaches 204,800 with 79 GiB peak against the 104 GiB GTT ceiling — 25 GiB spare. The arms are therefore comparable at the top length, which protocol §9 warns is not to be assumed. ⚑ **THE f16/q8_0 DIFFERENCE IS 2 GiB, NOT 12.** `cfg-0002` records `kv_gb: 12.4` for q8_0 at 200,000 context, implying ~25 GiB for f16. The measurement says otherwise: 79 vs 77 GiB peak. The likely explanation is architecture — this is a hybrid model, so only a minority of layers hold a full attention cache, exactly as clm-0015 found for Ling-3.0-flash (7 of 42 layers). **If so, cfg-0002's memory record is a substantial overestimate and the fit calculus behind several decisions was too conservative.** It also means quantised KV saves far less memory than assumed — which, combined with clm-0022 showing it costs nothing in speed once patched, makes the whole q8_0-vs-f16 question much less consequential than it looked. Both figures come from `capability-probe.sh`, which only produced them after being given `--load-mode none`; before that both arms timed out at 673 s and were recorded as non-loading. A probe that cannot load the model reports a fit cliff that does not exist.
clm-0028
- measured-here high ●●● volatility low · verified 2026-08-08
- Across four models on gfx1151, flash attention is worth 2.5x to 4.8x DECODE at 131k and its absence is catastrophic — a 35B loses 88% of its decode speed from empty context to 131k without it, against 44% with it. Separately, quantised KV **requires** flash attention: the q8_0 + fa=0 cell failed on 4 of 4 models. Neither fact is visible at shallow depth, where flash attention is irrelevant.
- aihydra, ROCm 7.1.0, stock llama.cpp 3653e6d, `--parallel 1`, `--load-mode none`, 3 reps, page cache dropped between cells. Depths 0 / 32,768 / 131,072. Ingested as run-0050..run-0097. ## Flash attention: decode at 131,072 (f16 KV) | model | fa=1 | fa=0 | gain | |---|---|---|---| | Qwen3.6-35B-A3B | **28.46** | 5.91 | **4.8x** | | Nemotron-3-Super-120B-A12B | **16.21** | 6.41 | **2.5x** | | Qwen3.5-122B-A10B (at 32k) | 18.17 | 9.87 | 1.84x | | gpt-oss-120b | 24.53 | **failed to run at all** | n/a | **The absence of flash attention is not a slowdown, it is a collapse.** Qwen3.6-35B-A3B goes 50.88 → 5.91 tok/s from empty context to 131k without it, an 88% loss. With it, 51.12 → 28.46, a 44% loss. Nemotron: 17.46 → 6.41 without (−63%) against 17.51 → 16.21 with (−7%). **And it is invisible at shallow depth.** At depth 0 every model is within 1% either way (35B: 51.12 vs 50.88; Nemotron: 17.51 vs 17.46). A benchmark that only measures empty context would conclude flash attention does not matter. It is the single most important setting on this hardware and only depth reveals it. ## Quantised KV requires flash attention — 4 of 4 models Every `q8_0 + fa=0` cell failed outright: Qwen3.5-122B, Qwen3.6-35B, Nemotron, gpt-oss. llama-bench exits with "failed to create context" before any inference. This is not a performance finding but a hard constraint, and it means the two levers are not independent — you cannot sweep them as a clean 2x2. gpt-oss is stricter still: **both** its fa=0 arms failed, f16 as well as q8_0. It cannot run without flash attention at all. ## q8_0 penalty on a STOCK build varies by model Decode at 131k, fa=1, q8_0 against f16: | model | f16 | q8_0 | penalty | |---|---|---|---| | gpt-oss-120b | 24.53 | 10.70 | **−56%** | | Qwen3.6-35B-A3B | 28.46 | 18.48 | −35% | | Qwen3.5-122B-A10B | 12.03 | 7.81 | −35% | | Nemotron-3-Super | 16.21 | 12.58 | −22% | All four pay a penalty, spanning 22% to 56%. This is the defect clm-0006 describes and clm-0022 measured a fix for (+70.3% recovery at production depth on the 122B). **The patch has not been tested on the other three**, but gpt-oss's −56% suggests it has the most to gain of anything measured. ## Guards Qwen3.6-35B-A3B and Nemotron-3-Super both pass 4/4 — coherence, tool call with real arguments, needle recovered at 8,000 tokens. gpt-oss fails retrieval with the stale-answer defect (clm-0025) and its rows are recorded but not capability-cleared. ## Practical **Leave `-fa on` everywhere on this hardware, unconditionally.** It wins or ties at every depth on every model tested, it is the difference between a 44% and an 88% decode loss at 131k, and three of four models cannot use quantised KV without it.
clm-0029
- measured-here high ●●● volatility high · verified 2026-08-08
- DIAGNOSED: `llama-perplexity` produces garbage on this build — PPL 532 on deliberately repetitive English that should score 2-5, 3,654 on wikitext for a model that generates coherently at 51 tok/s and passes tool-calling guards. Not flash attention, not the corpus, not context size. **The KL-divergence route to measuring KV-quant QUALITY is therefore blocked**, and clm-0022's open half needs a different instrument.
- Qwen3.5-122B-A10B UD-Q4_K_M, aihydra, ROCm 7.1.0, stock llama.cpp 3653e6d, `-fa 1`, `-c 32768`, corpus `wiki.test.raw` sha256 `430983e5…` (bench/CORPUS.md), q8_0 arm against an f16 `--kl-divergence-base`. **Core finding, by elimination across four controls (model, corpus, context, flash attention):** the defect is in `llama-perplexity` itself on this build (3653e6d) with these models — not the corpus (a 4-sentence repeating probe that should score 2-5 scored 532), not flash attention (disabling it moved 17,180 to 3,654, still absurd), not context size (4096 gave 56.2, no better than 32768), and not the model (the same 35B generates coherently and passes its tool-calling guard 4/4). Plausibly MoE or hybrid-architecture handling in the KL path, or the quant format — not established. **Consequence:** `kv-quality.sh`, built on `llama-perplexity --kl-divergence`, is unusable on this build, so the KV-quant QUALITY question could not be answered this way. It was answered instead via τ²-bench task completion (`clm-0035`, `clm-0038`). `clm-0022` covers what quantised KV costs in SPEED, which this instrument failure never touched. Full diagnostic method (the four controls, the two reasons the original KL numbers were implausible, and the alternatives considered) is written up as a standing methodology rule in `docs/methodology-lessons.md` §2 — read that before repeating this measurement.
clm-0030 superseded
- measured-here low ●○○ volatility medium · verified 2026-08-09
- SUPERSEDED: the Pass^1 = 1.000 reported by this run came from a 3-task subsample biased toward the domain's easiest tasks — the other two of the original five never terminated and were excluded as infrastructure errors. The sustained score across a realistic sample is clm-0037's 0.545 (n=22), which is the number to cite for this model on tau2 airline. This run was nonetheless the project's first genuine capability measurement rather than a throughput number, and it proved the harness and scoring path work.
- τ²-bench v1.0.1, airline domain (14 tools), aihydra, ROCm 7.1.0, stock llama.cpp 3653e6d, Qwen3.5-122B-A10B UD-Q4_K_M, `-fa on`, f16 KV, `-c 32768`, `--parallel 1`. ## Pins — changing ANY of these re-baselines the series (protocol §8) | pin | value | |---|---| | agent LLM | the local 122B | | user simulator | **the same local 122B** | | judge | **none — deterministic scoring only** | | domain | airline, 14 tools | | tasks requested | 5 | No external judge was available offline, so scoring rests on tau2's deterministic components: database state-hash comparison and tool-call trajectory match. That is the more objective half of tau2's scoring, but it is not the whole of it. ## Result | metric | value | |---|---| | **Pass^1** | **1.000** | | Read actions | 5/5 (100%) | | Write actions | **none exercised** | | DB match | ✓3 / ✗0 (100%) | | Termination | 3 normal, all user-initiated | | Infra errors | **2** | | task | reward | turns | duration | |---|---|---|---| | 0 | 1.0 | 16 | 31 min | | 1 | 1.0 | 22 | 36 min | | 2 | 1.0 | 21 | 44 min | | 3 | — | 0 | infra error | | 4 | — | 0 | infra error | The model held 16-22 turn conversations with correct tool use throughout and reached the right database state every time. Against our own stated failure mode — tool selection — that is the first direct evidence in either direction, and it is positive. ## What this does NOT establish · **Three tasks is a small sample.** Pass^1 = 1.000 over n=3 is consistent with a true pass rate anywhere from roughly 0.3 upward. It is a floor, not a score. · **Read paths only.** Write actions were never exercised — every completed task was a lookup. Mutation is where an agent can do damage, and it is untested. · **The two missing tasks are now explained, and it matters (clm-0032).** They were `infrastructure_error` in this run because tau2 defaults to THREE concurrent simulations against our deliberately single-slot server. But re-run serially they do not fail — they **never terminate**, running 4h15m and 2h10m without producing a result. They are precisely the two tasks whose user is scripted to escalate indefinitely, and on those the agent argues instead of emitting `###TRANSFER###`. **So Pass^1 = 1.000 is computed over exactly the subset the model handles well**, and the two excluded tasks are the ones testing the hardest behaviour. The headline is real but it is not the whole picture. · **Self-play.** The user simulator is the SAME model as the agent. That is tau2's default when no second model is available, but a model conversing with itself may be an easier interlocutor than a different one would be. · **Airline is 14 tools**, far below the tool counts Warden actually runs (~51) and far below where grammar or context limits bite (clm-0024). ## Cost, which matters for planning **31-44 minutes per task** on this model, and the earlier clipped run measured 110 minutes per task while sharing the box. A 5-task domain is a 3-4 hour job; the full airline suite would be far longer. Any future capability series needs that budgeted honestly rather than guessed — my first attempt at this run was killed by a 2-hour timeout I set without measuring a single task first.
clm-0031
- community low ●○○ volatility high · verified 2026-08-09
- A community DeepSeek-V4-Flash report independently confirms our mmap/GTT double-residency finding, demonstrates a 120 GiB GTT ceiling in production use, and — most consequentially — reports Vulkan BEATING ROCm on 3 of 4 cells including 55% faster decode at depth. It also documents a GPU ring-timeout failure mode we have never hit and were not watching for.
- Source: a single operator (Reddit, u/Neuromacmd) on an ASUS PX13 — Strix Halo / Radeon 8060S, gfx1151, 128 GB unified, Fedora 44, kernel 7.1.5, Mesa 26.1.5. Model `unsloth/DeepSeek-V4-Flash-0731-GGUF` UD-IQ3_XXS, 97.05 GiB. Same silicon as aihydra, so it transfers more directly than most community reports. ## 1. INDEPENDENT CONFIRMATION of the double-residency hazard Their words: *"`-mmp 0` is not optional at this size on my box: with mmap the page cache gets counted against GTT on top of the weights and it never loads."* That is precisely the defect that bit us four separate times before being fixed properly (protocol §0, `LLAMA_ARG_LOAD_MODE=none`). A second operator hitting it independently, on the same silicon, confirms it is a property of unified memory rather than anything specific to our setup — and that the fix belongs at the environment level, not per script. ## 2. A 120 GiB GTT ceiling works in production They run `ttm.pages_limit=31457280` (**120 GiB**) with a **512 MB** VRAM carve-out, and load a 97 GiB model to a 99 GiB footprint. We run 104 GiB GTT with a 1 GiB carve-out. So our 104 GiB was conservative — deliberately, as a guardrail against the aibeast unkillable-OOM class (clm-0022 notes the reasoning). This is evidence the headroom above it is usable if a model ever needs it. It does not change the guardrail argument; it removes the worry that 120 GiB is unstable. ## 3. ⚑ VULKAN AHEAD OF ROCm ON 3 OF 4 — which contradicts our working assumption | test | Vulkan | ROCm | |---|---|---| | shallow pp2048 | **124.16** | 114.37 | | shallow tg64 | **18.15** | 13.27 | | d24576 pp2048 | 63.67 | **74.31** | | d24576 tg64 | **14.36** | 9.25 | **Decode at depth: Vulkan 14.36 against ROCm 9.25 — 55% faster.** We have been treating HIP as the default on the strength of two garbage-output reports (clm-0014) and clm-0021's finding that the FA prefill cliff is Vulkan-only. This says the picture is more mixed. ⚠ BUT IT IS NOT A CLEAN COMPARISON, and the author does not claim it is. The Vulkan side is revision `4a1fb6c` (build 867) carrying **three of their own unsubmitted patches**; the ROCm side is `cd0fa60` (build 824). **Different commits, different patch sets.** That is exactly the confound clm-0017's author avoided with five same-base builds and that we avoided by cherry-picking a single commit for clm-0022. Their Vulkan advantage may be the backend, their patches, or the 43-build gap. ⚠ CONFIDENCE LOWERED TO LOW ON THIS BASIS: a 43-build gap between the two revisions tested, compounded by three of the author's own unsubmitted patches present on the Vulkan side only, is a structural confound severe enough that the Vulkan-ahead finding cannot be attributed to the backend until it is reproduced here on matched commits. ACTION: this is worth testing ourselves and we are unusually well placed — we have both backends, a controlled harness, and the discipline to build both from one base. **A clean Vulkan-vs-HIP comparison on identical commits is now the highest-value untested backend question**, and it was already on the coverage matrix as not-started. ## 4. A failure mode we were not watching for *"Slow enough at depth that the 10s compute-ring watchdog fires and the driver kills the context."* With dmesg signature: amdgpu: ring comp_1.1.1 timeout, signaled seq=7906, emitted seq=7907 amdgpu: Starting comp_1.1.1 ring reset amdgpu: [drm] device wedged, but no recovery needed **We have zero occurrences of this on aihydra** (checked). But it is a plausible hazard for exactly the work we do — a single op slow enough at depth trips a 10-second GPU watchdog and the context dies. It would present as an inexplicable mid-run failure, and we would not currently know to look for it. Added to what our watchers check. Root cause in their case (traced by khimaros, llama.cpp #25664): with no Vulkan `LIGHTNING_INDEXER`, `resolve_fused_ops` disables `fused_lid`, DeepSeek-V4 takes a fallback path with a dim0/dim2 permute plus `CONT` that misses the tiled-transpose path and lands in a generic strided copy with consecutive elements 4 MB apart. ## 5. Worth noting on provenance The author states plainly: *"Claude wrote the shader and the ggml-vulkan.cpp integration under my direction. I ran the testing and the benchmarks. I can't defend the shader line by line to a reviewer or maintain it, which is why there's no PR."* That is an honest disclosure and the right call — and it mirrors our own constraint that upstream contributions must be defensible by the person submitting them. It also means the patches are unlikely to land upstream, so any advantage they confer is not something to plan around.
clm-0032
- measured-here high ●●● volatility low · verified 2026-08-09
- The tau2 "runaway" tasks are a failure to escalate — and clm-0033 later established the failure is CAUSED BY THINKING, which this claim wrongly ruled out. Tasks that terminate do so via `###TRANSFER###`; the two that never terminate are exactly the two whose user is scripted to escalate indefinitely, and on those the agent argues rather than handing off. Reproduced across three independent runs. A model that will not hand off is an operational risk, not a benchmark artefact.
- τ²-bench airline, Qwen3.5-122B-A10B UD-Q4_K_M on aihydra, thinking on, `--parallel 1`. Observed across three runs (2026-08-08 x2, 2026-08-09 x1) with identical task selection. ## The split is perfectly clean | task | user's exit condition | outcome | terminated by | |---|---|---|---| | 0 | *"You don't want to cancel if you don't get a refund"* | reward 1.0, 16 msgs | **`###TRANSFER###`** | | 1 | *"You don't want to go ahead with the cancellation if..."* | reward 1.0, 22 msgs | **`###TRANSFER###`** | | 2 | topic change, accepts the policy | reward 1.0, 21 msgs | user: *"No, that's all for now"* | | 3 | **none** — *"ask to be transferred to a supervisor"* | **never terminates** | — | | 4 | **none** — *"insist... after you insisted 5 times"* | **never terminates** | — | Every task with a scripted exit condition converges and scores 1.0. Every task where the user is told to escalate indefinitely runs forever: 4h15m and 2h10m in two unbounded runs, still generating, never emitting a termination signal. ## Why this is a capability finding and not a harness artefact `###TRANSFER###` is τ²-bench's escalation signal, and the agent uses it correctly on tasks 0 and 1 — it recognises a request it cannot fulfil within policy and hands off. On task 3 the user **explicitly asks to be transferred to a supervisor**, which is the same action the agent already demonstrated it can take, and it does not take it. It keeps restating the policy instead. So the model can escalate, and chooses not to under sustained pressure. That is a behaviour, not a limitation. ⚠ **SELF-PLAY AMPLIFIES IT.** Both roles are the same model (no second model was available offline). So this is one model refusing to yield to itself: the agent will not transfer, the simulated user will not stop asking, and neither side has a mechanism to break the loop. A human user would eventually give up or a different simulator might. **How much of the non-termination is the agent versus the self-play pairing is not established**, and it is the obvious next control — run the user simulator on the 35B and see whether tasks 3 and 4 terminate. ## What it means for Warden This is the first tau2 result that maps directly onto production risk. Warden holds multi-turn conversations and has no supervisor to escalate to — but it does have the option to say "I cannot do this" and stop. A model that instead argues indefinitely under pressure would, in Warden's case, burn context and time rather than deferring. It is also a reminder of what Pass^1 = 1.000 (clm-0030) does and does not mean. Three of three COMPLETED tasks scored perfectly. The two that did not complete are excluded from that average — and they are the two testing the hardest behaviour. **The headline number is computed over exactly the subset the model handles well.** ## Corrections this forces to my own earlier reporting · I framed the long-running tasks as "thinking failing to converge" and treated the generation volume as the cause. It is not — the cause is conversational deadlock, and thinking merely makes each turn of that deadlock expensive. · I predicted the thinking-OFF arm would hang on the SAME two tasks, on the reasoning that the deadlock was structural. **That prediction was WRONG** — see clm-0033. Thinking-off completes both in about a minute each. The deadlock is caused by thinking, not by the task design, which makes this a much stronger finding than I expected and means the section above overstates the "structural" reading. · `--timeout` in tau2 is accepted and silently ignored (verified: flag present in `/proc/<pid>/cmdline`, simulation ran 2h10m past an 1800s bound). `--max-steps` (default **200**) is the lever that actually binds.
clm-0033
- measured-here high ●●● volatility medium · verified 2026-08-09
- On tau2-bench airline, thinking is a NET NEGATIVE for this model: thinking OFF solves 5/5 tasks at reward 1.0 in ~10 minutes, while thinking ON solves 3/5 and deadlocks indefinitely on the other two (4h15m and 2h10m in unbounded runs). Same reward on every task both complete. Thinking bought no measurable accuracy and caused total failure on 40% of the set.
- τ²-bench airline, Qwen3.5-122B-A10B UD-Q4_K_M, aihydra, ROCm 7.1.0, stock llama.cpp 3653e6d, `-fa on`, f16 KV, `-c 32768`, `--parallel 1`, `--max-concurrency 1`, `--max-steps 40`. Thinking gated SERVER-side with `-rea on|off`, verified per arm (`reasoning_content=1475 chars` with it on, absent with it off). **Both arms otherwise byte-identical** — same model, same flags, same tasks, same pins. | | thinking ON | thinking OFF | |---|---|---| | tasks completed | **3 / 5** | **5 / 5** | | mean reward (completed) | 1.00 | 1.00 | | task 3 | never terminates | 14 turns, **1 min**, reward 1.0 | | task 4 | never terminates | 12 turns, **1 min**, reward 1.0 | | arm wall time | 90 min (hit the backstop, still unfinished) | **~10 min** | Per-task, where both completed: | task | ON turns / time | OFF turns / time | |---|---|---| | 0 | 10 / 4 min | 20 / 2 min | | 1 | 24 / 5 min | 24 / 2 min | | 2 | 20 / 6 min | 29 / 4 min | **Thinking produces fewer, more expensive turns; without it the model takes more, cheaper ones and finishes sooner.** Reward is identical on every task both arms complete. ⚑ **I PREDICTED THIS WOULD GO THE OTHER WAY.** In clm-0032 I recorded the expectation that thinking-OFF would deadlock on the SAME two tasks, because I had concluded the deadlock was structural — a user simulator scripted to escalate forever against an agent that would not hand off. That was wrong. **Thinking-off terminates both tasks in about a minute each.** The deadlock is not structural; it is caused by thinking. ## Why this is the more troubling result Tasks 3 and 4 are the adversarial-pressure tests: the user insists indefinitely, asks to be transferred to a supervisor, demands compensation they are not owed. With thinking on, the agent argues in circles and never emits `###TRANSFER###` — despite using that same escalation signal correctly on tasks 0 and 1. With thinking off, it resolves them in 12-14 turns. A plausible mechanism, NOT established: extended reasoning gives the model more room to rationalise continuing to engage — to construct another framing, another policy restatement, another attempt to satisfy the user — where a shorter path reaches "I cannot do this" and stops. That is speculation. What is measured is that the behaviour flips cleanly with the lever. ## ⚑ SUPERSEDED IN SCOPE by clm-0035 — read that first This claim is correct for the 122B and **wrong as a generalisation**, which is how I went on to use it. `clm-0035` ran the same arms across all four models: thinking IMPROVES the other three (Nemotron 0.60→1.00, gpt-oss 0.00→0.60, 35B 0.40→1.00) and degrades only this one. The Warden recommendation below rests on that mistaken generalisation and should be read with clm-0035 alongside it. ## What it means for Warden Warden runs `thinking=high` on the synthesis-anticipatory cron (thinking-experiment, 2026-07-22), enabled as an A/B against a 168.9s no-think baseline and never revisited. **This is evidence that experiment should be re-run and probably reverted** — at least for anything agentic or multi-turn. Note the scope though: that cron is a single-shot synthesis pass, not an adversarial multi-turn conversation, so this result does not transfer to it automatically. It does make the assumption worth testing rather than carrying. ## Scope — one domain, five tasks · **Airline only**, 14 tools. Retail and telecom untested. · **Five tasks**, and the effect is a 2-task swing. Small, though the failures are total rather than marginal, and reproduced across four runs. · **Self-play** — user simulator is the same model. A different simulator might not escalate as relentlessly, which would shrink the gap. · **Deterministic scoring only** — no LLM judge was available offline, so reward reflects DB-state match and tool-call trajectory, not conversational quality. It is possible the thinking arm produces better *prose* while failing the task; nothing here would see it. · **Read paths only.** No write actions were exercised in any completed task.
clm-0034
- community med ●●○ volatility medium · verified 2026-08-09
- domdoss/Warden (unrelated project, coincidental name) implements the multi-model architecture we have been designing toward — a small local orchestrator routing to named specialists with per-agent model selection. Its most valuable idea for us is the SUPERVISION LOOP, which is the exact mechanism missing from the deadlock we measured in clm-0032/0033, and its orchestrator/executor split reframes what the 35B's 0.40 tau2 score actually rules out.
- Source: `github.com/domdoss/Warden` — TypeScript/Node 20, SQLite, MIT, 83 stars, 141 commits, actively maintained. A personal desktop assistant with shell/browser/desktop access and no sandbox. Read for architecture, not adopted. ## The architecture, briefly A **12B Gemma 4 local orchestrator** "reads your message, works out what you actually want, hands a clean brief to the right specialist, and then babysits that specialist until the job is done." Specialists are named by role — Atlas (shell/browser/web), Hephaestus (code), Iris (mail/calendar), Dexter (schedules, never executes), Artemis (audit), **Council (three independent deliberation seats)**, Sentry (security monitor). Each agent's model is independently configurable, local Ollama or cloud. ## ⚑ THE SUPERVISION LOOP — the thing we are missing *"The orchestrator supervises them on a fixed 30-second monitor tick. Failed jobs auto-retry with corrected briefs; cascading failures surface to the user after two identical failures."* That is precisely the mechanism absent from the failure we measured. `clm-0032`/`clm-0033` found the 122B deadlocks on adversarial-pressure tasks with thinking on — arguing in circles for 4h15m, never emitting `###TRANSFER###`, no internal mechanism to stop. **A model cannot reliably supervise its own termination.** An external tick that notices "this has been running 30 seconds past reasonable and has not converged" solves structurally what we tried to solve with a `--max-steps` bound. Two failures this week would have been caught by the same pattern: the deadlocking tau2 tasks, and the llama-server that hung in `futex_` while my wait loop watched forever. Both are the same shape — **no external observer with the authority to intervene**. ## ⚑ ORCHESTRATOR ≠ EXECUTOR, WHICH CHANGES WHAT 0.40 MEANS Their orchestrator is a **12B** model, smaller than anything we have benchmarked as a candidate. It works because its job is triage, brief-writing and supervision — NOT doing the task. The capability bar differs by role. We measured Qwen3.6-35B-A3B at **0.40 mean reward** on tau2 airline and I called that discouraging for a reflex tier (clm-0025, coverage.md). That judgement conflated two roles. tau2 measures *task execution* under adversarial pressure. It says almost nothing about whether a model can read a request, pick the right specialist and write a clean brief — which may well be within a 35B, and is demonstrably within their 12B. **So the 35B is not ruled out as an orchestrator by our data. It is ruled out as an executor of hard agentic tasks.** Those need separate measurement, and we have no benchmark for the routing role at all. ## Other transferable pieces · **Delegation discipline** — *"Never tell Atlas how to use the internet — no URLs, no search queries, no 'go to X then click Y.'"* Brief the WHAT, never the HOW. This matches what the delegated-build work already found independently (working brief template, 6-item stumble taxonomy), which is mild evidence the principle is real. · **Persistent agent runner** — one warm child process holding MCP connections across turns, IPC rather than cold starts. A different solution to the same problem our warm-lane work attacks from the cache side. · **Context compaction to ~1K after each turn** — the opposite strategy to ours. We keep a large stable prefix and exploit reuse (clm-0027: 9.8x, strictly prefix-anchored); they keep context tiny so there is little to re-process. Both are defensible; ours depends on the prefix never changing at the head, theirs does not. Worth knowing there is a second viable answer. · **File-based state** — MEMORY.md / TODO.md / HEARTBEAT.md loaded each turn, with a local model distilling the last ~30 messages into durable facts afterwards. Convergent with our own design, arrived at independently. ## What NOT to take No sandbox, no containers, full user-account access, with an explicit safety modal admitting it. That is a deliberate trade for a single-user desktop tool. Warden already runs with narrower actuation and a delivery gate, and the measured behaviours here — a model that will not stop arguing, a server that will not die — are arguments for keeping it that way rather than loosening it. NOT EVALUATED: whether any of it works well. This is a read of the design, not of the results. 83 stars is early-stage, and there are no published benchmarks to compare against ours.
clm-0035 retracted
- measured-here low ●○○ volatility medium · verified 2026-08-09
- RETRACTED: the per-model reward rankings and quantised-KV cost reported by this run do not hold — they came from 5-task tau2 arms whose ~0.40 run-to-run noise and 40-step cap bias were only characterised afterward (clm-0036), so the reward numbers below are not usable. The corrected 122B score is clm-0037's; the corrected cross-model comparison is clm-0039's. What survives is categorical, not scored: the 122B fails to terminate some tasks with thinking on, and gpt-oss fails the domain outright.
- ⛔ **RETRACTED IN LARGE PART — see clm-0036 before reading any number below.** The identical stock q8_0 configuration was subsequently run three times and scored 0.60, 1.00, 0.60. The run-to-run noise of this 5-task harness is ~0.40, which is the same size as nearly every gap reported here. Specifically: · **The KV quality finding (f16 1.00 vs q8_0 0.60) is retracted.** f16 is 2 for 2 and stock q8_0 is 1 for 3 — suggestive, not significant. Underpowered, not disproven. · **The Nemotron thinking result (0.60 → 1.00) is not established.** It is exactly one noise-width, and was the headline conclusion here. · **The per-model rankings are not established** for the same reason. What survives is the categorical outcomes rather than the scores: the 122B's failure to TERMINATE with thinking on (2 of 5, reproduced across four runs in clm-0033), and gpt-oss's 0.00 total failure alongside its independently reproduced staleness bug. Confidence dropped medium → low. The text below is left unedited as written so the reasoning that produced it stays inspectable. --- τ²-bench airline, 5 tasks, `--max-concurrency 1`, `--max-steps 40`, 60-min backstop per arm. aihydra, ROCm 7.1.0, stock llama.cpp 3653e6d, `-fa on`, f16 KV unless stated, `--parallel 1`. Thinking gated server-side with `-rea on|off`, and every ON arm probed for `reasoning_content` before running — all four models produce it, so none was skipped. ## The thinking matrix | model | thinking OFF | thinking ON | wall OFF / ON | |---|---|---|---| | Qwen3.5-122B-A10B | **5/5 @ 1.00** | 3/5 @ 1.00 (2 deadlock) | 10 / 90+ min | | Nemotron-3-Super-120B-A12B | 5/5 @ 0.60 | **5/5 @ 1.00** | 15 / 45 min | | Qwen3.6-35B-A3B | 5/5 @ 0.40 | 1/1 @ 1.00 (backstop) | 4 / 60+ min | | gpt-oss-120b | 5/5 @ **0.00** | 5/5 @ 0.60 | 15 / 10 min | ⚑ **THIS CORRECTS THE FRAMING OF clm-0033.** That claim measured the 122B and concluded thinking is a net negative — correctly, for that model. But I treated it as the likely general case and flagged Warden's `thinking=high` cron for reversion on that basis. **The opposite is true for the other three.** Nemotron goes from 0.60 to a perfect 1.00 with thinking on, completing every task. gpt-oss is unusable without it (0.00) and merely poor with it (0.60). So the lever is not "thinking good" or "thinking bad" — it is **per-model, and must be measured per-model**. Any placement decision that sets it globally is wrong for three models out of four whichever way it is set. **Nemotron-3-Super with thinking on is the standout of the whole set**: 5/5 completed at reward 1.00, the only configuration to match the 122B's best while also terminating on every task. It is the slowest model measured (17.51 tok/s decode, clm-0025) — which makes it a genuine speed-versus-reliability trade rather than a dominated option. ⚠ The 35B's ON arm completed only 1 task before the 60-minute backstop, so its 1.00 is over n=1 and means very little. It does show the same deadlock tendency as the 122B. ## q8_0 KV costs task success — clm-0022's open half, closed Same model, same tasks, thinking off, ONLY the cache type differing: | KV | completed | mean reward | failures | |---|---|---|---| | f16 | 5/5 | **1.00** | — | | q8_0 | 5/5 | **0.60** | tasks 2 and 4 hit `max_steps` at 42 and 41 turns, scored 0.0 | **Quantised KV is not free.** clm-0022 established it costs nothing in SPEED once the dequant patch is applied (+70.3% recovery at production depth). This is the other half: it costs **40% of task success** in this sample, and the failure mode is specific and legible — the two failed tasks did not produce wrong answers, they **failed to terminate**, running to the step limit exactly as the deadlock cases do. That is a more useful result than a perplexity delta would have been. The KL route was dead on this build (clm-0029, PPL 532 on repetitive English), and τ² turned out to measure the thing that actually matters — whether the model finishes the job. ## Scope, stated plainly · **One domain, five tasks.** Every cell is n=5. A 0.40 swing is two tasks. · **Self-play** — agent and user simulator are the same model throughout. · **Deterministic scoring only** — DB-state match and tool-call trajectory, no LLM judge. · **Read paths only**; no write actions exercised. · The q8_0 result is on the STOCK build. Whether the dequant patch (clm-0022) also restores quality, or only speed, is **untested and is now the obvious next run**.
clm-0036
- measured-here high ●●● volatility low · verified 2026-08-09
- τ²-bench at 5 tasks cannot resolve the differences drawn from it in this project's capability matrix. The IDENTICAL stock q8_0 configuration scored 0.60, 1.00 and 0.60 across three independent runs — a 0.40 spread from noise alone. That is the same size as most gaps in the capability matrix, so those gaps are not established. Three tasks always pass and two are coin-flips, which makes the 5-task mean effectively two Bernoulli trials. The domain has 50 tasks available; five were used.
- **This claim retracts the KV half of clm-0035 and puts every ranking in it in doubt.** **Core finding:** repeating the identical stock q8_0 configuration three times scored 0.60, 1.00, 0.60 — the patched-vs-stock "quality fix" that prompted the repeat lay entirely within that spread, and would have been published as a confirmed result without it. The failures are not scattered: in both 0.60 runs the same two tasks failed, both at the step-limit boundary, which means the 5-task mean is really two Bernoulli trials (values 0.60/0.80/1.00 only) — and the step limit itself was a 40-turn budget chosen for wall-clock convenience, cutting tasks off right where they naturally finish rather than scoring them wrong. **What survives:** categorical completion failures, not reward differences — the 122B deadlocking with thinking ON (`clm-0033`) and gpt-oss's stale-answer bug (`clm-0025`). **What does not survive:** q8_0-vs-f16 quality, whether the dequant patch restores it, and the Nemotron thinking-on result from `clm-0035` — all fall inside this noise band. The full worked diagnosis — the run table, the Bernoulli-trial argument, the fix (the airline domain ships 50 tasks; five were used) and the methodological lesson about confirming results that look *too* clean — is written up as a standing rule in `docs/methodology-lessons.md` §1.
clm-0037
- measured-here med ●●○ volatility medium · verified 2026-08-10
- The 122B's real τ²-bench airline score is 0.545 +/-0.208, not the 1.00 reported by 5-task runs, which sampled only the easiest tasks in the domain. At n=22 it passes 12 and fails 10. Nothing was cut at the 200-step budget, and tasks run a median of 22 turns and a maximum of 36 — so the model's long tasks are long in TOKENS per turn (~10k), not in turns.
- Qwen3.5-122B-A10B UD-Q4_K_M, f16 KV, `-fa on`, `--parallel 1`, thinking off, stock llama.cpp 3653e6d asserted at launch. τ² airline, `--max-concurrency 1`, `--max-steps 200` (tau2's default). aihydra, ROCm 7.1.0. ## The number | | value | |---|---| | mean reward | **0.545** +/-0.208 (95%) | | n | **22** — the arm hit its 6h bound at 360 min, incomplete of 50 | | split | 12 pass / 10 fail | | cut at max_steps | **0** | | turns | median 22, p90 30, max 36 | ## What it corrects Every 5-task run of this configuration returned **1.00**, and I quoted that as the model's score repeatedly — including as the reference point the whole KV-quality comparison was built on. At n=22 it is **0.545**. The first five tasks are simply the easy ones, so the 5-task figure was not a noisy estimate of 1.00; it was a confident measurement of the wrong thing. This is the concrete cost of the sampling error `clm-0036` identified, and it is larger than that claim predicted. clm-0036 argued the 5-task harness could not RESOLVE differences of ~0.4; this shows it was also BIASED, because the truncated set was not a random sample of the domain. ## Two findings that only appear at a realistic step budget **Nothing was cut at max_steps (0 of 22), and the longest task ran 36 turns.** So the 200-step budget is comfortably adequate, and my old `--max-steps 40` sat just above the observed maximum — close enough that it clipped the tail while looking generous. **The slow tasks are slow in tokens, not turns.** Two tasks in this arm each took over an hour while the GPU stayed pinned at 99%. The cause is generation length: ~10,000 tokens in a single turn, which at 19.5 t/s is ~8.5 minutes for one response. A task of 30 such turns takes hours without ever looping or stalling. ⚑ I misread this at the time and called it a retry loop, on a flat `n_tokens` field whose semantics I had not checked. The prompt-eval counts (41-397 tokens per request) showed healthy KV cache reuse of an advancing conversation the whole time. Corrected before it reached the record, but it was one step from being written down. ## Scope — read this before quoting 0.545 · **n=22 is a CONTIGUOUS PREFIX, not a random sample.** The arm ran tasks in order and was cut off by the time bound, so if task order correlates with anything, this is biased — in an unknown direction. A shuffled or completed run is needed before 0.545 is a property of the model rather than of tasks 0-21. · +/-0.208 is still wide. It excludes 1.00 decisively, which is the point; it does not pin the value. · Airline only, self-play, deterministic scoring, read paths only.
- evidence: run-0100
clm-0038
- measured-here med ●●○ volatility medium · verified 2026-08-10
- Paired on identical tasks, q8_0 KV costs TURN EFFICIENCY: 228 turns against f16's 164 over the same 9 tasks, +39%, taking more turns on 6 of 9 and fewer on 1. Reward barely moves (1.000 vs 0.889, a single task) because reward is binary and coarse — turn count is the sensitive instrument and shows a consistent direction the mean hides. This also explains the wall-clock divergence: q8_0 was 76% slower on one task while decoding only 8% slower.
- Qwen3.5-122B-A10B UD-Q4_K_M, thinking off, `-fa on`, `--parallel 1`, `--max-steps 200`, stock llama.cpp 3653e6d (NO dequant patch). τ² airline. Arms differ ONLY in `-ctk/-ctv`. ## The paired table Both arms ran the same tasks in the same order, so the overlap is a true paired design rather than two independent samples. | task | f16 reward | q8_0 reward | f16 turns | q8_0 turns | Δ turns | |---|---|---|---|---|---| | 0 | 1.0 | 1.0 | 20 | 14 | **-6** | | 1 | 1.0 | 1.0 | 24 | 24 | 0 | | 2 | 1.0 | 1.0 | 29 | 40 | +11 | | 3 | 1.0 | 1.0 | 14 | 14 | 0 | | 4 | 1.0 | 1.0 | 12 | **48** | **+36** | | 5 | 1.0 | 1.0 | 15 | 22 | +7 | | 6 | 1.0 | 1.0 | 10 | 10 | 0 | | 8 | 1.0 | **0.0** | 28 | 38 | +10 | | 9 | 1.0 | 1.0 | 12 | 18 | +6 | | **total** | **1.000** | **0.889** | **164** | **228** | **+39%** | ## Why this is the first useful reading of a question I kept mismeasuring clm-0035 claimed q8_0 cost 0.40 of reward; clm-0036 retracted it once the same stock configuration produced 0.60, 1.00, 0.60 on repeat. Both were UNPAIRED comparisons of 5-task means, and 5-task means turned out to be both noisy and biased (clm-0037). Pairing removes the sampling problem entirely: same tasks, same order, one variable. And it exposes that **I was reading the wrong metric.** τ² reward is 1.0/0.0 per task, so it can only move in steps of 1/n and needs a task to flip outright before it registers anything. Turn count is continuous, moves on every task, and here shows a consistent direction — q8_0 takes more turns on 6 of 9, ties on 3, and is shorter on 1. **Turn count should be the primary instrument for KV-quality work**, with reward as the coarse confirmation. That reverses how I have been using them. ## It also explains an anomaly I had flagged but not accounted for Arm 2 spent 4.4 h on task 8 where arm 1 spent ~2.5 h, while decoding only 8% slower (18.84 vs 20.41 t/s). An 8% speed difference cannot produce a 76% runtime difference. The paired data resolves it: q8_0 took **38 turns on task 8 against f16's 28**, and 48 against 12 on task 4. The extra wall-clock is extra WORK, not slower work. ## Scope · **n=9 paired.** The reward difference is one task and means little on its own; the turn difference is the substantive signal, and even that is 9 pairs. · A sign test on 6 improvements / 1 regression / 3 ties does not reach conventional significance. The +39% aggregate is driven substantially by task 4 (12 -> 48). · **Stock build only.** Whether the dequant patch (clm-0022) removes the turn penalty along with the speed penalty is untested and is now the obvious next run. · Both arms were cut short by the 6 h bound (22 and 9 of 50), which is why the overlap is 9 rather than 50. The pairing is sound; the sample is small. · Airline only, self-play, deterministic scoring, read paths only.
- evidence: run-0100 run-0101
clm-0039
- measured-here med ●●○ volatility medium · verified 2026-08-10
- On reward the 122B and Nemotron are indistinguishable, but reward is the wrong headline: on tasks both get RIGHT, the 122B needs 19 turns and 2.0 minutes against Nemotron's 26 and 3.7 — 37% fewer loops and 85% less wall-clock to the same correct answer. Separately, the turn distributions differ so much — 122B max 36, Nemotron max 84 — that the earlier --max-steps 40 cap sat above one model's entire distribution and sliced through the other's, making the earlier cross-model ranking biased rather than merely noisy.
- τ² airline, `--max-steps 200` (tau2 default), `--max-concurrency 1`, f16 KV, `-fa on`, thinking off, stock llama.cpp 3653e6d. Both arms bound-limited at 6h. ## The two arms, and the distributions underneath them | | mean | n | pass/fail | turns median | p90 | **max** | |---|---|---|---|---|---|---| | Qwen3.5-122B-A10B | 0.545 +/-0.208 | 22 | 12/10 | 22 | 30 | **36** | | Nemotron-3-Super | 0.625 +/-0.237 | 16 | 10/6 | 28 | 60 | **84** | Nothing was cut at max_steps in either arm, so 200 is genuinely adequate for both. ## ⚑ Why the old 40-step budget was worse than "noisy" `clm-0036` established that my `--max-steps 40` was chosen for wall-clock convenience and clipped tasks at 41-42 turns. What that claim treated as a source of NOISE is, with the distributions now visible, a source of BIAS: · The 122B's longest task takes **36** turns. A 40-step cap almost never binds. · Nemotron's p90 is **60** and its longest is **84**. A 40-step cap truncates something like a third of its tasks. So the same bound was nearly free for one model and punitive for the other. Every cross-model number in `clm-0035` was produced under it. That is why Nemotron appeared to score 0.60 against the 122B's apparent 1.00 — the comparison was measuring *how many turns a model takes to reach an answer* as much as whether it reaches one. This is the third instance of the same error class, and the most damaging: the 2-hour tau2 timeout, the 90 °C thermal kill, and now this. The first two cost time. This one produced a wrong ranking that I published and reasoned from. ## What the corrected comparison says — ON REWARD **Nothing separates these two models on reward.** 0.545 +/-0.208 against 0.625 +/-0.237 — intervals that overlap across most of their range. ## ⚑ BUT REWARD IS THE WRONG HEADLINE, AND THE PROJECT ALREADY KNEW THAT The operator's framing, 2026-08-10: *"a model that takes more turns to get something correct is likely to be a greater frustration than one that is slower, but arrives at the correct outcome with less back-and-forth."* That is not a preference — it restates this project's OWN founding metric. `docs/14-model-backend-benchmark.md` (June 2026) sets the headline as **"time-to-correct- result + loops-to-done, with tool-call success as a gate"**. I drifted to raw τ² reward and spent this session treating turn count as a nuisance variable to be controlled for, when the original plan had it as a primary outcome. Restricted to tasks each model actually got RIGHT — turns spent failing are a different question — the two are not close: | on successful tasks | turns median | turns mean | minutes median | minutes mean | worst | |---|---|---|---|---|---| | Qwen3.5-122B-A10B | **19** | 18.8 | **2.0** | 2.4 | 6.0 min | | Nemotron-3-Super | 26 | 28.0 | 3.7 | 4.9 | **14.2 min** | Nemotron needs **37% more turns** and **85% more wall-clock** to reach the same correct answer, and its worst successful case takes **2.4x longer**. On the metric that describes what using the thing feels like, the 122B wins clearly. So the summary goes: clm-0035 said Nemotron was the standout (wrong — an artefact of the 40-step cap). Earlier in this claim I said they were indistinguishable (true of reward, and the wrong metric). **The 122B is materially better at getting to a correct answer with less back-and-forth.** One incidental finding from the same cut: failed tasks run LONGER than successful ones in both models — 122B 26 turns against 19, Nemotron 29 against 26. Failure is preceded by flailing, not by giving up early. That suggests turn count could serve as a live early-warning signal for a task that is going wrong, which is exactly the supervision hook clm-0034 identified as missing. ## Scope · Both arms incomplete — 22 and 16 of 50 — because the 6h bound fired. Contiguous prefixes, not random samples (see clm-0037). · The turn distributions are the robust part here. They are measured over every task in each arm and do not depend on the reward metric at all. · Airline only, self-play, deterministic scoring, read paths only. · Nemotron still runs at **IQ4_XS while the 122B runs Q4_K_M** — the quant confound from `candidates/nemotron3-super` is UNRESOLVED and applies to this comparison too. If anything it handicaps Nemotron further, which makes the equal-performance finding conservative rather than generous. · ⚠ This is a cross-model comparison, so it carries the unpinned-simulator confound documented in `clm-0043` — the user simulator was the model under test on both sides, not a fixed third party — and should be read as provisional until re-run pinned.
- evidence: run-0100 run-0102
clm-0040 superseded
- measured-here low ●○○ volatility medium · verified 2026-08-10
- SUPERSEDED by clm-0042's per-task measurement, which found the true energy cost roughly 10x lower — this run's figures are whole-arm totals padded by model loading and non-scoring tasks, not the model's actual energy per answer, and must not be used to rank models. As measured here: a correct τ² answer cost 78.0 Wh on the 122B with f16 KV, 100.1 Wh on Nemotron, and 115.3 Wh on the 122B with q8_0 KV, but only 8-27% of each arm's wall time fell inside a scored task. aihydra's power envelope stands on its own: idle 10.1 W, 150-168 W under inference, peaking at 218 W.
- Source: Home Assistant recorder, `sensor.hardware_ai_hydra_energy` — a CUMULATIVE kWh counter at the wall socket, sampled every ~10 seconds at full float precision. Measured at the SMART PLUG, so this is whole-box draw — CPU, 128 GB RAM, two NVMe, fans and PSU losses included — not GPU package power. ## Method: counter differences, not averaged power Per-arm energy is the counter's value at the end minus its value at the start. No integration of noisy power samples, no assumption about duty cycle. At ~150 W a 10-second boundary error is ~0.0004 kWh, i.e. negligible. I first computed these from HOURLY statistics, which the operator rightly flagged as too coarse — over a 6-hour arm the boundary error is tolerable, but it is meaningless for the minutes-long performance runs. Re-measured at 10-second resolution the values moved by at most 1.2%, so the conclusions were not wrong, merely imprecise. The method now works at any run length, which is what matters for the backfill. ⚠ **Raw 10-second history is retained ~10 days.** Long-term statistics persist forever but only hourly. So the Aug 7-8 performance runs must be extracted before roughly Aug 17-18 or they drop to hourly resolution permanently. ## ⚑ This field was empty for 99 runs before today The `energy` field has been in the run schema since the beginning and every one of the 99 recorded runs carries `energy: null`. The operator asked whether it was being captured; it was not. Wall-measured energy is the thing this lab has that published benchmarks almost universally lack, and it was designed in and then never populated. ## The power envelope | state | wall power | |---|---| | idle | **10.1 W** | | 122B under τ² load | ~154 W mean | | Nemotron under τ² load | ~167 W mean | | peak observed | **218 W** | Nemotron draws about **8.5% more** than the 122B for the same work — consistent with the 86 °C spike observed on it against the 122B's steady 68-74 °C. ## ⛔ WHAT "PER CORRECT ANSWER" ACTUALLY MEASURES HERE — read before quoting The operator, 2026-08-10: *"you had the energy delta for a full run, so including some failed tasks right? this would mean that the 'energy per correct' would be higher than that model?"* Correct, and the problem is larger than failed tasks. The figures below are **total arm energy divided by correct answers**. Two consequences: **1. Failed-task energy is included.** Deliberate — one useful result should carry the cost of the wrong ones. But it means this is NOT the model's per-task energy, and the label invites that reading. **2. Most of the energy was not spent on scored tasks at all.** Checking the summed task durations against the 360-minute arms: | arm | arm length | time inside scored tasks | fraction | |---|---|---|---| | 122B f16 | 360 min | 62.8 min | **17%** | | 122B q8_0 | 360 min | 27.5 min | **8%** | | Nemotron off | 360 min | 98.5 min | **27%** | So 73-92% of the measured energy went on model loading and on tasks that never scored. Arm 2's 4.4-hour non-terminating task is **not among its nine scored results** — 4.4 hours of electricity yielding no measurable outcome, then attributed to the eight answers that did land. **3. The arms cover DIFFERENT TASK SETS.** Bound-limited at 6h, arm 1 reached tasks 0-21, arm 2 only 0-8, arm 3 0-15. Comparing them compares different mixes of work — the same contiguous-prefix confound `clm-0037` names, which I flagged there and then ignored here. **So these numbers answer "run the box six hours; what does each correct answer that falls out cost?" — a real question, but not "how efficient is this model".** They should not be used to rank models. The sound comparison is the paired one in `clm-0038`, and a sound energy comparison needs matched task sets with per-task windows, which requires the `started_at`/`ended_at` instrumentation added on 2026-08-10 and does not exist for these arms. ## Energy per correct answer, as measured (see caveats above) Six-hour arms, so total energy is nearly identical across them; what differs is how many correct answers each bought. | arm | counter start -> end (kWh) | delta kWh | tasks | correct | Wh/task | **Wh per CORRECT answer** | marginal* | |---|---|---|---|---|---|---|---| | 122B f16 | 4.804997 -> 5.740740 | **0.93574** | 22 | 12 | 42.5 | **78.0** | 72.9 | | Nemotron off | 6.662776 -> 7.663805 | **1.00103** | 16 | 10 | 62.6 | **100.1** | 94.0 | | 122B q8_0 | 5.740740 -> 6.662776 | **0.92204** | 9 | 8 | 102.4 | **115.3** | 107.7 | *delta = total minus the 10.1 W idle floor, i.e. energy attributable to inference rather than to the machine merely being on. `docs/lab-site-design.md` calls this "the only figure that means anything", because a naked wattage reading mostly measures the idle floor. ## In money, which is the point of using kWh At the Grid Import Price observed today, **30.3 p/kWh**: | arm | Wh per correct answer | **pence per correct answer** | delta-only | |---|---|---|---| | 122B f16 | 78.0 | **2.36 p** | 2.21 p | | Nemotron off | 100.1 | **3.03 p** | 2.85 p | | 122B q8_0 | 115.3 | **3.49 p** | 3.26 p | So the q8_0 KV penalty is about **1.1 p per correct answer** — small per answer, and the kind of number that only becomes visible when denominated in something a person can price. ⚠ Tariff caveat: this uses the instantaneous import price. The house has solar and a battery, so the marginal cost of a run is lower — sometimes zero — when it lands in a solar or cheap-rate window. `docs/lab-site-design.md` reserves a `tariff_window` field for exactly this; these figures are grid-import-equivalent, not what was actually paid. ## ⚑ UNITS: kWh / Wh / pence, never joules `docs/lab-site-design.md` already specified this — *"mWh per task, Wh per 1,000 tasks, kWh per month. Joules is not a home-energy unit and doesn't map to /kWh tariffs."* I wrote "denominated in joules" and "J/token" anyway. Same failure as drifting from doc 14's time-to-correct headline: the project had decided, and I did not check. ## ⚑ Quantised KV is an energy REGRESSION Same model, same tasks, only `-ctk/-ctv` differing: **78.0 Wh per correct answer with f16 against 115.3 Wh with q8_0 — a 48% penalty**. Quantised KV saves memory and spends electricity. That is a second, independent instrument agreeing with `clm-0038`, which found q8_0 needs **+39% more turns** on paired tasks. Extra turns are extra work, and work is electricity. Two measurements of different quantities pointing the same direction is much stronger evidence than either alone — and neither was visible in the reward metric, which moved by a single task. It also reframes `clm-0022`. That claim established the dequant patch recovers +70.3% throughput, making quantised KV look free. At the wall it is not free: on the stock build it costs half again as much energy per useful result. Whether the patch removes the energy penalty along with the speed penalty is now a well-posed question with a way to answer it. ## Scope · Arms were bound-limited at 6h and are contiguous task prefixes, not random samples (clm-0037). The q8_0 arm completed only 9 tasks because one ran 4.4 h, which inflates its Wh/task — but that IS the cost, not an artefact. · Boundaries are resolved to ~10 s from the raw recorder, not to the hour. The residual error is ~0.0004 kWh per boundary, far below the effects being compared. · Whole-box measurement includes anything else the machine was doing. The arms ran with the box otherwise quiet, but model downloads overlapped part of the first arm's window. · The 10.1 W idle figure is the box powered but not serving. A model held resident in memory would sit higher, so marginal energy here slightly overstates the true increment for an always-warm server. · Wall measurement excludes nothing on the machine but does exclude network gear and the client side.
- evidence: run-0100 run-0101 run-0102
clm-0041
- measured-here med ●●○ volatility medium · verified 2026-08-10
- The KV dequant patch cuts energy 42% at 200k context — 146.1 Wh unpatched against 85.3 Wh patched for the same throughput benchmark — and patched q8_0 (85.3 Wh) beats f16 (89.6 Wh). That INVERTS clm-0040's agentic finding, and both are correct: on fixed-token throughput work quantised KV wins once patched, while on agentic work it loses because it spends extra TURNS. The workload decides, not the flag.
- EVIDENCE TRAIL: `eng-0017` (stock q8_0), `eng-0018`/`run-0037` (kvfix q8_0), `eng-0019` (stock f16), and `eng-0020` (kvfix f16) are the four cited windows. Three of the four carry `run: []`; they are source-JSON-matched/reconstructed windows rather than joined raw run records. Wall energy from `sensor.hardware_ai_hydra_energy`, a cumulative kWh counter at ~10s resolution, differenced across each benchmark invocation's window. Units Wh / pence. ## The 200k-context comparison Four invocations of the same deep-context benchmark, Aug 8, differing only in build and KV type: | build + KV | wall time | **Wh** | mean W | pence | |---|---|---|---|---| | stock q8_0 | 51 min | **146.08** | 171.7 | 4.43 p | | **kvfix q8_0** | 29 min | **85.25** | 178.5 | 2.58 p | | stock f16 | 30 min | 89.55 | 177.7 | 2.71 p | | kvfix f16 | 31 min | 90.43 | 177.7 | 2.74 p | Note the mechanism: power draw is essentially identical across all four (172-179 W). The energy difference is **entirely wall time**. The patch does not make the machine sip less; it makes it finish sooner. That is `clm-0022`'s +70.3% decode recovery, expressed in the unit that appears on a bill. **Patched q8_0 is now the cheapest option at depth** — 85.3 Wh against f16's 89.6 Wh — so once the patch is applied, quantised KV saves both memory and electricity on this workload. ## ⚑ This inverts clm-0040, and both results stand `clm-0040` found q8_0 costs **48% MORE** energy per correct answer on τ² agentic tasks. This claim finds it costs **42% LESS** on a throughput benchmark. Not a contradiction — they measure different workloads: · **Throughput work** processes a FIXED token count. Faster decode means less wall time means less energy. Quantised KV wins once the dequant path is fixed. · **Agentic work** has a VARIABLE turn count. `clm-0038` measured q8_0 taking +39% more turns on paired tasks, so it does more total work — and the extra work outweighs any per-token saving. So the correct statement is not "q8_0 is efficient" or "q8_0 is wasteful", it is: **q8_0 is efficient per token and can be wasteful per task, and which dominates depends on whether the workload's length is fixed or emergent.** A benchmark that only measured tok/s would have reported the first half and missed the second entirely. ⚠ The agentic measurement (clm-0040) was on the STOCK build. Whether patched q8_0 also recovers the agentic energy penalty is untested — the turn inflation may or may not be a consequence of the same broken dequant path. That is the obvious next run. ## Method and provenance Windows are **RECONSTRUCTED, not recorded**: each llama-bench invocation wrote a JSON whose mtime is that invocation's end, so the window is [previous end, this end]. Run records only ever stored a `date`, never timestamps — the gap this exercise exposed. · Windows include model load time, which is part of the cost but not part of inference. · Where invocations are separated by more than an hour, a 5-minute lead-in is assumed rather than attributing hours of idle to a run. · One artefact is visible and left in deliberately: `deepkv200-080145/stock-q8_0` reads 0.84 Wh at 10.1 W — exactly the idle floor, i.e. an aborted run. Reconstruction that surfaces its own failures is worth more than one that hides them. · Peak draw is remarkably consistent across every benchmark on this box: 205-215 W.
- evidence: eng-0017 eng-0018 eng-0019 eng-0020 run-0037
clm-0042
- measured-here med ●●○ volatility medium · verified 2026-08-10
- Per-task energy windows, matched to a common task set across arms, put a correct τ² answer at 6.81 Wh on the 122B with f16 KV, 9.48 Wh with q8_0, and 12.75 Wh on Nemotron. That is 0.21 p, 0.29 p and 0.39 p at 30.3 p/kWh. The q8_0 penalty is +39%, the same figure clm-0038 measured for turns, which is the mechanism. These supersede clm-0040's arm-level numbers, which were ~10x too high because ~90% of an arm's energy went on model loading and tasks that never scored.
- EVIDENCE TRAIL: `run-0100`/`eng-0174`, `run-0101`/`eng-0175`, and `run-0102`/`eng-0176` are whole-arm records. The headline 6.81 Wh, 0.21 p, and 30.3 p/kWh values below are matched-task derived calculations over the common scored task set, recorded in this claim note; they are not the `wh_per_task` fields on the whole-session energy records, which include load, gaps, and unscored work. ## Method — and why it took three attempts Each τ² simulation records `start_time` and `end_time`. Energy is the cumulative wall counter differenced across each TASK's window and summed, so it excludes model load, inter-task gaps, and tasks that never scored. Restricted to the **9 tasks common to all three arms**, which removes the contiguous-prefix confound (`clm-0037`) — bound-limited arms reached different depths into the task list. Three passes at this number, each wrong for a different reason: 1. **Hourly statistics** — too coarse; the operator flagged it. Fixed by using the 10-second raw counter. 2. **Whole-arm energy ÷ correct answers** — the operator asked whether failed tasks were included. They were, and worse: only 8-27% of each arm's wall time was inside scored tasks at all, so the figures mostly measured loading and a 4.4-hour non-terminating task that never scored. 3. **This.** Per-task windows, matched task set. ⚑ The transcripts carried `start_time`/`end_time` all along. I concluded per-task windows were unavailable and built new instrumentation for future runs — correct to do, but the data for these arms was already there and I had not looked. The operator asked "the tau2 transcripts don't contain timestamps?" and they did. ## The measurement | arm | correct | in-task time | total Wh | Wh/task | **Wh per correct** | **pence per correct** | |---|---|---|---|---|---|---| | 122B f16 | **9 / 9** | 22.4 min | 61.30 | 6.81 | **6.81** | **0.21 p** | | 122B q8_0 | 8 / 9 | 27.5 min | 75.84 | 8.43 | **9.48** | **0.29 p** | | Nemotron off | 8 / 9 | 36.3 min | 102.00 | 11.33 | **12.75** | **0.39 p** | ## What it says **The 122B with f16 KV is the cheapest per useful result, by a clear margin** — 6.81 Wh against Nemotron's 12.75, so Nemotron costs **87% more electricity per correct answer**. On the same nine tasks it also took 62% longer in-task (36.3 min against 22.4). **q8_0 costs +39% per correct answer against f16 on the identical model.** `clm-0038` measured +39% TURNS on this same paired task set. Energy and turns agreeing to the percentage point is strong evidence that the extra turns *are* the mechanism — quantised KV does not draw more power, it does more work. This also completes the picture with `clm-0041`, which found the dequant patch cuts energy 42% on fixed-token throughput runs and makes patched q8_0 cheaper than f16. Both hold: **quantised KV is cheaper per token and dearer per task**, and which dominates depends on whether the workload's length is fixed or emergent. The agentic measurement here is on the stock build; whether the patch also removes the per-task penalty is untested and is the obvious next run. ## ⚑ SELF-PLAY OVERHEAD — the user simulator is not useful work The operator, 2026-08-10: *"the model is picking up both sides of the conversation, the actual useful work is the LLM side, not the 'user' side. Do you account for this already?"* I had flagged it in scope and NOT corrected for it. Correcting now. τ² self-play runs the agent and the user simulator on the same model and the same box, so the measured energy includes generating the customer's half of the dialogue — scaffolding, not work anyone would pay for. Only `assistant` messages carry `generation_time_seconds`, but that bounds it: | arm | agent generation | share of task wall time | remainder | |---|---|---|---| | 122B f16 | 1031.4 s | **76.8%** | 23.2% | | 122B q8_0 | 1353.4 s | **82.0%** | 18.0% | | Nemotron off | 1726.1 s | **79.3%** | 20.7% | Apportioning energy by that share: | arm | measured Wh/correct | **agent-only Wh/correct** | vs f16 | |---|---|---|---| | 122B f16 | 6.81 | **5.23** | — | | 122B q8_0 | 9.48 | **7.77** | **+49%** | | Nemotron off | 12.75 | **10.11** | **+93%** | **The ranking is unchanged and the gaps widen.** So the conclusions hold, but the headline numbers were ~20-25% too high as a measure of useful work. Treat the agent share as a LOWER bound on its share of energy. The non-agent remainder is user-simulator generation PLUS tool execution and framework overhead, and tool calls are local database operations that draw near-idle power. So the agent's share of energy ABOVE IDLE is higher than its share of wall time — the true correction is smaller than 20-25%, and the measured figures are conservative rather than optimistic. A cleaner design would run the user simulator on a different host, or on a small model whose cost is separately accounted. That is a benchmark-harness change, not an analysis one, and worth doing before energy figures are quoted as model properties. ## Scope · 9 matched tasks, one domain, self-play, deterministic scoring, read paths only. · Wall measurement, so whole-box: CPU, RAM, NVMe, fans, PSU losses. Idle floor 10.1 W is included in these figures, not subtracted — at ~160 W active it is ~6% of the total. · Tariff is the observed 30.3 p/kWh grid import. Solar and battery mean realised cost is lower, sometimes zero. · ⚠ The 122B-vs-Nemotron comparison carries the unpinned-simulator confound documented in `clm-0043` and should be read as provisional; the f16-vs-q8_0 comparison does not, since both arms ran the same model as its own simulator.
- evidence: run-0100 run-0101 run-0102 eng-0174 eng-0175 eng-0176
clm-0043
- measured-here high ●●● volatility low · verified 2026-08-10
- Every τ² arm was run with --user-llm set to the same model as --agent-llm, so cross-model comparisons changed the agent AND the user simulator together — exactly what the runbook forbids ("hold both --user-llm and the judge fixed across comparisons, or results re-baseline silently"). The 122B-vs-Nemotron comparisons are therefore confounded. The f16-vs-q8_0 comparisons are NOT, because both arms ran the same model on both sides.
- **Core finding:** every τ² arm ran `--agent-llm "openai/$MODEL" --user-llm "openai/$MODEL"` — the user simulator was always the model under test, which `docs/benchmark-runbook.md` already said not to do. Comparisons where the simulator changed alongside the agent (122B-vs-Nemotron: `clm-0039`, `clm-0042`; thinking on/off: `clm-0033`, `clm-0035`) are confounded and should be read as provisional. Comparisons where the same model ran both sides in every arm (f16-vs-q8_0 KV: `clm-0038`, `clm-0042`; the dequant-patch throughput result, `clm-0041`, which has no simulator at all) are sound — the simulator was held constant by accident rather than by design. **The fix:** pin `--user-llm` to a single independent model for all future arms and record it in the run's config fingerprint. Re-running the confounded comparisons under a pinned simulator is the only way to de-confound them. The full argument for why an unpinned simulator has no predictable bias direction, the correlated-blind-spot problem, and the eight properties a user simulator actually needs are written up as a standing rule in `docs/methodology-lessons.md` §3.
clm-0044
- community med ●●○ volatility high · verified 2026-08-11
- A second independent Strix Halo source (llama.cpp PR #26856 + its Reddit write-up) reports Vulkan ahead of ROCm on decode at depth by ~10.5% on a clean same-binary comparison — same direction as clm-0031's +55% but a fifth the magnitude, confirming that figure was mostly build-gap and private patches. The PR itself adds a native BF16 flash-attention path for RDNA3+ that inverts the PREFILL gap (patched ROCm ~40% ahead of Vulkan) and delivers near-F32 KV quality (+0.04% PPL vs F16's +5.8%). Unmerged, no maintainer review yet.
- Source: r/StrixHalo post by u/Look_0ver_There (GitHub stew675 — strongly implied same person, not confirmed), llama.cpp PR #26856 "bf16-tile-packed-q", head 9921e01, base dd1ea524 (2026-08-10) — AFTER both our builds (3653e6d, min-62bf73d), so none of this is in anything we run. Verified open/unmerged via GitHub API; only procedural comments so far. ## The lever llama.cpp silently converts BF16 KV to F16 before flash-attention on all backends. The PR adds a native BF16 tile path gated on `V_DOT2_F32_BF16_AVAILABLE` (RDNA3/3.5/4 only, gfx110x/115x/120x — includes our gfx1151), selected automatically when both cache types are bf16. No new flags. NVIDIA unaffected. Author's numbers (Qwen3.6-35B-A3B-Q8_0 unless noted): · F16 KV, both backends stock, same binary: Vulkan pp1024 717.93 vs ROCm 618.31 (+16%), tg256 46.86 vs 42.42 (+10.5%) — the CLEANEST of the three community comparisons. · BF16 KV, ROCm patched: ROCm pp1024 684.02 vs Vulkan 484.00 (~40% ahead); decode gap unchanged (Vulkan 47.59 vs 42.51). · PPL @32k (Qwen3.5-4B, wikitext-2): F32 8.6368, patched BF16 8.6403 (+0.04%), F16 9.1400 (+5.8%) — BF16 KV as a near-free QUALITY upgrade over F16 is the sleeper finding, directly relevant to our KV-quality thread (clm-0038/0041/0042). ## The decode-gap mechanism (unverified, no artifact) Author's per-op breakdown claims ROCm kernels match Vulkan's (20.93 vs 20.9 ms summed) but lose ~3.82 ms/token to inter-kernel dispatch gaps because HIP graphs never stabilise. UPDATE 2026-08-13: the author PUBLISHED the fix and RETRACTED that mechanism — "HIP_GRAPHS=ON was, in fact, working. It just wasn't really providing any real benefit." Corrected root cause: ROCm's HSA AQL dispatch floor (~2-3 us/kernel, ~974-1624 kernels/token on the 35B MoE) vs Vulkan's cheaper command-processor path — a library-level difference, mitigated by KERNEL FUSION, not graph repair. Public artifact: github.com/stew675/llama.cpp branch rdna-boosts (consolidated, includes the PR-26856 BF16-KV path, the fusion campaign, a gfx1151 mmvq table for Q8_0 decode, and a revert of upstream #24233 — which he identifies as the root of the Strix Halo async-race KV corruption, superseding the HIP_LAUNCH_BLOCKING PSA on his branch; mmap stays broken either way). Claimed results: ROCm-vs-Vulkan decode gap 12.2% -> 3.2% at d32768 (BF16 KV, Qwen3.6-35B) and ROCm prefill +36.5% over Vulkan; residual gap attributed to the dispatch floor itself. Now TESTABLE against clm-0050's stock matrix — queued. Source: r/StrixHalo thread (PR-26856 write-up, edits 11-12 Aug). ## What it does to clm-0031 Supports the direction (Vulkan genuinely ahead on decode-at-depth on this silicon), undercuts the magnitude (~10.5% clean vs 55% confounded). clm-0031's own hedge — "the backend, their patches, or the 43-build gap" — resolves as: mostly the latter two. ## Test plan (queued under the anchor-pair policy) One dual-backend binary (GGML_HIP=ON + GGML_VULKAN=ON) at base dd1ea524 with the PR branch cherry-picked; backend selection becomes the only variable. Anchor pair vs 3653e6d first to isolate the 3-day upstream drift. Cells: 35B (the post's own model, already screened here), pp512/1024 + tg128/256, shallow + d32768, f16 AND bf16 KV both backends, -fa 1, --parallel 1. Predictions on record: F16 Vulkan +16/+10.5%; BF16 ROCm +40% prefill, decode gap persists. ## ⚠ Side-findings from the same author's linked PSA (separate post, unverified) · `HIP_LAUNCH_BLOCKING=1` reportedly required on some recent ROCm versions to avoid SILENT KV-cache corruption on Strix Halo. Not in our protocol.json; our ROCm 7.1.0 may or may not be in the affected range — version labels ambiguous. · **MTP draft verification may pass corrupted tokens** — "not even a depth of 1 is truly safe". Production's 122B runs MTP (clm-0022). Nothing in our records addresses MTP output-correctness. Flagged for its own investigation before aibeast's production restore; a correctness question, not a performance one.
clm-0045
- measured-here med ●●○ volatility medium · verified 2026-08-11
- The KV dequant patch removes most of quantised KV's agentic cost, not just its speed cost: on identical seeded tasks, patched q8_0 takes +9.3% more turns than patched f16 (234 vs 214 over 11 paired tasks) where the stock build cost +39% (clm-0038). Reward is near-identical (10/11 vs 11/11). Separately, the non-termination marathons that were attributed to q8_0 strike f16 too — this run's 3.3-hour blowup was on f16 while q8_0 solved the same task in 18 turns — so task blowups look stochastic, not KV-caused.
- First run under the fixed methodology: PATCHED build (ce7689f asserted), identical 12-task subset via --task-ids 0-6,8-12, --seed 42, --max-steps 200, protocol-checked, windows recorded by instrumentation (run-meta.jsonl), Qwen3.5-122B both sides (within-model lever test — simulator confound does not apply, clm-0043). ## Paired table (reward, turns, minutes) | task | patched f16 | patched q8_0 | |---|---|---| | 0 | 1.0, 20, 2.0 | 1.0, 12, 1.6 | | 1 | 1.0, 24, 2.0 | 1.0, 24, 3.2 | | 2 | 1.0, 29, 3.5 | 1.0, 22, 3.2 | | 3 | 1.0, 14, 1.2 | 1.0, 18, 1.7 | | 4 | 1.0, 12, 1.4 | 1.0, 24, 2.7 | | 5 | 1.0, 15, 2.1 | 1.0, 24, 3.3 | | 6 | 1.0, 10, 1.5 | 1.0, 12, 1.4 | | 8 | 1.0, 26, 3.3 | 1.0, 32, 3.8 | | 9 | 1.0, 12, 2.3 | 1.0, 16, 2.6 | | 10 | 1.0, 30, 6.3 | 1.0, 26, 11.9 | | 11 | 1.0, 22, 2.9 | **0.0**, 24, 3.3 | | **Σ common** | **214** | **234 (+9.3%)** | | 12 | unscored — killed at the arm's 4h bound after ~3.3h | **1.0, 18, 3.4** | ## The three findings **1. The patch closes ~75% of the agentic turn gap.** Stock: +39% turns (clm-0038, matched to the percentage point by energy, clm-0042). Patched: +9.3%. Direction on signs: q8 longer on 6 tasks, shorter on 3, ~tie 2 — consistent but weak at n=11; the residual may be real or may be noise. Combined with clm-0041 (patch = 42% energy saving on throughput work), the patched picture is: quantised KV costs ~nothing in speed, little in reward, and possibly a single-digit-percent turn overhead on agentic work. **2. Task blowups are not a q8_0 property.** Stock runs kept drawing multi-hour non-terminating tasks on q8_0 arms (4.4h, clm-0040), which fed a "quantised KV fails to terminate" narrative. Here f16 drew the blowup — 3.3h on task 12 without converging — while q8_0 finished it correctly in 3.4 minutes. Same model, same seed, same subset. Non-termination looks like a stochastic simulator-path phenomenon that any arm can draw, which also means single-arm wall-times are a poor basis for KV conclusions. **3. Reward stayed flat.** 11/11 vs 10/11 — the one q8 miss was a wrong answer at normal length (24 turns, user_stop), not a runaway. Within small-n noise. ## Scope · n=11 pairs, one domain, self-play (fine for a within-model lever), one seed. The +9.3% residual needs repeats (--num-trials) before it is a number rather than a direction. · Wall-clock asymmetry (f16 4h bound-cut vs q8 62 min) is dominated by the single f16 blowup and must NOT be read as f16-is-slower. · Per-task energy from the recorded windows to follow as eng- records; arm-level energy is marathon-skewed and deliberately not quoted here (clm-0042's lesson).
- evidence: run-0103 run-0104
clm-0046
- measured-here med ●●○ volatility medium · verified 2026-08-11
- The community BF16 flash-attention predictions (clm-0044) reproduce on this hardware under single-binary methodology: at 32k depth, stock Vulkan leads ROCm +17.4% prefill / +17.1% decode at f16 KV, and the PR-26856 patch inverts prefill to ROCm +46.9% with bf16 KV. At shallow depth every gap collapses (+2.6% / -1.0%), so backend comparisons are depth-statements or they are nothing. The fleet build delta is +2.8% pp / +0.9% tg with identical tau2 capability (anchor pair).
- One binary (77a9a66eb, HIP+Vulkan), device selection the only variable — tighter than either community source. Vulkan bf16 cells succeeded despite the device banner reporting bf16:0 — treated as a cast/fallback path, not native compute. The depth cells were re-run after a SIGPIPE/pipefail flag-detection bug silently substituted a prefill proxy (which would have shown +24.6% instead of +46.9% — the bug class is documented in methodology-lessons). PR 26856 remains unmerged and unreviewed upstream: adopt-watch, with the anchor-pair policy governing any adoption.
- evidence: run-0105 run-0106
clm-0047
- measured-here med ●●○ volatility medium · verified 2026-08-11
- Nemotron's quant confound resolves cleanly: on 12 identical seeded tasks, UD-Q4_K_M and UD-IQ4_XS produce IDENTICAL reward on every task (0.583 both), but IQ4_XS takes +11% more total turns (630 vs 567). Quantisation cost this model efficiency, never correctness — prior cross-model comparisons at mismatched quant understated Nemotron's speed, not its quality. At fair quant its turn median (25) still trails the 122B's (~20) on the same subset: the incumbent's efficiency lead narrows but survives.
- Stock 3653e6d, seeded subset (tasks 0-6,8-12), --max-steps 200, self-play (sound for a within-model lever, clm-0043). Energy followed duration, not draw: E1 ran 1.86x as long as E2 at near-identical wattage (eng-0061/0062). One task cut per arm. AUDIT ANNOTATION (2026-08-12 re-execution spot-check): the absolute 0.583 mean is partly a simulator artifact and MUST NOT be compared against pinned-simulator runs of other models. Re-running the E1 arm on tasks 0-3 under the protocol-pinned Haiku simulator flipped tasks 1 and 2 from 0.0 to 1.0 (trajectories: the self-play user talked the agent into DB-mutating policy violations the independent simulator never induces; total turns 158 -> 76). The within-configuration quant equivalence claimed here is unaffected. Cross-model comparison at fair quant requires re-running both arms under the pinned simulator. See audits/repeatability-2026-08.md.
- evidence: run-0107 run-0108
clm-0048
- measured-here med ●●○ volatility medium · verified 2026-08-11
- Simulator identity changes what tau2 measures. On identical seeded tasks with the same 122B agent: self-play and a Haiku-4.5 simulator produce identical rewards (9/9) but Haiku lengthens conversations (162 to 210 total turns) — the self-play customer volunteers what its twin needs, the clm-0043 correlated-blind-spot effect now measured. A Sonnet-4.5 simulator failed the agent on 2 of 9 tasks; forensics split them — one simulator persona violation (discounted), one genuine agent bug (clm-0049). The registered prediction that an off-box simulator would cut wall time FAILED: self-play was faster per task (median 121s vs 173s) and cheaper in local energy (4.4x less per task), because the independent simulator's longer conversations mean more agent computation — not because of wait-state draw, which resident-quiet measurement rules out (13.6 W).
- Haiku pin CONFIRMED as primary: adequate on rewards, independent by family, and the cheapest option with evidence. Sonnet proved less persona-constrained rather than more discriminating (task 5: its customer pivoted to demanding the cancellation the scenario explicitly forbids). Periodic Sonnet secondary passes remain worth considering — its unconstrained pressure DID surface clm-0049's genuine bug, even if by accident. The self-play speed/energy advantage is real but buys measurements of an easier, flattered benchmark: correctness of measurement beats cost here. Provider routing is account-guardrailed to Anthropic (verified first-probe leak to Bedrock, then dashboard- confirmed). Wall-time comparison cannot separate cache effects from API latency — per-turn agent-side timing would; registered and closed as a failed prediction.
- evidence: run-0109 run-0110 run-0105
clm-0049
- measured-here high ●●● volatility medium · verified 2026-08-11
- The 122B has a reproducible rule-precedence bug in policy application: on tau2 airline task 9 it cancels a partially-flown reservation in 7 of 8 trials, every failure with the identical signature — cancel_reservation called without ever checking flight status. It applies the permissive business-class exception without gating on the hard already-flown precondition. Under self-play the task passed consistently — the twin customer steered around the trap — so the defect was invisible until an independent simulator was pinned.
- Stage G: --num-trials 8, seed 42, pinned Haiku simulator, temperature-0 serving; 7/8 failures, 7/7 signature matches, detection logic validated against the original F2 failure transcript first. The passing trial is the outlier (simulator-side sampling variance). The error class — a permissive exception short-circuiting a hard precondition — generalises beyond airline policy and joins MTP output-verification on the production-restore checklist: any agentic deployment of this model wants a hard-precondition guard pattern in its policy prompts or tooling. Candidate for a targeted guard probe in the standard screen.
- evidence: run-0111
clm-0050
- measured-here med ●●○ volatility medium · verified 2026-08-14
- AMENDED 2026-08-14 — the rule below is PER-MODEL AND PER-PHASE, not fleet-wide; see the correction history and the amendment note. On gfx1151 at f16 KV, stock Vulkan beat stock ROCm in EVERY cell of a matched matrix measured on Qwen3.6-35B-A3B (one binary commit 3653e6d, depths 0 to 131,072): decode +19-21% at every depth, prefill +4% to +20% growing with depth to 65k. That finding was correct as measured and remains the reference case for this model. It does NOT generalise: the 2026-08-14 perf-matrix sweep found gpt-oss-120b's PREFILL inverts hard in ROCm's favour at depth — Vulkan −52% at d65536, −84% at d131072 (23.1 t/s vs ROCm's 144.3, tight across 3 reps) — while gpt-oss-120b's own DECODE still favours Vulkan (+10 to +16%), same direction as Qwen3.6-35B. Nemotron-3-Super is now a THIRD supporting data point rather than an absence: the OOM first recorded against it (twice, including at an expanded 122 GiB GTT ceiling) was a benchmarking-harness artifact of running stock ROCm under mmap on a model within ~40 GiB of the box's full memory, corrected 2026-08-14 — see clm-0053. On the identical binary with `--load-mode none`, Nemotron-3-Super UD-Q4_K_M at d32768 measures 244.05 pp / 16.43 tg (median of 3 post-reboot reps; an earlier single-rep check the same day landed within 2% at 248.79 pp / 16.39 tg) against Vulkan's own 139.8 pp / 17.56 tg at the same cell: ROCm +74.6% prefill, Vulkan +6.4% decode — the same prefill-favours-ROCm, decode-favours-Vulkan split gpt-oss-120b shows, now measured on a third architecture. The r/LocalLLaMA claim that ROCm leads Vulkan 3.5x at 65k depth — measured on an RDNA2 V620 — still inverts on gfx1151 for Qwen3.6-35B decode/prefill and for gpt-oss-120b decode, but NOT for gpt-oss-120b's own prefill, which is the live counter-example. The community fork (v0.6.1) over stock Vulkan remains a prefill-only win at f16 KV on Qwen3.6-35B: +13/+13/+1/+2.5/+16% by depth, decode unchanged (±1%), and no stride bug (41/41 layers on GPU, CPU utilisation identical to stock) — untested on the other two models.
- METHOD — three arms, one model, matched flags. A = stock llama.cpp 3653e6d ROCm, B = stock 3653e6d Vulkan, C = fork v0.6.1 commit 3be50ccc2 with its bundled RADV 26.3.0-devel. Qwen3.6-35B-A3B UD-Q4_K_XL, fa on, pp1024/tg256, llama-bench defaults otherwise (-b 2048 / -ub 512), median of 3 fresh-process reps with page cache dropped between reps; 72/72 reps rc=0 across both KV matrices. One benign outlier rep in the matrix (median absorbs it). f16 KV, pp/tg by depth: | depth | A ROCm | B Vulkan | C fork | |---|---|---|---| | 0 | 1068.1 / 51.1 | 1110.5 / 61.6 | 1251.5 / 61.8 | | 4,096 | 962.6 / 49.7 | 1020.0 / 59.3 | 1149.5 / 59.9 | | 32,768 | 595.8 / 42.3 | 697.7 / 50.4 | 705.4 / 51.0 | | 65,536 | 412.5 / 36.4 | 495.1 / 43.5 | 507.3 / 43.9 | | 131,072 | 253.6 / 28.5 | 278.3 / 34.2 | 321.6 / 34.5 | The counter-claim this answers: a r/LocalLLaMA report of ROCm leading 3.5x at 65k depth was measured on a V620 (RDNA2 — a different GPU family). At 65k on gfx1151 the same comparison reads Vulkan +20% prefill / +19.5% decode. Backend verdicts do not transfer across GPU families; this claim is scoped to gfx1151. FORK ATTRIBUTION CAVEAT: arm C differs from arm B in two ways at once — the fork's kernels AND the newer bundled Mesa (RADV 26.3.0-devel vs system RADV). The d0/d4096 prefill edge is plausibly part-Mesa; the cells are not single-variable the way clm-0022's cherry-pick was. Decode being flat (±1%) at f16 is expected per the fork docs for an hd256 MoE — hd128 models are documented to gain more, so these fork deltas are a FLOOR, not the fork's headline. Single model tested; medium confidence until a second model (ideally hd128 dense) fills the matrix. Stride-bug check: v0.6.1 shows no stride bug — 41/41 layers verified on GPU, CPU% identical to stock during runs. RELIABILITY COUNTERWEIGHT (2026-08-13): llama.cpp issue 25664 (open, field reports incl. Strix Halo) documents Vulkan/RADV DeviceLostError at ~80k context on this hardware class - our matrix ran clean to 131k, but a naive "switch to Vulkan" operational conclusion should carry this open issue until it resolves. CHALLENGER ON RECORD (2026-08-13): stew675's public rdna-boosts branch (kernel-fusion campaign, see clm-0044) claims to cut stock ROCm's decode deficit to ~3-5% and take prefill +36.5% past Vulkan with BF16 KV. Untested here; a rdna-boosts-vs-stock-Vulkan rerun of this matrix is queued. AMENDMENT 2026-08-14 — the fleet-wide reading of this claim was wrong; the per-model reading above (everything through "rerun of this matrix is queued") was and remains correct. The 2026-08-14 perf-matrix sweep ran the same fa-on, pp1024/ tg256, 3-fresh-process-rep methodology against two more models, ROCm vs Vulkan, stock 3653e6d, f16 KV: gpt-oss-120b, pp1024/tg256 by depth: | depth | ROCm | Vulkan | Vulkan delta | |---|---|---|---| | 32,768 | 308.2 / 39.0 | 294.8 / 45.1 | pp −4.3% / tg +15.7% | | 65,536 | 226.7 / 33.0 | 108.8 / 36.6 | pp −52.0% / tg +10.7% | | 131,072 | 144.3 / 24.0 | 23.1 / 26.3 | pp −84.0% / tg +9.8% | gpt-oss-120b's prefill does the OPPOSITE of Qwen3.6-35B's as depth grows: ROCm pulls further ahead at every depth measured, reaching a 6.2x lead at 131,072 (144.3 vs 23.1 t/s, medians of 3 reps each, tight — samples in run-0162/run-0180). Decode keeps the Qwen3.6-35B direction (Vulkan ahead, +10 to +16%), so the per-model reversal is prefill-specific, not a wholesale gpt-oss-120b exception — which is exactly why the rule has to be stated per-phase as well as per-model. Nemotron-3-Super, stock ROCm f16 and q8_0, d32768: both OOM before producing a number (run-0234, run-0235; "HSA exception: BadAlloc" x93/x5 in the process stderr, but the kernel-level cause captured via journalctl -k at the same timestamps is "amdgpu: SVM mapping failed, exceeds resident system memory limit" — a host-RAM residency ceiling, not a GTT-window one: re-running at an expanded 122 GiB GTT ceiling did not rescue either cell). Vulkan completed both KV variants cleanly (matrix-medians.json on aihydra). Nemotron "favours Vulkan" only in the degenerate sense that ROCm could not run at all — this is a reliability finding, not a throughput comparison, and should not be read as a third data point for the fleet-wide direction either way. CORRECTED RULE: backend advantage on gfx1151 is per-model AND per-phase. Qwen3.6- 35B's "Vulkan wins everywhere" result was real and is now the exception the rule has to explain, not the rule itself. A fleet verdict needs the phase named and the model named; "Vulkan is faster on this chip" is not a sentence this lab can stand behind without both. SECOND AMENDMENT 2026-08-14 — the Nemotron paragraph directly above ("should not be read as a third data point... either way") is superseded. That OOM was traced to a benchmarking-harness defect, not the hardware: the perf-matrix and perf-matrix- gtt120 llama-bench queues never carried --load-mode none, so stock ROCm was benchmarked under mmap on a model (~82 GB of weights) within ~40 GiB of the box's full 122 GiB. Full mechanism, detection-lag numbers, and the mmap-vs-no-mmap throughput control on a model that fits are in clm-0053; the summary needed here is the corrected cell. Nemotron-3-Super UD-Q4_K_M, stock ROCm 3653e6d, --load-mode none, d32768: | rep | pp1024 | tg256 | |---|---|---| | single (pre-reboot check, run-0243/run-0244) | 248.788416 | 16.390674 | | median of 3 (post-reboot, run-0236/run-0237) | 244.051835 | 16.434903 | The two checks agree within 2%. run-0236/run-0237 is the fully-backed N=3 record (stddev 3.76 pp / 0.03 tg across the 3 post-reboot reps) and is the one this claim quotes. Against Vulkan's own run-0190/run-0191 at the same cell (139.810384 pp / 17.559139 tg, median of 3, stock 3653e6d, mmap default — mmap is not a throughput confound for Vulkan or for any model that fits, per clm-0053's control pair), the ROCm figures read: pp +74.6%, tg −6.4%. (The single pre-reboot check alone would read +77.9%/−6.7% — the same conclusion, slightly noisier.) ROCm's prefill lead and Vulkan's decode lead put Nemotron-3-Super in the SAME prefill-favours-ROCm/decode-favours-Vulkan pattern as gpt-oss-120b, not the degenerate no-data case the first amendment recorded. CORRECTED RULE, restated: the per-model, per-phase rule from the first amendment stands unchanged. What moves is Nemotron-3-Super's place in it — from "no data, reliability finding only" to a third model confirming the same phase split gpt-oss-120b showed. Three of three multi-model cells measured on gfx1151 now show prefill and decode capable of favouring opposite backends on the same model; no cell measured here has shown a backend win both phases at once except Qwen3.6-35B at f16 KV, which remains the outlier the rule exists to explain.
- evidence: run-0114 run-0115 run-0120 run-0121 run-0130 run-0131 run-0136 run-0137 run-0146 run-0147 run-0152 run-0153 run-0162 run-0163 run-0166 run-0167 run-0180 run-0181 run-0184 run-0185 run-0190 run-0191 run-0234 run-0235 run-0236 run-0237 run-0243 run-0244
clm-0051
- measured-here med ●●○ volatility medium · verified 2026-08-13
- Quantised KV on gfx1151 splits three ways by build. Stock Vulkan q8_0 prefill COLLAPSES (-26% to -32% vs its own f16) even as its decode gains. Stock ROCm q8_0 decode CRATERS with depth (-16.8/-25.0/-35.1% at 32k/64k/131k) — the same unpatched-build failure clm-0022 measured and patched on the 122B, reproduced here on a build without the ce7689f kvfix. The community fork rescues the Vulkan collapse completely (q8_0 prefill within 3% of f16 at every depth) while keeping and extending the decode gain (+7.8/+14.1/+23.2% over its own f16). Fork + q8_0 is the best long-context configuration measured on this chip: 42.5 tok/s decode at 131,072 — +23% over the best f16 arm and 2.3x stock ROCm q8_0 — at roughly half the KV memory.
- Same arms, model and methodology as clm-0050 (A = stock 3653e6d ROCm, B = stock 3653e6d Vulkan, C = fork v0.6.1 3be50ccc2 + bundled RADV 26.3.0-devel; fa on, pp1024/tg256, llama-bench defaults -b 2048 / -ub 512, median of 3 fresh-process reps, drop-caches; one benign outlier rep in the matrix). Deep cells only — q8_0 effects are depth effects. q8_0 KV, pp/tg by depth: | depth | A ROCm | B Vulkan | C fork | |---|---|---|---| | 32,768 | 595.7 / 35.2 | 517.8 / 53.8 | 701.1 / 55.0 | | 65,536 | 407.8 / 27.3 | 334.5 / 48.1 | 495.4 / 50.1 | | 131,072 | 251.8 / 18.5 | 197.4 / 40.2 | 330.9 / 42.5 | q8_0-vs-f16 deltas, per arm (pp / tg at 32k, 64k, 131k): - A stock ROCm: pp ~0% at every depth; tg -16.8% / -25.0% / -35.1%. This is clm-0022's stock-build decode crater — repeated KV dequantisation during inference — reproduced on a second model, second chip-generation build. This binary lacks the ce7689f kvfix; the fork lineage carries it. - B stock Vulkan: pp -25.8% / -32.4% / -29.1% (the collapse); tg +6.7% / +10.6% / +17.5%. On stock Vulkan, q8_0 trades a third of your prefill for a decode gain. - C fork: pp -0.6% / -2.3% / +2.9% — collapse fully rescued, q8_0 prefill ≈ f16; tg +7.8% / +14.1% / +23.2%. Quantised KV becomes free-or-better, echoing clm-0022's patched-build result. OPERATIONAL READING: fork + q8_0 at 131k gives 42.5 tok/s decode and 330.9 pp — both the best of any arm/KV combination at that depth — while halving KV memory. CAVEATS: one model, and it is an hd256 MoE — fork docs say hd128 gains more, so fork deltas are a floor. Arm C bundles a newer Mesa with the fork kernels (not single-variable; the d0 prefill edge in clm-0050 is plausibly part-Mesa). No quality measurement anywhere in this matrix — this is throughput only, and clm-0006's source made the same disclaimer.
- evidence: run-0122 run-0123 run-0126 run-0127 run-0138 run-0139 run-0142 run-0143 run-0154 run-0155 run-0158 run-0159
clm-0052
- measured-here med ●●○ volatility medium · verified 2026-08-14
- On gfx1151, stock ROCm 3653e6d's own BF16 KV decode is SLOWER than its own F16 KV decode: 29.26 t/s vs 42.35 t/s at 32,768 depth (Qwen3.6-35B-A3B, fa on) — a 31% penalty for switching to the KV type the community recommends for its quality, with prefill roughly unchanged (587.4 vs 595.0). stew675's rdna-boosts branch (commit ed89854) fixes it with a native BF16 flash-attention tile kernel: 47.31 t/s BF16 decode, +61.7% over stock's own BF16 and stock's best config on this chip. Upstream commit a94d563 — six days ahead of 3653e6d and the exact base ed89854 is rebased on — reproduces the same BF16 penalty (29.12 t/s), so the fix is stew675's kernel work, not upstream drift absorbing it for free.
- METHOD — same rdna-boosts-challenger matrix as the anchor-pair calibration in bench/protocol.json: Qwen3.6-35B-A3B UD-Q4_K_XL, ROCm, fa on, pp1024/tg256, llama-bench defaults otherwise (-b 2048 / -ub 512), median of 3 fresh-process reps with page cache dropped between reps, 45/45 reps rc=0. Four arms at d32768: | build | KV | pp1024 | tg256 | |---|---|---|---| | 3653e6d (stock baseline) | f16 | 595.0 | 42.35 | | 3653e6d (stock baseline) | bf16 | 587.4 | 29.26 | | a94d563 (upstream, 6 days ahead) | f16 | 593.7 | 42.12 | | a94d563 (upstream, 6 days ahead) | bf16 | 588.3 | 29.12 | | ed89854 (stew675/rdna-boosts) | f16 | 607.3 | 44.96 | | ed89854 (stew675/rdna-boosts) | bf16 | 651.4 | 47.31 | BUILDS — 3653e6d: fleet stock baseline, all pre-Aug-10 runs (bench/protocol.json). a94d563: upstream llama.cpp a94d563ed801d1da1b8c2432946de07d0231bb3d (2026-08-13, PR #27026), built on-box 2026-08-14 with the same ROCm flags as every other arm; exists solely to de-confound stew675's branch from ordinary upstream drift — it is the exact commit rdna-boosts is rebased on. ed89854: stew675/llama.cpp branch rdna-boosts, head ed89854b2aeb0e333dd61424f14af2aedaca126e (2026-08-13T21:09Z), built on-box 2026-08-14 with the stock ROCm flags, like-for-like with 3653e6d except the branch's kernel-fusion changes (see clm-0044's "what it does to clm-0031" thread — this is the consolidated branch that PR-26856's BF16 flash-attention tile path became part of). THE TRAP — a94d563 is proof the bf16 penalty is not something upstream fixed in the six days between it and 3653e6d: its own bf16 decode (29.12 t/s) matches stock's (29.26 t/s) within reps. The fix lives ENTIRELY in stew675's fork, which is unmerged and carries a divergence disclosure the moment any config depends on it (src/content.config.ts `runtime.divergence`) — this is a CARRIED FORK, not something a fresh `git pull` on 3653e6d or on any near-future upstream commit will give you. PR-26856, the BF16 flash-attention tile path this branch builds on, is itself still unmerged and unreviewed upstream (clm-0044, clm-0046) — adopt-watch policy applies to any production use of ed89854, not just the PR. CROSS-REFERENCE TO clm-0046 — that claim measured BF16 KV as a near-free QUALITY upgrade over F16 (+0.04% PPL vs F16's +5.8%, PR-26856 author's numbers on a 4B model). The quality case for adopting BF16 KV is sound and this claim does not contest it. What this claim adds is the throughput side the community advice usually omits: on a STOCK build, taking that quality upgrade silently costs roughly a third of your decode speed. "BF16 KV is a free quality win" is only true once ed89854's kernel (or PR-26856's, once/if it lands) is actually in your binary — on stock ROCm it is a real trade, not a free lunch. CAVEAT — single model tested (Qwen3.6-35B-A3B, an hd256 MoE), matching clm-0050/ 0051's own scope caveat; this is the same matrix and the same fleet-wide-vs- per-model lesson clm-0050 was just corrected on, so this claim is deliberately NOT generalised past this one model until a second one fills the matrix. ed89854 also lifts f16 decode over stock (44.96 vs 42.35, +6.2%) and f16/bf16 prefill modestly (607.3/651.4 vs stock's 595.0/587.4) — the bf16 fix is the standout result, not the only one, and it is what makes bf16 ed89854's best ROCm configuration on this chip rather than merely a repaired one.
- evidence: run-0204 run-0205 run-0210 run-0211 run-0212 run-0213 run-0218 run-0219 run-0220 run-0221 run-0226 run-0227
clm-0053
- measured-here med ●●○ volatility medium · verified 2026-08-14
- A benchmarking-harness defect, not a hardware limitation, produced this lab's earlier "Nemotron-3-Super cannot allocate on ROCm at any depth" verdict (clm-0050's first amendment, corrected 2026-08-14). House standard on this hardware is `--load-mode none` (no mmap): it is in the llama-server serving flags and in every tau2/capability queue script, but four llama-bench throughput queue scripts never carried it, so every throughput number taken through them inherited llama-bench's own default of mmap. The gap sat between the serving configuration and the benchmark configuration — same lab, same box, two different flag sets. Consequence measured: Nemotron-3-Super (~82 GB of weights on a 122 GB box) failed on stock ROCm 3653e6d at every depth tested under mmap with an HSA `BadAlloc` (run-0234, run-0235), and was recorded as a categorical backend limitation. Under `--load-mode none`, on the identical binary, it runs: 244.05 pp / 16.43 tg at d32768, median of 3 reps (run-0236, run-0237; an earlier single-rep check the same day landed within 2% at 248.79 pp / 16.39 tg, run-0243/run-0244 — see clm-0050's second amendment for the full comparison against Vulkan). The capability verdict was an artifact of the harness disagreeing with the serving configuration, not a finding about the hardware. Cost quantified: mmap is not a throughput confound for a model that fits. A matched `--load-mode none` vs mmap pair on Qwen3.6-35B-A3B, same box, same boot session, 3 reps each (comfortably inside memory) measured +0.72% prefill / −0.12% decode (run-0238, run-0239 against run-0240, run-0241) — noise. An independent, two-days-earlier measurement of the same mmap arm (run-0116, run-0117) landed within 0.25% of this session's mmap median, so the result is not session-specific. Published relative numbers taken under mmap survive; mmap is a correctness lever only for a model near the memory ceiling, not a throughput one for a model that isn't. Mechanism: under mmap the page cache holds a duplicate of the model (43 GiB observed resident) that competes with the KFD system-memory gate for the same host RAM. The kernel logs the refusal (`amdgpu: SVM mapping failed, exceeds resident system memory limit`) from roughly 2 minutes into the run; the userspace `BadAlloc` operators actually watch for does not surface for 55 to 72 minutes. Raising GTT from 84 to 122 GiB did not rescue the failing cell (run-0234, run-0235) — GTT and the KFD system-memory gate are separate accountants, so it was the wrong knob. `amdgpu.no_system_mem_limit=1` is inert once mmap is off (the postboot no-mmap reps behind run-0236/run-0237 ran with it set and were unaffected), and with mmap on it converts the fast, visible `BadAlloc` into an indefinite `svm_range_restore_work` hang with no watchdog-visible signal — measured directly in run-0242: a 900-second wall-clock bound was what ended that attempt, not the process, and the kernel log shows no `BadAlloc` at all, only the restore workqueue hogging the CPU with increasing frequency for the full window. Worse than the failure it was meant to fix. The lesson: a capability verdict is only as good as the harness's agreement with the serving configuration it claims to describe. The way to catch this defect class is to diff benchmark flags against production serving flags before trusting a categorical result — a check this lab had not run until this diagnosis.
- METHOD — the defect was found by diffing the flag sets of every queue script against the standing serving config. `docs/benchmark-runbook.md` and `bench/protocol.json` both name `--load-mode none` as mandatory on this unified-memory hardware, and it is present in every llama-server serving invocation and every tau2/capability queue script (e.g. `bench/queues/queue-e-nemotron-quant.sh`, cfg-0027/cfg-0028's serving flags). It was absent from the four llama-bench throughput queue scripts behind the perf-matrix and perf-matrix-gtt120 suites, which instead ran llama-bench's own default (mmap on). All four were rebuilt to carry `--load-mode none` after this finding. THE OOM, AS FIRST RECORDED — cfg-0057 (ROCm, f16 KV) and cfg-0058 (ROCm, q8_0 KV) both failed at d32768: run-0234 hit a `BadAlloc` 93 times over a 55-minute window before the watchdog terminated it, run-0235 hit it 5 times over a longer window before terminating. journalctl -k at the same timestamps shows the real cause — "amdgpu: SVM mapping failed, exceeds resident system memory limit" — recorded continuously from within minutes of process start in both cases, 44,862 lines total across the two cells. Raising the boot-time GTT ceiling from the fleet default (~84 GiB) to 122 GiB (`amdgpu.gttsize=122880`) and re-running did not rescue either cell — the failure tracks host-RAM residency, not the GTT window. THE FIX — re-running cfg-0057's arm with `--load-mode none` added and nothing else changed (cfg-0059, same binary, same box) produced a clean single-rep pass (248.788416 pp / 16.390674 tg at d32768, run-0243/run-0244) confirmed by a post-reboot 3-rep median (244.051835 pp / 16.434903 tg, run-0236/run-0237, stddev 3.76 pp / 0.03 tg) within 2% of the single check. run-0236/run-0237 is the fully-backed N=3 record and the one quoted as this cell's result. Against Vulkan's own run-0190/run-0191 at the same cell (139.810384 pp / 17.559139 tg, unaffected by the mmap question because it never hit the resident-memory ceiling ROCm did) that is ROCm +74.6% prefill / Vulkan +6.4% decode — see clm-0050's second amendment for the full reading of what that means for the per-model/per-phase rule. THE CONTROL — before generalising "mmap doesn't matter for throughput," it needed a check on a model that fits, because Nemotron-3-Super's cell only demonstrates that mmap matters for CORRECTNESS near the ceiling, not that it is throughput- neutral away from it. The check was built as a true matched pair: cfg-0032 (ROCm, Qwen3.6-35B-A3B, f16 KV, d32768, mmap default — the same arm behind clm-0050's Qwen3.6-35B numbers) was re-measured in the SAME boot session as its `--load-mode none` companion, immediately before it (run-0240/run-0241: 597.271098 pp / 42.41958 tg, median of 3), rather than reusing the two-days-earlier run-0116/run-0117 (595.763132 pp / 42.326925 tg) as the baseline — same-session removes timing/thermal/session drift as a possible explanation for the delta. cfg-0060 is the identical arm with `--load-mode none` added: run-0238/run-0239, 601.555743 pp / 42.368806 tg, median of 3 — +0.72% pp / −0.12% tg over the same-session mmap median (+0.72%/-0.12% again if measured against the two-days-earlier run-0116/run-0117 instead: the two mmap baselines agree within 0.25% of each other). Both readings are inside the ~3% scatter line the runbook already treats as noise for this suite. Every relative number this lab has published under mmap on a model that fits survives this finding unmodified — the corpus's models sit comfortably under the box's 122 GiB pool with the single measured exception of Nemotron-3-Super's ~82 GB weight footprint. What does not survive is any categorical "cannot run"/"cannot allocate" verdict reached under mmap on a model near that ceiling — those need a `--load-mode none` re-check before they can be trusted, and Nemotron-3-Super's was the one this lab had made. THE HANG VARIANT — a second postboot session added `amdgpu.no_system_mem_limit=1` to the kernel cmdline (cfg-0061) and re-ran both arms of the Nemotron cell. `--load-mode none` was unaffected (run-0236/run-0237, above) — that flag is inert once mmap is off. The mmap arm was not: run-0242 was killed by the queue's own 900-second bound (rc=124) with no `BadAlloc` ever surfacing; journalctl -k for the same window shows no "SVM mapping failed" either, only `workqueue: svm_range_restore_work [amdgpu] hogged CPU for >10000us` recurring with increasing frequency (4 times at +2m20s, 67 times by +15m, when the bound cut it off). Setting that flag while mmap stays on trades a failure an operator can at least detect for one a fixed-duration bound is the only thing that ends. KERNEL VS USERSPACE DETECTION LAG — the kernel-side "SVM mapping failed" message is legible from roughly 2 to 5 minutes into a doomed run; the userspace `BadAlloc` an operator would actually be watching a process exit code or stderr for does not land for 55 to 72 minutes (run-0234's window: 55 min; run-0235's: 72 min). `dmesg`'s ring buffer rotates out under the message flood well before that; only `journalctl -k` (the persistent journal) still has the onset by the time the userspace symptom finally shows. See `docs/methodology-lessons.md` for this as a general detection-lag rule, not just the one incident. CROSS-REFERENCE — this is the same failure class documented in `docs/aihydra-first-boot.md`'s "mmap double-residency" note and `docs/storage-and-lanes.md`'s `LLAMA_ARG_LOAD_MODE=none` rule and `bench/protocol.json`'s `--load-mode none on unified memory` entry, which records that the underlying defect "was hit FOUR separate times before being fixed properly" in serving code. This claim is the same defect class recurring a fifth time, in benchmark tooling rather than serving code, which is why it survived those four fixes untouched — they patched call sites that served requests, and none of them were a `llama-bench` invocation.
- evidence: run-0116 run-0117 run-0190 run-0191 run-0234 run-0235 run-0236 run-0237 run-0238 run-0239 run-0240 run-0241 run-0242 run-0243 run-0244
clm-0054
- measured-here med ●●○ volatility medium · verified 2026-08-15
- On gfx1151 at f16 KV, Qwen3.8-27B (dense) is a fourth model measured under the per-model, per-phase backend rule (clm-0050), and it lands emphatically on the prefill-favours-ROCm side: DECODE is backend-independent on this model (every matched cell within 5%), but Vulkan PREFILL collapses with depth — 0.55x ROCm at d32768 on Q8_0 (87.82 vs 159.98 t/s) and 0.48x on UD-Q4_K_XL (99.80 vs 207.61) — and at d131072 stock Vulkan cannot complete the cell at all: 2 of 2 reps in BOTH quants aborted with vk::DeviceLostError, the kernel logging an amdgpu ring timeout and recovering the device by ring reset, while ROCm completed every cell it was offered (83.44 pp / 6.11 tg Q8_0, 95.59 pp / 8.14 tg UD-Q4_K_XL at d131072). This REVERSES the screen's own shallow-data backend pick: a decode-only d0 probe chose Vulkan (+0.8% Q8_0, +5.0% UD-Q4_K_XL), and depth then showed prefill swinging to ROCm by +82.2% (Q8_0) and +108.0% (UD-Q4_K_XL) at d32768 and Vulkan failing outright where ROCm ran. For this model on this chip, ROCm is the only backend that holds to 131k, and the d0 pick was wrong in the way clm-0050 predicts shallow picks to be wrong: backend advantage is per-model AND per-phase, and it moves with depth.
- METHOD — qwen38-screen phase B: stock llama.cpp 3653e6d, one build per backend arm, f16 KV, fa on, -ngl 999, --load-mode none (clm-0053's rule, carried by the queue from its first rep), pp1024/tg256, llama-bench defaults otherwise (-b 2048 / -ub 512), median of 3 fresh-process reps with page cache dropped between reps. Depth-major order, d131072 last. Widest rep spread in any completed cell: 3.08% pp (vulkan Q8_0 d32768); every other cell under 1%. pp1024 / tg256 by depth: | quant | depth | ROCm | Vulkan | vk/rocm pp | vk/rocm tg | |---|---|---|---|---|---| | Q8_0 | 0 | 226.22 / 7.81 | 249.63 / 7.86 | 1.103 | 1.006 | | Q8_0 | 32,768 | 159.98 / 7.29 | 87.82 / 7.27 | 0.549 | 0.998 | | Q8_0 | 131,072 | 83.44 / 6.11 | DEVICE LOST | — | — | | UD-Q4_K_XL | 0 | 346.49 / 11.44 | 361.69 / 12.03 | 1.044 | 1.052 | | UD-Q4_K_XL | 32,768 | 207.61 / 10.36 | 99.80 / 10.69 | 0.481 | 1.031 | | UD-Q4_K_XL | 131,072 | 95.59 / 8.14 | DEVICE LOST | — | — | THE DEVICE-LOST CELLS are recorded as failed runs with the process and kernel evidence verbatim (run-0265, run-0266), the same discipline as the OOM cells of the Nemotron matrix (run-0234/run-0235): a cell that failed is a result, and no invented number stands in for it. The failure is recoverable — the kernel reset the compute ring each time ("Ring comp_1.2.0 reset succeeded ... device wedged, but recovered through reset") and the box needed no reboot — but four out of four attempted reps across two quants at d131072 is deterministic enough for the queue's fast-fail rule. This is the first in-lab reproduction of the failure class llama.cpp issue 25664 documents for Vulkan/RADV at deep context on this hardware class, which clm-0050 has carried as a reliability counterweight since 2026-08-13: the Qwen3.6-35B matrix ran Vulkan clean to 131k, this model does not. SCOPE — the d204800 cell was not run in this screen (the queue's depth list ended at 131072); whether ROCm's prefill lead extends there is measured on other models but only extrapolated here. The decode parity is this model's OWN result: on Qwen3.6-35B Vulkan leads decode by 19-21% at every depth, on gpt-oss-120b by 10-16%, on this model by at most 5.2% at d0 and effectively nothing at depth — another way the rule is per-model. The backend verdict for serving this model at useful context is ROCm on both phases: prefill by a wide margin, decode by indifference, and completion-at-depth by necessity. The community's first-throughput rows for this model (r/StrixHalo, Vulkan fork b10283, amd_iommu=off, pp512/tg128, 2 reps) are adjacent evidence, not the same measurement — different build lineage, IOMMU state, and phase lengths; their own A/B puts amd_iommu=off alone at up to +34-38% on dense prefill. No public ROCm figures for this model on this silicon existed before this matrix.
- evidence: run-0245 run-0246 run-0247 run-0248 run-0249 run-0250 run-0251 run-0252 run-0253 run-0254 run-0255 run-0256 run-0257 run-0258 run-0259 run-0260 run-0261 run-0262 run-0263 run-0264 run-0265 run-0266
clm-0055
- measured-here med ●●○ volatility medium · verified 2026-08-15
- draft-mtp speculation on Qwen3.8-27B (stock 3653e6d, gfx1151) peaks at spec-draft-n-max=3 over CLEAN cells — Vulkan Q8_0 17.75 t/s (2.26x its 7.86 no-speculation floor, acceptance 0.626), Vulkan UD-Q4_K_XL 26.68 t/s (2.23x, 0.623), ROCm Q8_0 18.25 t/s (2.33x, 0.6235) — and at n_max >= 4 the feature is BROKEN on this model: after accumulated generation volume in a live session (sequential VARIED prompts at full length; not fresh servers, not one repeated prompt, not short generations), generations start terminating at 1 token with <|im_end|> (id 248046). The decisive measurement is that the TARGET model's own pre-sampling distribution puts im_end at logprob -0.084 (~92%) on a prompt the same server answers normally with speculation off — the speculative path is corrupting the target's forward pass, not merely mis-accepting drafts — and ignore_eos:true restores both correct text AND draft accounting. The hazard compounds into a measurement artifact: llama.cpp reports a 1-token generation as 1,000,000 tokens/s, so unfiltered throughput averages FLATTER exactly the broken cells; a community sweep on this silicon reporting a peak at n_max=5 sits on what is measured here as one of the two worst cells (11/15 degenerate on Q8_0), under a workload shape (single repeated codegen prompt) that this trigger analysis shows cannot reproduce the failure. Recommendation from measurement: draft-mtp ON, n_max hard-capped at 3 until the mechanism is understood upstream.
- METHOD (sweep) — llama-server harness (llama-bench has no speculative support; spec params are force-set to 0 there), 5 varied prompts x 3 reps = 15 generations per cell, greedy (temperature 0, seed 42), n_predict 500, cache_prompt off, -c 32768, --load-mode none, stock 3653e6d, IOMMU on. The floor (spec off) travels in-session with every cell; the RATIO is the deliverable, absolutes are not comparable to llama-bench numbers. "degenerate" = predicted_n <= 1, and the per-cell degenerate fraction is part of the record: Vulkan sweep, decode t/s median of clean samples (degenerate/attempted): | n_max | Q8_0 | degen | UD-Q4_K_XL | degen | |---|---|---|---|---| | off | 7.86 | 0/15 | 11.98 | 0/15 | | 2 | 16.50 | 0/15 | 23.74 | 0/15 | | 3 | 17.75 | 0/15 | 26.68 | 0/15 | | 4 | (artifact) | 12/15 | 25.55 | 5/15 | | 5 | (artifact) | 11/15 | 24.73 | 7/15 | | 6 | 34.41 | 3/15 | 43.72 | 2/15 | The Q8_0 n_max=4/5 "medians" computed naively are 1,000,000 t/s — the 1-token sentinel value, not a throughput; that is the artifact hazard in one line. The n_max=6 cells are faster on their clean samples and markedly LESS degenerate than 4/5 — unexplained, and left as an open question rather than a recommendation. The clean-cell peak is n_max=3 on every arm measured. The n_max >= 4 cells are deliberately NOT recorded as run records; the clean cells behind this claim are (run-0268 ROCm Q8_0, run-0269 Vulkan Q8_0, run-0270 Vulkan UD-Q4_K_XL). THE EOS CLIFF, characterised (full evidence bundle: ~/bench-results/qwen38-screen/eos-cliff-evidence.md on the bench box, prepared for upstream review): - Trigger needs accumulated generation volume in one server session: fresh server per request 0/6 degenerate; one repeated prompt sequentially 0/10; varied prompts at n_predict 48 (~480 tokens total) 0/10; varied prompts at n_predict 500 hit the cliff by request 2-3 (~1,400 tokens) — 11/15 degenerate at n_max=5. Recovery is possible mid-session (a later request produced a clean 500), so the state is not latched. - Affected requests show tokens_cached == tokens_evaluated (the KV slot WAS reset — not stale prompt cache) and draft_n/draft_n_accepted ABSENT entirely, vs 666-847 drafted on healthy requests in the same session. - Not a sampler artifact: reproduces under greedy (seeds 42 and 1234, bit-identical sequences) and under the vendor thinking sampler (temp 1.0, top_p 0.95, top_k 20, min_p 0) at seeds 42 and 7 with seed-varying rates. Bit-identical greedy failure sequences across four independent sessions, across server restarts and drop_caches. - Worse on Q8_0 (12/15 at n_max=4) than UD-Q4_K_XL (5/15) — the OPPOSITE of what a quantisation-noise story predicts. - ignore_eos:true control: same server, same n_max, same prompt — correct text returns AND draft accounting reappears (draft_n 341, accepted 129). The corruption expresses entirely through the EOG token. - CHECKED against the nearest upstream issues; none matches this signature: 23302/23335 are draft-mtp token DIVERGENCE with complete generations (Metal); 25618 is quant-dependent divergence (and the Q8-worse observation here cuts against it); 26750 is a CUDA acceptance collapse — HIP acceptance here (0.6235) is statistically indistinguishable from Vulkan (0.6261), so it does not generalise to this box. CONFIDENCE NOTES — what is and is not established. Established by measurement: the trigger conditions, the n_max threshold, the pre-sampling distribution shift onto im_end, the ignore_eos control, backend-indifferent acceptance at n_max=3, and determinism. NOT established: the root cause (the MTP head's internal state is the natural suspect but was not instrumented), why a single repeated prompt escapes, and why n_max=6 partially recovers. ROCm at n_max >= 4 was not probed for the symptom (the clean ROCm n_max=3 cell is not evidence either way) — see the 2026-08-15 amendment below, which is the first ROCm probe at n_max >= 4. UPSTREAM: not yet filed — the evidence bundle is written for review first; this claim should be re-verified against whatever the eventual issue thread establishes. Confidence is medium on the characterisation and the n_max<=3 operating rule; the mechanism language above is deliberately interpretation-free. COMMUNITY CONFLICT, stated fairly: the r/StrixHalo sweep reporting draft-mtp peaking at n_max=5 for this model family ran a different build lineage (the strix-halo-vulkan fork, b10283/b10397 era), amd_iommu=off, and a single repeated 500-token codegen prompt — the exact workload shape measured here as unable to trigger the cliff — and llama-bench-style unfiltered averages cannot distinguish a fast cell from a broken one once 1-token generations enter the mean. Their numbers may be correct for their harness; the recommendation that follows from them is what this claim disputes for live serving. Guard evidence for the recommended operating point: 25 sequential full-length vendor-sampler generations at n_max=3 on the tau2 serving config, 0/25 degenerate, ~11,000 generated tokens in one session (run-0267) — deeper session volume than any failing cell needed. ENERGY (retroactive join, computed 2026-08-15 against the same nmaxoff/nmax3 sub-windows behind the sweep table above; eng-0077/0078/0079/0080). MTP's decode speedup does NOT translate 1:1 into energy savings, because mean power draw rises under speculation rather than holding constant — the extra draft/verify compute costs watts, not just wall-clock: | arm | floor Wh/1000tok | n_max=3 Wh/1000tok | energy improvement | decode speedup | mean W floor -> n3 | |---|---|---|---|---|---| | Vulkan Q8_0 | 5.03 | 2.42 | 2.08x | 2.26x | 142.4 -> 154.7 (+8.6%) | | Vulkan UD-Q4_K_XL | 3.41 | 1.88 | 1.82x | 2.23x | 147.2 -> 180.5 (+22.6%) | | ROCm Q8_0 | 5.21 | 2.70 | 1.93x | 2.33x | 146.8 -> 177.6 (+21.0%) | Every arm still burns meaningfully less energy per output token with speculation on — this is not a wash — but the win undershoots the raw decode ratio by 8-23 percentage points depending on backend/quant, entirely because power draw is not flat. UD-Q4_K_XL shows the largest power increase and therefore the smallest energy win despite a comparable decode speedup, which is the opposite of what "MTP is free, only wall-clock changes" would predict. Wh/1000tok = wh_total / (decode_tps * window_seconds) — an approximation that assumes the window is generation-dominated (verified true: prompts are ~20 tokens, prefill negligible against 500-token generations). Recommendation unchanged (n_max=3, hard-capped) — the energy case reinforces rather than revises it — but "2.3x faster" should not be read as "2.3x less energy"; it is closer to 1.8-2.1x less energy, per this measurement. AMENDMENT 2026-08-15 — EOS-CLIFF x q8_0 KV, and the first ROCm probe at n_max >= 4. Prompted by a r/Qwen_AI thread (1vorjo7) recommending n_max=7 + q8_0 KV for this model, and by upstream #25618 (a q8_0 V-cache + MTP divergence interaction on a DIFFERENT model) — two open questions this claim had not addressed: does the cliff reproduce under quantised KV, and is our own recommended combo (n_max=3, q8_0 KV — clm-0057/clm-0058's serving config) free of that interaction. Same 5-prompt-sequence x3-reps harness as the sweep above, greedy seed 42, n_predict 500, cache_prompt off, ROCm, UD-Q4_K_XL weight quant, `-ctk q8_0 -ctv q8_0`: | n_max | KV | degenerate | evidence | |---|---|---|---| | 3 | q8_0 | 0/15 | run-0275 | | 7 | q8_0 | 0/15 | run-0276 | Q1 ANSWERED FOR OUR SERVING COMBO: n_max=3 stays clean under q8_0 KV (run-0275), same as it already was under f16 KV (run-0267's 0/25 at the tau2 serving config). The recommended Warden combination is not exposed to this interaction by this test. Q2 — DOES NOT REPRODUCE, BUT THE TEST IS CONFOUNDED. n_max=7 under q8_0 KV on ROCm produced zero degenerate generations (run-0276) — no im_end-as-first-token signature anywhere in the 15-request sequence, at an n_max deep inside the zone that broke Vulkan/f16 KV (12/15 and 11/15 at n_max=4/5). But this cell changes TWO variables from the original repro at once: backend (ROCm, not Vulkan) AND KV type (q8_0, not f16). It therefore does NOT cleanly show that q8_0 KV suppresses the cliff — it is equally consistent with ROCm simply not exhibiting this symptom at any KV type, which is exactly the gap this claim already flagged ("ROCm at n_max >= 4 was not probed for the symptom"). This is the first data point on that gap, and it is clean, but it does not close it: a matched ROCm + f16-KV + n_max=7 control is the cell that would decompose backend from KV-quant as the protective factor, and it has not been run. PRACTICAL READING: this does not change the standing recommendation (n_max=3, hard-capped) — that recommendation was never contingent on n_max=7 being safe or unsafe, and the clean n_max=3 result under q8_0 KV is confirmatory, not novel. What it does do is narrow the open backend question without closing it, and it gives a direct, sourced answer to the r/Qwen_AI thread's n_max=7+q8_0 suggestion for THIS box's ROCm build: no cliff observed in 15 requests, but the test that would show whether that is the KV type or the backend doing the protecting has not been run. AMENDMENT 2026-08-15 (SECOND) — CAUSE-ISOLATION MATRIX: backend is the dominant single-variable gate, session accumulation is a real secondary factor, closes the "matched ROCm + f16-KV control has not been run" gap from the amendment above. Full method, raw matrix, and reasoning: eos-cliff-evidence.md section 11 (bench box + local project copy) and ~/bench-results/eos-cliff-isolation-20260815/matrix-summary.md on the box. One variable changed per cell from the failing baseline (Vulkan, Q8_0, f16 KV, n_max=5), 15 requests each (same 5-prompt x3-rep greedy harness as the sweep above), plus 45-request extended confirmation of the two clean results, plus four avoidance probes on the still-failing config: | cell | backend | quant | KV | n_max | n | degenerate | short-stop | |---|---|---|---|---|---|---|---| | 0 baseline (positive control) | Vulkan | Q8_0 | f16 | 5 | 15 | 11/15 | 1/15 | | B backend-only flip | ROCm | Q8_0 | f16 | 5 | 15 | 0/15 | 0/15 | | C KV-only flip | Vulkan | Q8_0 | q8_0 | 5 | 15 | 0/15 | 1/15 | | E new-evidence combo at n5 | ROCm | UD-Q4_K_XL | q8_0 | 5 | 15 | 0/15 | 3/15 | | B-ext (45 req) | ROCm | Q8_0 | f16 | 5 | 45 | 0/45 | 0/45 | | N7-ext (45 req) | ROCm | UD-Q4_K_XL | q8_0 | 7 | 45 | 0/45 | 9/45 | | avoid-1 n_max=3 | Vulkan | Q8_0 | f16 | 3 | 15 | 0/15 | 0/15 | | avoid-2 p-min 0.9 | Vulkan | Q8_0 | f16 | 5 | 15 | 1/15 | 0/15 | | avoid-3 cache_prompt on | Vulkan | Q8_0 | f16 | 5 | 15 | 12/15 | 1/15 | | avoid-4 restart every 3 | Vulkan | Q8_0 | f16 | 5 | 15 (5 sessions) | 1/15 | 3/15 | "degenerate"=predicted_n<=1 (this claim's established definition); "short-stop" = any other premature stop, the same "intermediate case" already present in the original characterisation's minimal repro (394/500). Evidence: run-0277 (cell 0), run-0278 (cell B), run-0279 (cell C), run-0280 (cell E), run-0281 (B-ext), run-0282 (N7-ext), run-0283 (avoid-1), run-0284 (avoid-2), run-0285 (avoid-3), run-0286 (avoid-4). LOCALISATION: backend (ROCm vs Vulkan) is the dominant single-variable gate at this session depth. Flipping it alone took 11/15 degenerate to ZERO anomalies of any kind at 3x the request volume (B-ext). KV-only and quant-only flips both leave the model still perturbed (KV-only: 1 short-stop; quant-only, already in this claim's own sweep table: 7/15 degenerate). The "new evidence" triple-flip combination, re-tested at the actually-diagnostic n_max=5 instead of the untested-in-the-failure-zone n_max 3/7, shows a HIGHER short-stop rate (3/15) than the KV-only flip — combining the backend flip with quant+KV changes does not add protection over the backend flip alone, and may reintroduce a partial echo of the mechanism. This is a mitigation gradient, not a clean/broken binary: even N7-ext (fully clean of true collapse) shows a small, perfectly periodic short-stop tied to one prompt position, recurring every cycle, never escalating with volume. SESSION ACCUMULATION IS REAL BUT PARTIAL. Restarting the server every 3 requests on the FAILING Vulkan config cut the degenerate rate 11/15 -> 1/15 — the single largest mitigation measured — but one fresh 3-request sub-session still produced a full collapse on its 3rd request; a targeted pre-sampling replay of that exact case (post_sampling_probs:false, this claim's decisive- measurement method) shows im_end at logprob -0.148 (~86.2%) vs the #2 candidate at -2.115 (~12.1%) — as strong a shift as the original 92% characterisation. Raising --spec-draft-p-min to 0.9 shows the same profile (11/15 -> 1/15, not zero); when it fires, im_end sits at -0.391 (~67.6%) vs -1.175 (~30.9%) — weaker but the same direction. cache_prompt:true is NOT a mitigation — it made the baseline WORSE (12/15 vs 11/15). REPETITION-LOOP CHECK (motivated by llama.cpp #26425 "MTP retains inter-request state... model degradation" and #23577 "Qwen3.6-27B MTP outputs repeated //// after long session" — near-miss reports sharing the accumulated-session-state trigger class but a different reported symptom). All 26 full-raw JSON captures from this matrix were grepped for repeated- character (8+) and repeated-short-token (7+) runs. Zero genuine hits — every match was a legitimate markdown separator inside a correct, complete answer. This box's failure signature stays exclusively in the EOS-collapse family; it never drifts into a repetition-attractor mode under any configuration tested. AVOIDANCE ENVELOPE: --spec-draft-n-max<=3 remains safe regardless of backend (avoid-1 reconfirms 0/15 on the original failing Vulkan combo) and is still the only zero-anomaly result achieved on Vulkan without changing backend, quant, or KV. ROCm held clean at 3x this matrix's depth (45 req / ~22,500 generated tokens) at n_max=5 — the strongest result in the matrix — but this scope is 15-45 requests, NOT production-scale session depths; whether it holds at hundreds-to-thousands of requests is not established and should not be assumed. On Vulkan at n_max>=4-5, no cheap mitigation measured here reaches zero: restart-every-3 and p-min=0.9 both cut the rate ~10x without eliminating it; cache_prompt:true is counterproductive. WHAT THIS DOES NOT ESTABLISH: the mechanism itself is still uninstrumented — this matrix localises WHICH variable gates the symptom, not WHY. "ROCm doesn't show it here" is not the same claim as "ROCm is immune," and this claim should not be read as making the stronger one. Standing recommendation UNCHANGED: n_max=3, hard-capped, on any backend — that recommendation was never contingent on this result and remains the only zero-anomaly, any-backend, any-depth-tested operating point measured to date.
- evidence: run-0267 run-0268 run-0269 run-0270 run-0275 run-0276 run-0277 run-0278 run-0279 run-0280 run-0281 run-0282 run-0283 run-0284 run-0285 run-0286
clm-0056
- measured-here med ●●○ volatility low · verified 2026-08-15
- Qwen3.8-27B's KV cache on llama.cpp costs exactly 64.00 KiB per token — the hybrid-attention allocation working as designed, measured byte-exact on this box: only 16 of the 64 layers carry a conventional KV cache (the 3:1 linear-to-full-attention layout), and 16 layers x 4 KV heads x 256 head_dim x 2 (K+V) x 2 bytes (f16) = 65,536 bytes/token. The community-reported 8.1 GiB KV footprint at 32k context for this model is 64-layer arithmetic and is wrong as a llama.cpp planning number: the measured allocation at 32k is 2.00 GiB. Full serving footprints measured at 262,144 context (f16 KV, MTP head not loaded): Q8_0 42.2 GiB, UD-Q4_K_XL 32.6 GiB — either quant fits a 120 GiB GTT window at FULL declared context with more than 75 GiB to spare, so context length is not a fit constraint for this model on 128 GB hardware. Loading the MTP head for speculation adds 430.4 MiB (Q8_0) or 272.6 MiB (UD-Q4_K_XL) of weights.
- METHOD — qwen38-screen phase A: llama-server load-only probes, both backends, both quants, n_ctx 4096 / 32768 / 262144, GTT usage read from the amdgpu telemetry before and after each load (gtt_used delta), server torn down between probes. Raw records: aihydra ~/bench-results/qwen38-screen/phaseA/footprint.jsonl (12 probes, all loaded clean) — measured footprints, not runs in the corpus sense, because a load probe produces no throughput or capability metric; the evidence lives in the phase A files and the arithmetic below is checkable from them. THE PER-TOKEN NUMBER, derived two independent ways that agree exactly: - Measured: ROCm GTT delta between c4096 and c262144 is 16,128 MiB for BOTH quants (Q8_0 27,118.5 -> 43,246.5 MiB; UD-Q4_K_XL 17,274.5 -> 33,402.5 MiB) over a 258,048-token context difference = 0.0625 MiB = 64.00 KiB/token, exact to the reported MiB. (Vulkan totals run ~0.4-0.8 GiB higher at large context — compute-buffer overhead, not KV; the ROCm deltas are the clean read.) - Arithmetic: the GGUF declares 64 layers, but only the 16 full-attention layers allocate KV (llama.cpp implements the hybrid layout for this arch); 16 x 4 kv_heads x 256 head_dim x 2 (K and V) x 2 bytes = 65,536 B/token. At 32,768 tokens: 2.00 GiB. At 262,144: 16.0 GiB. THE COMMUNITY CORRECTION — the 8.1 GiB-at-32k figure circulating for this model (r/StrixHalo, same thread as the first Vulkan throughput rows) matches allocating a full-attention cache across ALL 64 layers (~8.6 GiB by the same arithmetic, within reporting error of the 8.1 read). Whatever produced that reading, current llama.cpp does NOT allocate it: the hybrid layout is honoured and the planning number for fit maths is 64 KiB/token, 4x smaller. This was flagged as an open instrumentation question on the candidate record before the screen ran, and the screen settled it on the first phase. FIT — at the full declared 262,144 context: Q8_0 total 42.2 GiB (43,246.5 MiB GTT), UD-Q4_K_XL 32.6 GiB (33,402.5 MiB) against the box's 120 GiB GTT window (gttsize=122880). Weights alone: ~27.1 GiB Q8_0, ~17.3 GiB UD-Q4_K_XL (c4096 loads, which carry only 0.25 GiB of KV). The MTP-head weight deltas are from the GGUF tensor tables (mtp_num_hidden_layers 1: 430.4 MiB at Q8_0, 272.6 MiB at UD-Q4_K_XL); phase A probed serving loads without speculation, so those tensors were not resident in the probes above.
clm-0057
- measured-here med ●●○ volatility medium · verified 2026-08-15
- Qwen3.8-27B met its registered agentic prediction: on the standing 5-task tau2 airline smoke subset it scored 1.000 (5/5, 21 tool-call messages, 0 empty assistant turns, VALID under the smoke gate) against the >= 0.80 bar recorded in the candidate record on 2026-08-14, BEFORE any measurement — and against the incumbent qwen36-27b-mtp's 0.80 on the identical tasks under the identical pinned-simulator protocol, where the incumbent failed task 2 and this model passed it. At n=5 one task is worth 0.20: that is a margin met, not a significance claim. The full run behind it: mean 0.577 over an effective n=26 of the 50-task set (the run hit its 11,400 s wall bound at rc=124; all 26 completed tasks scored, 165 tool-call messages, 0 empty turns, 0 max-steps cuts), on ROCm Q8_0 with draft-mtp n_max=3, reasoning_effort pinned to medium (positive-control verified on the baked template), agent max_tokens 4096, pinned haiku-4.5 simulator, seed 42. The first five tasks are the easy end of the set — this model's own full-run mean is 0.577 against 1.000 on them — which is exactly why smoke and full are reported as separate numbers rather than blended. The max_tokens 4096 output cap is a fingerprint deviation: the historical tau2 series ran uncapped, so 0.577 is not directly comparable to those means.
- METHOD — tau2-bench airline, harness 668d3bc, --num-trials 1 --seed 42 --max-steps 200 --max-concurrency 1, user simulator pinned per protocol (openrouter/anthropic/claude-haiku-4.5, temperature 0, Anthropic-only routing). Agent: local llama-server endpoint on cfg-0066 (ROCm Q8_0, draft-mtp n_max=3 under the clm-0055 hard cap, -c 32768, --parallel 1, --load-mode none), agent sampling temperature 1.0 / top_p 0.95 / max_tokens 4096, and reasoning_effort pinned to medium via --chat-template-kwargs — llama.cpp discards the OpenAI reasoning_effort request field (PR 26941 pending), so the template kwarg is the only route that reaches every request, and the queue PROVED it lands before any task ran (xhigh and medium renders differ; the medium render carries no xhigh; run-0267's comment). Q8_0 is outside protocol.json quants.expected; the override is recorded in run-meta per the protocol's override discipline (Q8_0 is the candidate record's designated primary screening quant, chosen to remove the quant confound from a capability screen). THE PREDICTION — recorded on the candidate yaml 2026-08-14, before release-day measurement: "Qwen3.8-27B meets or beats 0.80 on the same smoke" (the vendor's headline claims are all agentic, which is why the smoke was pre-registered as the independent check). Outcome: MET, 1.000 vs the bar of 0.80. The incumbent comparison is same-tasks, same-protocol, same-simulator (qwen36-27b-mtp re-screen of 2026-08-13, recorded in that candidate's history): 0.80, failing task 2 — the task this model completed in 576 s, its longest smoke task. THE FULL RUN — 26 of 50 tasks inside the hard 11,400 s bound (tasks arrive in id order, so the effective set is tasks 0-25, not a random draw; per-task rewards, turns and wall times are in run-0271's trace). Passed 15/26. Median turns 24, p90 36, max 44, none cut at max_steps. Wall time per task ranged 50 s to 1,176 s. SMOKE VALIDITY GATE (protocol, added 2026-08-14): tool_call_messages recorded alongside both means — 21 (first five) and 165 (all 26), zero empty assistant turns, zero infrastructure errors, so neither number is a do-nothing artifact. CAVEATS — (1) the output cap: max_tokens 4096 bounds each assistant turn; historical uncapped arms allowed longer reasoning, so cross-series comparison of the 0.577 must carry this asterisk (the smoke-vs-incumbent comparison is NOT affected the same way — the incumbent smoke ran under its own screen-tier settings, and both are n=5 cliff-detectors, not rankings). (2) Effective-n truncation: 26 of 50 by wall bound is a completed-prefix comparison in clm-0037's sense if set against a full-50 mean; compare only matched prefixes. (3) Speed is the cost of the capability: ~18 t/s decode under n_max=3 speculation (clm-0055) put the median passing task at several minutes of wall time — the capability verdict and the latency verdict are separate findings and the candidate record carries both. ENERGY (retroactive join, computed 2026-08-15, eng-0081) — the wh_total for the full run's window is 574.4 Wh (mean 181.4 W, delta 171.3 W over the 10.1 W box-idle floor), across 15 correct of 26 scored: **38.29 Wh per correct answer**, 1.16 pence at 30.3 p/kWh tariff. This is WHOLE-SESSION energy — the window runs start-to-finish across the arm including simulator (user_llm) wait time and a ~19-minute gap between task 9 and task 10 — not an agent-only or decode-only figure, so it should not be compared directly to a wh_per_task computed from a decode-only bench cell. The published bench went out on 2026-08-14 with no energy join at all; this closes that gap the day after.
- evidence: run-0267 run-0271
clm-0058
- measured-here med ●●○ volatility medium · verified 2026-08-15
- clm-0035's retracted quality-drop story does not reappear at n=14 on the PATCHED (ce7689f) build: task-matched against an f16 control, quantised KV on Qwen3.6-35B-A3B-UD-Q4_K_XL scores q8_0 0.857 against f16 0.786 — q8_0 AHEAD, not behind. 13 of the 14 matched tasks scored identically in both arms; the single exception is the one f16 failed and q8_0 passed. Combined with clm-0022's speed result, quantised KV on the patched build costs nothing measurable in speed or quality on the models measured so far.
- METHOD — hard-split, task-matched tau2 airline arms on the same server binary (`~/src/llama.cpp-kvfix/build/bin`, commit ce7689f — the KV-dequant cherry-pick behind clm-0022's speed result), differing only in `-ctk`/`-ctv` (f16 vs q8_0). Qwen3.6-35B-A3B-UD-Q4_K_XL, ROCm 7.1.0, `-fa on -c 32768 --parallel 1 --load-mode none`, pinned Haiku-4.5 simulator (openrouter/anthropic, temperature 0), seed 42. `bench/queue-kv-quality-patched.sh`'s design: the run window is split HARD in half up front, f16 runs first against a 50-task target list capped at its slice, and whatever it completed and scored (14 of 16 sims — tasks 8 and 14 hit `infrastructure_error`) becomes the EXACT task-id list q8_0 then runs. Both arms are task-matched by construction, not post-hoc intersection — cfg-0070/run-0272 (f16), cfg-0071/run-0273 (q8_0). | arm | scored | mean_reward | paired_mean | tool_call_msgs | infra_err | verdict | |---|---|---|---|---|---|---| | f16 | 14/16 | 0.786 | 0.786 | 110 | 2 (tasks 8, 14) | VALID | | q8_0 | 14/14 | 0.857 | 0.857 | 83 | 0 | VALID | Both arms pass the SMOKE validity gate (protocol.json): nonzero tool-call messages, zero empty assistant turns, so neither mean is a do-nothing artifact (armcompare.json, aihydra `~/bench-results/kv-quality-patched-20260815T052214Z/`). **13 of the 14 matched tasks scored identically across arms** (task 15 failed in both; task 7 failed in both; the other 11 passed in both). The sole exception is **task 4**: f16 scored 0.0 (`user_stop`, 9 tool calls, 176.5s), q8_0 scored 1.0 (`user_stop`, 7 tool calls, 45.7s) — q8_0 passed the task f16 failed, not the reverse. That single task is the entire gap between the two paired means (0.857 − 0.786 = 0.071, exactly 1/14). **THIS IS THE OTHER HALF OF clm-0022, AND IT REVERSES clm-0035's DIRECTION.** clm-0022 established the dequant patch restores q8_0 KV *speed* to parity with (and past) f16, on the STOCK-vs-patched 122B comparison. clm-0035 measured a quality cost for q8_0 on the STOCK build (5-task tau2, f16 1.00 vs q8_0 0.60) — a finding clm-0036 then retracted for being fully inside this harness's ~0.40 run-to-run noise band at n=5. Nobody had run the quality question on the PATCHED build until now. At n=14, patched q8_0 does not merely avoid costing quality — it is nominally ahead of its own f16 control, by a margin the size of one task. **Combined with clm-0022, the reading is: on the patched build, quantised KV costs nothing measurable in either speed or quality — on the models measured so far.** CAVEATS, stated plainly: - **n=14, one task's worth of the whole result.** A single matched task is 0.071 of the paired mean — the entire f16/q8_0 gap here is one task's outcome, not a distributed trend. This is not the n=5/noise-band problem clm-0036 diagnosed (14 is closer to a decision than 5 is), but it is still a single trial per task, not a repeated-arm measurement, and should be read as suggestive rather than dispositive on its own. - **Different subject than clm-0035, deliberately.** clm-0035's retracted quality-drop story was measured on Qwen3.5-122B-A10B; this run substitutes Qwen3.6-35B-A3B for time (the box was needed for a release window, `bench/queue-kv-quality-patched.sh`'s header). This is a NEW measurement on a smaller model, not a replication of clm-0035 on the patched build — the 122B's own patched-build KV quality is still unmeasured. - **max_tokens-capped, not comparable to the historical uncapped tau2 series.** Both arms ran with `max_tokens: 4096` (tau2 `--agent-llm-args`, mirrored by `-n 4096` server-side) after an earlier attempt on this same suite ran an agent turn unbounded into a stuck `<think>` block and stalled the queue. The cap is identical across both arms here, so the f16-vs-q8_0 comparison itself is unaffected, but neither arm's raw mean is directly comparable to clm-0035/clm-0036's uncapped numbers or to the historical tau2 series generally. - **f16's two infrastructure_errors are the known empty-content harness quirk, not a KV effect.** Tasks 8 and 14 both failed with zero tool calls and zero duration — the same harness-vs-serving-config discrepancy class clm-0053 documents (trust the harness only as far as it agrees with the serving configuration it claims to describe). They are excluded from both `tasks_total` and the paired comparison, not counted against f16's score. ENERGY (retroactive join, computed 2026-08-15, eng-0082/eng-0083 — neither arm had an energy join at the time this claim was first written). Wh-per-correct- answer over the identical 14 task-matched ids: f16 14.88 Wh/correct (163.696 Wh total, 11/14 passed), q8_0 8.66 Wh/correct (103.874 Wh total, 12/14 passed) — a 42% reduction. This is not merely the quality-neutral result clm-0058's `text` reports: q8_0's smaller KV cache also cut wall time nearly in half over the same tasks (2755s vs 4303s, window-for-window), so it costs less energy AND scores higher on this pair, not a trade between them. Both figures are WHOLE-SESSION energy (simulator wait time included, not agent-only decode), same caveat as every other tau2 arm's Wh-per-correct on this site. CROSS-REFERENCE — clm-0022 (the speed result this claim's other half rests on, measured on the 122B); clm-0035 (the retracted stock-build quality-drop story this claim answers, on the same subject clm-0022 used); clm-0036 (the noise characterisation that retracted clm-0035, and the reason n=14 rather than n=5 was worth the box time here).
- evidence: run-0272 run-0273
clm-0059 superseded
- measured-here med ●●○ volatility medium · verified 2026-08-15
- On gfx1151 at f16 KV, DeepSeek-V4-Flash-0731 UD-IQ3_XXS is a fifth model measured under the per-model, per-phase backend rule (clm-0050), and it lands on the same side as Qwen3.8-27B (clm-0054): ROCm swept the full throughput matrix clean (d0 through d262144) while Vulkan lost the GPU device on both allowed attempts at every depth >=32768 (rc=134/SIGABRT, vk::DeviceLostError; kernel evidence: amdgpu ring timeout, ring reset, "device wedged, but recovered through reset" — 12 reset cycles total across the four failed cells). Even at the one depth Vulkan completed, d0, ROCm decode was already ahead (15.30 vs 12.44 t/s, +23%) while prefill was close (140.80 vs 134.73, ROCm +4.5%); ROCm's prefill lead widens with depth on every model measured so far on this chip. Serving this model at any useful context on this box requires ROCm — Vulkan cannot be trusted to survive a conversation that grows past 32k.
- METHOD — deepseek-v4-flash-fullbench matrix: stock llama.cpp 3653e6d (ROCm) / 3653e6d6d (Vulkan, prefix-matches per house convention), f16 KV mandatory (candidate's COMMUNITY HAZARD note — quantised K trips an incoherence-rotation bug on this model), fa on, -ngl 999, --load-mode none, pp1024/tg256, llama-bench defaults otherwise (-b 2048/-ub 512). N=3 fresh-process reps at d0/d32768 (stddev <=0.6 t/s on every cell, well inside the 3% scatter-flag threshold — a healthy code path per protocol §0), N=1 above per protocol Tier1 ("the number does not move" past 32K). Vulkan capped at 2 attempts per cell per house policy; both attempts failed identically at every depth tried. pp1024 / tg256 by depth (median of N): | depth | ROCm | Vulkan | vk/rocm pp | vk/rocm tg | |---|---|---|---|---| | 0 | 140.80 / 15.30 | 134.73 / 12.44 | 0.957 | 0.813 | | 32,768 | 81.28 / 12.35 | DEVICE LOST | — | — | | 65,536 | 58.62 / 10.50 | DEVICE LOST | — | — | | 131,072 | 36.86 / 8.86 | DEVICE LOST | — | — | | 262,144 | 21.52 / 6.86 | DEVICE LOST | — | — | DEVICE-LOST CELLS carry the process and kernel evidence verbatim (run-0300 through run-0303; full journalctl -k excerpt at aihydra ~/bench-results/deepseek-v4-flash-fullbench/matrix/vulkan-devicelost-journalctl-k.txt), the same discipline as qwen38-27b's d131072 cells (clm-0054) and the Nemotron OOM cells — a failed cell is a result, not a gap to paper over. The failure recovers without a reboot (kernel resets the compute ring each time) but is deterministic enough at 2/2 across four depths that the fast-fail policy applied. THIS MODEL RAN CLEAN TO d262144 ON ROCm — the deepest cell measured in this programme so far (qwen38-27b's matrix stopped at d131072). See clm-0060 (fit) for why this model tolerates depth so much more cheaply than a conventional hybrid-attention model: its KV cache is ~9x smaller per token.
- evidence: run-0288 run-0289 run-0290 run-0291 run-0292 run-0293 run-0294 run-0295 run-0296 run-0297 run-0298 run-0299 run-0300 run-0301 run-0302 run-0303
clm-0060
- measured-here med ●●○ volatility low · verified 2026-08-15
- DeepSeek-V4-Flash-0731's KV cache costs approximately 7.13 KiB per token on llama.cpp — about a ninth of Qwen3.8-27B's 64.00 KiB/token hybrid-attention cache (clm-0056) — measured from a two-point GTT delta at --parallel 1, f16 KV: 884 MiB between a c=4096 and a c=131072 load probe (99,464 vs 100,348 MiB gtt_used), over 126,976 additional tokens. The fixed (weights + compute-graph) component this implies, 97.10 GiB, matches the GGUF's own reported model_size (104,202,502,492 bytes = 97.06 GiB) almost exactly, which cross-checks the KV-delta method. THE FIT CONSEQUENCE IS LARGE: at this model's full declared 1,048,576-token native context, total GTT usage projects to ~104.1 GiB (fixed 97.1 GiB + ~7.1 GiB KV) — comfortably inside the 120 GiB (122,880 MiB) GTT boot window with ~16 GiB to spare. This model is nowhere near memory-bound on this hardware at ANY context it natively supports; the wall-clock cost of prefilling that far (clm-0059's decode curve implies hours) is the real ceiling, not fit.
- METHOD — two llama-server load-only probes on aihydra, stock ROCm 3653e6d, UD-IQ3_XXS, f16 KV, -ngl 999, -fa on, --load-mode none, --parallel 1 (explicit — the house hazard-class flag; NOT how the 2026-08-15 screen's own c=32768 fit probe was run, which omitted --parallel and defaulted to n_slots=4/kv_unified — so that number, gtt_used=100,422 MiB at c=32768, is NOT directly comparable to this pair and is not used in the arithmetic here). gtt_used read from /sys/class/drm/card0/device/mem_info_gtt_used before/after each load, server torn down between probes. Raw values: c=4096 -> 99,464 MiB (delta 99,446 MiB from an ~18 MiB idle baseline); c=131,072 -> 100,348 MiB (delta 100,330 MiB). Not captured as formal run/config records — a load probe produces no throughput or capability metric in this schema's sense, the same treatment clm-0056's phase-A footprint probes received. ARITHMETIC — (100,348 - 99,464) MiB / (131,072 - 4,096) tokens = 884 MiB / 126,976 tokens = 905,216 KiB / 126,976 tokens = 7.1296 KiB/token. Fixed component at c=4096: 99,464 MiB - (4,096 * 7.1296 KiB) = 99,464 - 28.51 MiB = 99,435.5 MiB = 97.10 GiB. Cross-check against the GGUF's declared model_size (llama-bench JSON): 104,202,502,492 bytes = 97.06 GiB — agrees to within 0.04 GiB, i.e. the "fixed" component IS essentially just the resident weights; this quant's compute-graph/KV overhead at low context is near zero, consistent with an MLA-style (multi-head latent attention) compressed cache rather than a conventional per-head KV allocation. FIT PROJECTION — GTT total 128,849,018,880 bytes = 122,880.9 MiB. At the model's declared 1,048,576-token native context: 99,435.5 MiB (fixed) + 1,048,576 tokens * 7.1296 KiB/token / 1024 = 99,435.5 + 7,301.4 = 106,736.9 MiB = 104.24 GiB, leaving ~16.05 GiB headroom against the 122,880.9 MiB boot budget. Solving the same equation for the ceiling depth the GTT budget alone would permit (ignoring a safety margin) gives ~3.4M tokens — more than 3x the model's own native context — so the GTT window was never going to be the binding constraint for this model on this hardware; native context is. This directly informed the matrix depth selection (clm-0059): depths were chosen for wall-clock usefulness, not because anything deeper risked OOM.
clm-0061
- measured-here med ●●○ volatility medium · verified 2026-08-16
- DeepSeek-V4-Flash-0731 UD-IQ3_XXS completed the FULL standard 26-task tau2 airline set cleanly (rc=0, not a wall-bound cut) at 0.846 mean reward (22/26), 272 tool-call messages, 0 empty assistant turns, 0 infrastructure errors -- VALID under the SMOKE gate. This is well above qwen38-27b's comparable-shape result on the same task-id range (0.577 over an effective 26 of a requested 50, itself a wall-bound cut) -- NOT a controlled comparison (different model, quant, backend, no speculation vs qwen38's draft-mtp arm, and a different agent temperature/output regime), but a genuine data point that this MoE candidate, at its forced IQ3_XXS quant, is not merely "screened," it is competent at the standard task set end to end. The first five tasks -- the standing smoke subset -- scored 5/5 = 1.000, matching the 2026-08-15 screen's own smoke result on the identical seed and task ids exactly, which cross-confirms the screen was not a fluke. Energy: 534.28 Wh across the whole 11,190 s window (mean 171.9 W, delta 161.8 W over the 10.1 W idle floor) -- 24.29 Wh per correct answer, 0.74 pence at 30.3 p/kWh, roughly 1.6x CHEAPER per correct answer than qwen38-27b's 38.29 Wh (clm-0057), even though this model ran with no speculation at a plain ~10-15 t/s decode floor -- consistent with an MoE model activating a small fraction of its 284B total parameters per token against a dense 27B model decoding every parameter every step.
- METHOD -- tau2-bench airline, harness 668d3bc, --num-trials 1 --seed 42 --max-steps 200 --max-concurrency 1, user simulator pinned per protocol (openrouter/anthropic/claude-haiku-4.5, temperature 0, Anthropic-only routing). Agent: local llama-server endpoint on cfg-0084 (ROCm, f16 KV mandatory per the candidate's COMMUNITY HAZARD note, -c 32768, --parallel 1, --load-mode none, no speculation -- this candidate has no native draft/MTP head in the stock build), agent sampling temperature 0.0 / max_tokens 4096 -- identical to the 2026-08-15 screen's SMOKE arm, so this run shares that fingerprint rather than opening a new one. Task-ids 0-25 explicit (the STANDARD 26-task set this programme adopted after qwen38-27b's wall-bound experience, per the operator brief for this job). THE RUN COMPLETED IN FULL: 22:06:31Z to 01:13:01Z, 11,190 s, rc=0. Contrast qwen38-27b's run-0271, which hit an 11,400 s hard wall bound (rc=124) after requesting 50 tasks and completing 26 -- this arm requested exactly 26 and finished all of them with time to spare, at a MUCH slower raw decode floor (10-15 t/s here vs qwen38's 7.8 t/s floor / 18.25 t/s under MTP): the MoE's small active-parameter count evidently keeps per-task wall time competitive despite the missing speculation. FAILED TASKS (4 of 26, reward 0): task 3 (16 turns, 104s, shortest reward-0 task -- an early miss, not a runaway), task 7 (48 turns, 496s), task 21 (30 turns, 550s), task 23 (46 turns, 1,302s -- the longest task in the run). No pattern of cut_at_max_steps (0 across all 26) or infrastructure error; every failure resolved via user_stop, i.e. the simulator ended the conversation judging the task unresolved, not a harness timeout. ENERGY (same-session join, computed immediately after completion, eng-0107): 534.28 Wh total, mean 171.9 W active (delta 161.8 W over the 10.1 W box-idle-all-empty baseline), 22 correct of 26 scored: 24.2853 Wh per correct answer, 0.7358 pence at the standing 30.3 p/kWh tariff. WHOLE-SESSION energy -- includes simulator (user_llm) wait time end to end, not an agent-only or decode-only figure, so it is directly comparable to qwen38-27b's eng-0081 (same denominator: wh_total / tasks_passed, whole session) but not to a decode-only bench-cell wh figure. 534.28 / 574.37 = 0.930 -- this run's TOTAL energy was already lower than qwen38-27b's despite running for a comparable wall time (11,190 s vs qwen38's 11,400 s window), because DeepSeek-V4-Flash's active parameter count per token is far smaller (per clm-0059/clm-0060's fit discussion) even though it decodes at a fraction of qwen38's MTP-boosted speed. CAVEAT ON THE CROSS-MODEL COMPARISON -- this is NOT a controlled A/B: different model, different quant tier (IQ3_XXS forced by fit vs qwen38's Q8_0), different backend history (qwen38 ran under draft-mtp; this candidate has no MTP path at all), and a different agent temperature (0.0 here vs 1.0 for qwen38's thinking-mode default). The comparison is reported because it is the only same-shape (26-task, same simulator, same protocol) data point available in this programme, not because the confound list has been controlled for.
- evidence: run-0287
clm-0062
- measured-here med ●●○ volatility low · verified 2026-08-16
- Nemotron-3-Super-120B-A12B's KV cache costs 8.00 KiB per token at UD-Q4_K_M on llama.cpp — measured byte-exact from a two-point GTT delta, ROCm, f16 KV, --load-mode none, --parallel 1: 224 MiB between a c=4096 and a c=32768 load probe (78,754 vs 78,978 MiB gtt_used), over 28,672 additional tokens (229,376 KiB / 28,672 tok = 8.00 KiB/token exactly). The fixed (weights + compute-graph) component this implies, 76.88 GiB (78,722 MiB), matches the GGUF's own reported model_size (82,533,249,024 bytes = 76.87 GiB) to within 0.01 GiB — a clean cross-check of the method, same discipline as clm-0060's DeepSeek-V4-Flash measurement. THIS DIRECTLY SUPERSEDES THE EARLIER "cannot allocate past d0" READING for this candidate: perf-matrix-gtt120's d32768 cell (2026-08-14) recorded a BadAlloc storm under mmap and inferred d65536/d131072 as OOM without running them — that was clm-0053's mmap double-residency defect (llama-bench's four throughput queue scripts lacked --load-mode none), not a memory-footprint fact about this model. Under the fix, this job's own throughput matrix (run-0308..run-0315) completed d0 and d32768 cleanly on BOTH backends, 3/3 fresh-process reps each, and this KV-cost measurement shows why: at 8.00 KiB/token the model is nowhere near the 122,880 MiB GTT boot budget at any depth this bench tested — headroom past d32768's 78,978 MiB is ~43.9 GiB, enough for several million more tokens by the KV-cost arithmetic alone (before compute-graph/batch-buffer growth at depth is accounted for, which this two-point measurement cannot isolate). The real ceiling for this candidate, like DeepSeek-V4-Flash (clm-0060), is very unlikely to be GTT fit — it is wall-clock: at ~16-18 t/s decode with no speculation path (Mamba rejects draft-mtp, candidate history 2026-06-27), a much deeper matrix cell would cost hours per rep for a number this job's scope did not call for.
- METHOD — two llama-server load-only probes on aihydra, stock ROCm 3653e6d, UD-Q4_K_M, f16 KV, -ngl 999, -fa on, --load-mode none, --parallel 1, gtt_used read from /sys/class/drm/card0/device/mem_info_gtt_used before/after each load, server torn down between probes. Raw values: c=4096 -> 78,754 MiB (delta from a fresh-boot-equivalent stop 78,737 MiB); c=32,768 -> 78,978 MiB (delta 78,961 MiB). Not captured as formal run/config records — a load probe produces no throughput or capability metric in this schema's sense, the same treatment clm-0056 and clm-0060's phase-A/load-probe measurements received. ARITHMETIC — (78,978 - 78,754) MiB / (32,768 - 4,096) tokens = 224 MiB / 28,672 tokens = 229,376 KiB / 28,672 tokens = 8.00 KiB/token exactly. Fixed component at c=4096: 78,754 MiB - (4,096 * 8.00 KiB / 1024) = 78,754 - 32 MiB = 78,722 MiB = 76.88 GiB. Cross-check against llama-bench's own model_size field (verified on this exact GGUF via this job's matrix JSONs): 82,533,249,024 bytes = 76.87 GiB — agrees to within 0.01 GiB. COMPARISON TO PEERS — 8.00 KiB/token sits close to DeepSeek-V4-Flash's 7.13 KiB/token (clm-0060, an MLA-style compressed cache) and far below Qwen3.8-27B's 64.00 KiB/token (clm-0056, hybrid 3:1 linear-to-full-attention layout). Two architecturally very different models (hybrid Mamba-2 + LatentMoE MoE here, MLA MoE there) land within 12% of each other on KV cost per token, both far cheaper than a conventional hybrid-attention allocation — worth noting as a pattern, not yet a claim about why. SUPERSEDES — perf-matrix-gtt120 (2026-08-14): nemotron3-super-rocm-kvf16-faon- d32768-rep1.stderr shows a sustained "HSA exception: BadAlloc" storm (dozens of lines) starting near process start; the d65536/d131072 cells were never run and instead recorded as `oom-inferred` with basis "measured BadAlloc OOM at d32768... at GTT120". clm-0053 root-caused this defect class (mmap double-residency competing with the KFD system-memory gate for host RAM, measured on this exact candidate) and fixed the four throughput queue scripts that lacked --load-mode none. This job's matrix used the fixed queue script throughout (see cfg-0087/cfg-0088's provenance) and both d0 and d32768 completed clean on both backends — the direct confirmation the fix holds for a fresh nemotron3-super matrix, not just the fix's own original test cell.
- evidence: run-0308 run-0309 run-0310 run-0311 run-0312 run-0313 run-0314 run-0315
clm-0063
- measured-here med ●●○ volatility medium · verified 2026-08-16
- On gfx1151 at UD-Q4_K_M / f16 KV, Nemotron-3-Super-120B-A12B is the FIRST model in this programme where Vulkan decode measurably beats ROCm: +8.97% at d0 (18.2043 vs 16.7065 t/s) and +8.19% at d32768 (17.7386 vs 16.3962 t/s), a stable signature by depth. ROCm keeps its usual prefill lead (+27.1% at d0, +26.9% at d32768) — so the per-model, per-phase backend rule (clm-0050) holds again, but this is the first candidate where the two phases point to DIFFERENT backends being "the better one" rather than the same backend winning both or Vulkan losing outright to a device-loss crash. Unlike qwen38-27b (clm-0054) and deepseek-v4-flash (clm-0059), Vulkan showed ZERO device-loss at d32768 here — 3/3 fresh-process reps clean on both depths tested, no vk::DeviceLostError, no amdgpu ring reset. Because decode dominates wall-clock in a multi-turn agentic conversation (prefill is amortised across turns via prefix-cache reuse; decode is not), this job picked Vulkan as the serving backend for the 26-task tau2 arm — a genuinely new per-model call, not an inherited default.
- METHOD — nemotron3-super-fullbench-matrix: stock llama.cpp 3653e6d (ROCm) / 3653e6d6d (Vulkan, prefix-matches per house convention), UD-Q4_K_M, f16 KV (candidate's serving config, cfg-0072/cfg-0089 — not swept here), fa on, -ngl 999, --load-mode none, pp1024/tg256, llama-bench defaults otherwise (-b 2048/-ub 512, -t 16). N=3 fresh-process reps at d0/d32768 (this job's brief-specified depth set — "up to the fit ceiling"; see clm-0062 for why d32768 is nowhere near this candidate's actual memory ceiling post-fix). Every cell's coefficient of variation <=0.92% — clean per protocol §0's ~3% scatter-flag threshold. pp1024 / tg256 by depth (median of 3): | depth | ROCm | Vulkan | vk/rocm pp | vk/rocm tg | |---|---|---|---|---| | 0 | 271.5460 / 16.7065 | 213.7281 / 18.2043 | 0.787 | 1.090 | | 32,768 | 244.7670 / 16.3962 | 192.8306 / 17.7386 | 0.788 | 1.082 | THE DECISION — cfg-0089 (the tau2 serving config) was set to Vulkan on this measurement, not the 2026-08-15 screen's ROCm pick (cfg-0072, made when no throughput data existed and both backends looked equally healthy at the FIT stage). This job's house 4-item guard (run-0307) then ran fresh against the Vulkan config and passed 4/4 clean before the tau2 arm launched, per protocol's guard-substitutions rule — the backend switch got its own capability check rather than inheriting the screen's ROCm-side guard evidence. WHY THIS DIFFERS FROM THE PRIOR PATTERN — qwen38-27b and deepseek-v4-flash both showed Vulkan collapsing or crashing with depth, so ROCm was the only real option past d0. Nemotron-3-Super's fit ceiling in this matrix (d32768) is shallower than where those models' Vulkan arms broke (>=32,768/>=131,072 respectively) only in the sense that this job did not test deeper cells (clm-0062) — whether Vulkan would eventually lose the device on this candidate at a much greater depth is UNTESTED, and the tau2 arm's serving context (32768) sits exactly at the deepest cell this matrix cleared, not comfortably past it. This is noted as an open question on the model page, not glossed over.
- evidence: run-0308 run-0309 run-0310 run-0311 run-0312 run-0313 run-0314 run-0315
clm-0064
- measured-here med ●●○ volatility medium · verified 2026-08-16
- Nemotron-3-Super-120B-A12B UD-Q4_K_M completed the FULL standard 26-task tau2 airline set cleanly (rc=0, not a wall-bound cut) at 0.7692 mean reward (20/26), 263 tool-call messages, 429 assistant messages, 0 empty assistant turns, 0 infrastructure errors -- VALID under the SMOKE gate. This is the FIRST tau2 result for this candidate, at ANY quant, run under the protocol's PINNED independent simulator (openrouter/anthropic/claude-haiku-4.5) -- every prior full tau2 number on this candidate (the IQ4_XS acquisition-era run-0102, and the quant-flip pair run-0107/run-0108) ran SELF-PLAY, agent_llm == user_llm, which the candidate's own open-questions flag as measuring something weaker ("self-play flattens what the benchmark measures"). The 2026-08-15 screen's 5-task SMOKE (4/5 = 0.800) was the only other pinned-simulator number this candidate had; this run's first five tasks (task-ids 0-4) score 4/5 = 0.800 again on the identical seed and task set, cross-confirming the screen was not a fluke -- SAME reward, though not necessarily the same per-task pattern, since the screen's own per-task breakdown is not separately recorded. VULKAN, not the screen's ROCm serving pick -- this job's own throughput matrix measured Vulkan +8-9% faster decode with zero device-loss at this candidate's fit ceiling (clm-0063), so the capability-establishing run itself ran on the newly-measured better backend, cfg-0089. Energy: 648.60 Wh across the whole 15,968 s (4.44 h) window, mean 146.23 W (delta 136.13 W over the 10.1 W idle floor) -- 32.43 Wh per correct answer, 0.98 pence at 30.3 p/kWh, whole-session (includes simulator wait time end to end). Against this programme's other two full-bench candidates on the SAME denominator: 1.34x qwen38-27b's 38.29 Wh/correct answer (draft-mtp dense, cheaper) but 1.15x more than deepseek-v4-flash's 24.29 Wh/correct answer (a small-active-param MoE) -- despite this candidate having the LARGEST active-parameter footprint measured on this box (12B active vs deepseek's ~6B-equivalent MoE activation and qwen38's dense 27B all-active), it lands in the middle of the three on Wh/correct, not at either extreme, which is a genuinely interesting finding worth noting rather than assuming active-parameter count predicts energy ranking cleanly.
- METHOD -- tau2-bench airline, harness 668d3bc (same version as qwen38-27b's run-0271 and deepseek-v4-flash's run-0287), --num-trials 1 --seed 42 --max-steps 200 --max-concurrency 1, user simulator pinned per protocol. Agent: local llama-server endpoint on cfg-0089 (Vulkan, f16 KV mandatory per the screen's serving config, -c 32768, --parallel 1, --load-mode none, no speculation -- Mamba rejects draft-mtp), agent sampling temperature 0.0 / max_tokens 4096 -- identical to the 2026-08-15 screen's SMOKE arm. Task-ids 0-25 explicit (the STANDARD 26-task set this programme adopted after qwen38-27b's wall-bound experience). THE RUN COMPLETED IN FULL: 02:32:56Z to 06:59:04Z, 15,968 s, rc=0, well inside the 28,800 s (8 h) safety ceiling. Slowest task (10) took 1,629.6 s (27.2 min); several tasks ran past 700-800 s. No pattern in the six failures (tasks 7, 10, 14, 20, 23, 24) suggesting a systematic cliff -- all resolved via user_stop, spread across both short (tc=6) and long (tc=19) tool-call counts, and across the full duration range (404.8 s to 1,629.6 s). No cut_at_max_steps, no infrastructure error, in any of the 26 tasks. SIMULATOR PROVENANCE MATTERS HERE MORE THAN USUAL: this is a genuinely new measurement, not a re-run of an existing number under a different simulator -- there is no earlier pinned-simulator FULL run to compare 0.7692 against for this candidate. The only same-simulator anchor is the 2026-08-15 SMOKE's 5-task subset (0.800), which this run reproduces exactly on tasks 0-4. ENERGY (same-session join, eng-0116): counter interpolated from raw ~10s HA history samples at each run boundary (not the coarser 5-minute statistics bucket, given the multi-hour window's need for precision at the exact run-meta timestamps): start 20.99128201 kWh, end 21.63988203 kWh, delta 0.64860002 kWh = 648.60 Wh. Mean power 146.23 W sits between this job's own throughput-matrix cells at the same depth (rocm-d32768 147.74 W / eng-0113, vulkan-d32768 149.47 W / eng-0115) -- a sanity-check the tau2 arm's power draw is not anomalous relative to raw decode-bound throughput at the same serving context. CROSS-MODEL ENERGY CAVEAT -- not a controlled A/B: three different quants (Q4_K_M here, Q8_0 for qwen38-27b, IQ3_XXS for deepseek-v4-flash), three different architectures (hybrid Mamba-2+LatentMoE MoE / dense hybrid-attention / MLA MoE), different backends (Vulkan here, ROCm for both peers), and qwen38-27b ran with draft-mtp speculation while this candidate and deepseek-v4-flash both ran plain decode. Reported because it is the only same-shape (26-task, same simulator, same protocol, same wh_per_task denominator) energy comparison available in this programme, not because the confound list has been controlled for.
- evidence: run-0316
clm-0065
- measured-here med ●●○ volatility medium · verified 2026-08-16
- On gfx1151, Laguna-S-2.1 Q4_K_M is a sixth model measured under the per-model, per-phase backend rule (clm-0050), and unlike nemotron3-super and deepseek-v4-flash (both Vulkan-favoured on this box), ROCm wins outright here: ROCm swept the full throughput matrix clean at every depth tried (d0, d32768, d131072 -- 300.7/23.68 t/s down to 116.43/7.19 t/s pp/tg), while Vulkan completed only d0 and d32768 (256.9/13.29 and 51.6/12.27 t/s pp/tg, ROCm ahead on every figure at every depth both backends completed) and lost the GPU device on both allowed attempts at d131072 (rc=134, vk::DeviceLostError; kernel evidence: amdgpu ring timeout, ring reset FAILED, a MODE2 GPU reset, "device wedged, but recovered through reset" -- the same device-loss class as clm-0054/clm-0059's precedent, now a fourth confirmed occurrence on this silicon). ROCm's decode lead is largest at d0 (23.68 vs 13.29 t/s, +78%) and narrows but does not close at d32768 (15.53 vs 12.27 t/s, +26%); ROCm's prefill lead widens sharply with depth (0.957x at d0 -> 4.17x at d32768, ROCm ahead both times). Serving this model at any useful context on this box requires ROCm -- Vulkan cannot be trusted to survive a conversation that grows past 32k, and even where it survives it is already behind on both phases.
- METHOD -- laguna-s21-fullbench matrix: stock llama.cpp 3653e6d (ROCm) / 3653e6d6d (Vulkan, prefix-matches per house convention), f16 KV (matches the 2026-08-10 screen's serving config), fa on, -ngl 999, --load-mode none, pp1024/tg256, llama-bench defaults otherwise (-b 2048/-ub 512, not passed explicitly). N=3 fresh-process reps at d0/d32768 (stddev <=1.01 t/s on every cell, CoV <=0.88% on every cell -- well inside the 3% scatter-flag threshold, a healthy code path per protocol §0), N=1 at d131072 per protocol Tier1 / deepseek-v4-flash convention. Vulkan capped at 2 attempts per cell per house policy; both attempts at d131072 failed identically. pp1024 / tg256 by depth (mean of N): | depth | ROCm | Vulkan | vk/rocm pp | vk/rocm tg | |---|---|---|---|---| | 0 | 300.67 / 23.68 | 256.92 / 13.29 | 0.855 | 0.561 | | 32,768 | 215.11 / 15.53 | 51.64 / 12.27 | 0.240 | 0.790 | | 131,072 | 116.43 / 7.19 | DEVICE LOST | -- | -- | DEVICE-LOSS CELL (run-0327): attempt 1 (08:21:07Z-08:38:22Z, 1035s) and attempt 2 (08:38:22Z-08:53:03Z, 881s) both terminated identically -- "radv/amdgpu: The CS has been cancelled because the context is lost. This context is innocent." -> vk::Queue::submit: ErrorDeviceLost, rc=134. The job's own per-cell dmesg capture (matrix/vulkan-d131072-rep1.dmesg.txt) is 0 BYTES -- a capture bug in queue-laguna-s21-fullbench-matrix.sh: its on-failure `dmesg` call ran without `sudo` and silently produced nothing under this host's permissions (the aihydra user has passwordless sudo per /etc/sudoers.d/90-aihydra; the script just never invoked it for this call site). This is a gap in the SCRIPT, not evidence the failure did not happen. The kernel evidence was independently re-pulled this session via `sudo dmesg -T` scoped to the failure window and confirms the device-loss class exactly: 08:37:54Z ring comp_1.2.0 timeout, signaled seq=7255308 -> ring reset -> recovered 08:38:25Z ring comp_1.2.0 timeout, signaled seq=7255338 -> ring reset -> recovered 08:44:46Z Fence fallback timer expired on ring comp_1.2.0 08:53:05Z ring comp_1.2.0 timeout, signaled seq=7302878 -> ring reset FAILED -> GPU reset begin! (MODE2 reset) -> GPU reset succeeded, trying to resume -> SMU resumed successfully -> GPU reset(21) succeeded -> "device wedged, but recovered through reset" -> "*ERROR* Failed to initialize parser -125!" The box self-recovered via the MODE2 reset on the second failed attempt -- no reboot needed, no lingering wedge, matching the clm-0054/clm-0059 precedent exactly. Fix needed in queue-laguna-s21-fullbench-matrix.sh for future runs: prefix its on-failure `dmesg` call with `sudo`, or its evidence file silently stays empty every time. Laguna's SWA-heavy layer mix (36/48 sliding-window + 12/48 global attention, per the candidate record) does not spare it from this class -- unlike deepseek-v4-flash's MLA-compressed cache, which ran clean on ROCm through d262144 and was never tested to the point of Vulkan failure at comparable depth in this programme's matrix design.
- evidence: run-0317 run-0318 run-0319 run-0320 run-0321 run-0322 run-0323 run-0324 run-0325 run-0326 run-0327
clm-0066
- measured-here med ●●○ volatility medium · verified 2026-08-16
- Laguna-S-2.1's KV cache costs approximately 48.0 KiB per token on llama.cpp at f16 -- reportedly measured from a two-point GTT load probe (c=4096 -> 92,131.38 MiB gtt_used, c=131,072 -> 98,083.38 MiB gtt_used; delta 5,952.00 MiB over 126,976 tokens). This is LARGER than nemotron3-super's 8.00 KiB/token and deepseek-v4-flash's 7.13 KiB/token despite this model's SWA-heavy layer mix (36/48 sliding-window + 12/48 global attention, per the candidate record): the 12 full-attention layers still dominate the aggregate KV cost, so "SWA-heavy" only discounts the bill relative to an all-global model of the same size, not in absolute terms. At the matrix's deepest tested depth (c=131,072), projected GTT use is ~98.1 GiB against the 120 GiB (122,880 MiB) boot window -- comfortably inside, with ~24.2 GiB headroom, and nowhere near the fit ceiling this job's matrix was designed around.
- RE-MEASURED 2026-08-16 (bounded add-on, same day, slotted between ornith-35b screen arms): the two-point GTT probe was re-run properly on stock 3653e6d ROCm, laguna model, --load-mode none -np 1, f16 KV, gtt_used sampled from /sys/class/drm/card0/device/mem_info_gtt_used 10s after each server reported healthy (settling window), server torn down between points -- this time capturing the actual readings to a file (kvprobe-c4096-redo.log / kvprobe-c131072-redo.log in ~/bench-results/laguna-s21-fullbench/ on aihydra), unlike the two broken prior attempts this record's confidence:low was about. RAW: c=4096 -> gtt_used_after_settle=96,380,256,256 bytes (91,915.75 MiB); c=131,072 -> gtt_used_after_settle=102,621,380,608 bytes (97,867.75 MiB). Delta = 6,241,124,352 bytes = 5,952.00 MiB over 126,976 tokens = 6,094,848 KiB / 126,976 tok = 48.0000 KiB/token EXACTLY -- confirms the previously-uncorroborated figure to four significant figures. No correction needed. This re-measurement's MiB pair (91,915.75 / 97,867.75) matches the "matrix script planning comment" pair cited in the original note below (91,915 / 97,867) to within rounding -- effectively an exact reproduction -- and differs from the "session's MODEL-PAGE.md draft" pair (92,131.38 / 98,083.38) by ~215.6 MiB at both points while producing an IDENTICAL 5,952.00 MiB delta, consistent with the original note's idle-baseline-drift-between-two-real-measurements read rather than fabrication. CROSS-CHECK: fixed (weights+graph) component at c=4096 = 91,915.75 MiB - (4,096 tok * 48.0 KiB/tok / 1024) = 91,915.75 - 192.00 = 91,723.75 MiB = 89.574 GiB, against the GGUF's own model_size (96,031,829,760 bytes = 89.428 GiB) -- agrees to within 0.146 GiB (compute-graph overhead), the same order of agreement clm-0060's method produced. Confidence raised low -> medium: this is now a directly re-derivable two-point probe with a surviving raw artifact, matching clm-0060's evidentiary standard. Not raised further because clm-0060 additionally cross-validated a fit projection against the model's full native context; that extension is out of scope for this bounded add-on and is left as a follow-up if a fit projection is wanted for laguna-s-21 too. Per house convention (clm-0060, clm-0056 phase-A), a load-only GTT probe produces no throughput or capability metric in the schema's sense and is not captured as a formal run/config record -- documented here instead, matching prior practice for this exact measurement class. --- ORIGINAL confidence:low investigation note, preserved for provenance --- CONFIDENCE DOWNGRADED FROM THE USUAL "medium" FOR THIS METHOD (contrast deepseek-v4- flash's clm-0060, same method, confidence medium) -- per house policy against trusting prose over raw JSON, this figure could NOT be independently re-derived from any surviving artifact this session, despite being explicitly named as source evidence for this job. What was pulled from aihydra: - kvprobe-c4096.log / kvprobe-c131072.log (the two files this job's brief names as the source for the 48.0 KiB/token figure): both are plain llama-server startup/shutdown logs (15 lines each, server load_model -> model loaded -> listening -> cleaning up). NEITHER FILE CONTAINS A SINGLE gtt_used READING. The `cat /sys/class/drm/card0/device/mem_info_gtt_used` calls this probe method requires (per fitprobe-c4096.log/fitprobe-c131072.log's own sibling probes, and per deepseek-v4-flash's clm-0060 METHOD note) were evidently run but their output went to a terminal, not to any file that survived -- confirmed by byte-for-byte review (wc -l = 15 on all four probe logs, no MiB/gtt string anywhere in any of them). - power.jsonl (the box's periodic power-telemetry log) does not cover this window (last sample 2026-08-14T10:59:34Z, two days before this job ran) and does not record GTT usage in any case (fields are gpu_w/edge_c/use_pct only). - The job's own matrix script (queue-laguna-s21-fullbench-matrix.sh) cites the SAME 91,915 -> 97,867 MiB pair the delegate's MODEL-PAGE.md draft calls "the prior delegate's cited figures" -- as a COMMENT written when the matrix was planned, not as a captured measurement file. This is the only trace of that number anywhere outside prose. So there are now TWO cited MiB pairs for this measurement (91,915/97,867 from the matrix script's planning comment, and 92,131.38/98,083.38 from this session's MODEL-PAGE.md draft) and NEITHER has a surviving raw capture. The two pairs differ by ~200-216 MiB at each point (plausible idle-baseline drift between two real measurements taken on different days, per the draft's own framing) and both derive the same ~48.0 KiB/token rate -- which argues mildly against outright fabrication (a fabricated number would more likely be copy-pasted exactly, not reproduced from two slightly different readings), but is not a substitute for a raw artifact either. PARTIAL CROSS-CHECK: subtracting KV-at-c=4096 (4,096 tok * 48.0 KiB/tok / 1024 = 192.0 MiB) from each cited c=4096 reading gives an implied fixed (weights + graph) component of 91,723 MiB (89.57 GiB) for the "prior delegate" pair and 91,939.4 MiB (89.78 GiB) for this session's pair, against the GGUF's own model_size (96,028,095,488 bytes = 89.42 GiB, read directly from every matrix cell JSON this session, e.g. rocm-d0-rep1.json). Both are within ~0.4 GiB of the weights figure -- plausible, and consistent with the method deepseek-v4-flash's clm-0060 used successfully -- but this is a consistency check against a THIRD independently-verified number (model_size), not a verification of the two GTT readings themselves. Published at confidence:low rather than suppressed, per protocol's "publish the frontier, not the winner" / "risk flagged rather than the row suppressed" -- the number is plausible and useful context, but this record should not be read as having the same evidentiary weight as clm-0060's fully-recoverable two-point probe.
clm-0067
- measured-here med ●●○ volatility medium · verified 2026-08-16
- Laguna-S-2.1 Q4_K_M completed the FULL standard 26-task tau2 airline set cleanly (rc=0, not a wall-bound cut) at 0.6923 mean reward (18/26), 187 tool-call messages, 0 empty assistant turns, 0 infrastructure errors -- VALID under the SMOKE gate. Wall time 3,444s (57.4 min) against a 28,800s (8h) safety ceiling never approached. Energy: 130.66 Wh across the run window (mean 136.6 W, delta 126.5 W over the 10.1 W idle floor), whole-session, joined same-session from HA history -- 7.2588 Wh per correct answer (0.2199 pence at 30.3 p/kWh), the LOWEST Wh-per-correct-answer of any full standard-26-task tau2 comparison run on this box to date: cheaper than deepseek-v4-flash's 24.29 Wh (clm-0061), nemotron3-super's 32.43 Wh, and qwen38-27b's 38.29 Wh (clm-0057) -- roughly a third of deepseek-v4-flash's figure and under a fifth of qwen38-27b's, despite plain decode throughout (no speculation/MTP exists for this architecture) at a raw floor of 15.53 t/s (d32768, the serving depth).
- METHOD -- tau2-bench airline, harness 668d3bcd135c02aa3438f987ef45735b7c163ee3 (same version as qwen38-27b/deepseek-v4-flash/nemotron3-super's full-bench arms), --num-trials 1 --seed 42 --max-steps 200 --max-concurrency 1, user simulator pinned per protocol (openrouter/anthropic/claude-haiku-4.5, temperature 0, Anthropic-only routing). Agent: local llama-server endpoint on cfg-0092 (ROCm, f16 KV, -c 32768, --parallel 1, --load-mode none, no speculation), agent sampling temperature 0.0 / max_tokens 4096. Task-ids 0-25 explicit (the standard 26-task set). Guarded by run-0328 (house 4-item guard, 4/4, against this exact serving config, run fresh immediately before this arm). ALL NUMBERS IN THIS CLAIM VERIFIED DIRECTLY FROM RAW ARTIFACTS, not copied from the delegate's MODEL-PAGE.md draft: tool_calls/assistant_msgs/empty/mean_reward summed and averaged from tau2-full-20260816T091034Z/armcompare.json's per_task records (187/388/0/0.692307692...); duration range (46.9s task 19 to 370.6s task 23) read from the same file; termination_reason == "user_stop" on all 26 confirmed. Energy independently recomputed from aihydra's sensor.hardware_ai_hydra_energy raw ~10s HA history for the exact run-meta.jsonl window (09:11:47Z-10:09:11Z), NOT copied from the draft's "resume delegate" join -- see eng-0124. This run's own numbers reproduced the draft's headline arithmetic (7.26 Wh/correct, 5.03 Wh/task, 136.6 W mean) to within rounding; the one discrepancy found anywhere in this job's energy join was the much smaller guard window (eng-0123), not this headline figure. FAILED TASKS (8 of 26, reward 0.0): task 1 (7 tc, 71.2s), task 5 (6 tc, 105.2s), task 7 (12 tc, 148.1s), task 10 (12 tc, 184.3s), task 14 (8 tc, 242.1s), task 16 (5 tc, 151.5s), task 21 (10 tc, 166.6s), task 23 (8 tc, 370.6s -- the longest task in the run). Every failure resolved via user_stop -- the simulator ending the conversation judging the task unresolved, not a harness timeout or agent error. EFFICIENCY-LEADER CLAIM SCREENED AGAINST EVERY OTHER MODEL PAGE'S wh_correct_energy BEFORE PUBLISHING (per the job brief's explicit instruction): deepseek-v4-flash eng-0107 = 24.29 Wh/correct, nemotron3-super eng-0116 = 32.43 Wh/correct, qwen38-27b eng-0081 = 38.29 Wh/correct -- laguna-s-21 leads all three on the SAME protocol (full standard 26-task set, pinned haiku-4.5 simulator, same energy-join method). ONE LOWER NUMBER EXISTS ON THE SITE: qwen35-122b's eng-0034 records 5.713 Wh/correct, but that run is a 5-task SMOKE window from 2026-08-09 -- a different scale (5 vs 26 tasks) on a different, pre-"standard-26-task-convention" protocol (the standard set was adopted post-qwen38, per run-0287/run-0329's own notes) -- not a like-for-like comparison, and this claim deliberately does NOT assert an unqualified site-wide superlative on the strength of it. The comparison above is scoped to full standard-26-task tau2 runs only. Speculation/MTP: not applicable -- llama.cpp source has no nextn/draft-tensor wiring for LLM_ARCH_LAGUNA (checked src/llama-model.cpp, src/models/laguna.cpp directly). Laguna has no MTP architecture support at all in this tree, unlike qwen38-27b's draft-mtp arm -- this candidate's efficiency lead is achieved at a raw plain-decode floor (15.53 t/s at the serving depth), the same class of result deepseek-v4-flash's clm-0061 noted for its own MoE-vs-dense comparison: a small active-parameter footprint (~8B active of 117.6B total) apparently costs less energy per correct answer than faster dense/speculated decode, even before accounting for absolute speed. CAVEAT ON EVERY CROSS-MODEL COMPARISON ABOVE -- none of these are controlled A/B: different models, different architectures (hybrid SWA/global MoE here vs deepseek's MLA-compressed MoE, nemotron's Mamba-hybrid, qwen38's dense+MTP), different quant tiers, different backends (ROCm here and for deepseek-v4-flash; Vulkan for nemotron3-super). Reported because it is the only same-shape (26-task, same simulator, same protocol, same energy-join method) set of data points available in this programme, not because the confound list has been controlled for.
- evidence: run-0329
clm-0068
- measured-here med ●●○ volatility medium · verified 2026-08-16
- Ornith-1.0-35B (ornith-ai/deepreinforce-ai, arch qwen35moe, A3B-class MoE, MIT) passed the house standard screen on release-adjacent acquisition, stock 3653e6d ROCm, Q8_0, -np 1, c=32768: FIT loaded in 12s at 35,607 MiB GTT (comfortable against the 122,880 MiB boot window), GUARD 4/4 (coherence, tool call, needle at 8000 tokens, isolation skipped under -np 1), SMOKE 5/5 with mean_reward 1.000 on the 5-task tau2 airline smoke (21 tool-call messages, 0 empty assistant turns, 0 infra errors -- VALID per the smoke gate; run-0330, run-0331). This is a joint-strongest smoke result on this box to date, and directionally consistent with the vendor's own agentic-coding benchmark table (Terminal-Bench 2.1 64.2 vs Qwen3.6-35B's 52.5, SWE-bench Verified 75.6 vs 73.4).
- ACQUISITION: Q8_0 from the OFFICIAL repo (ornith-ai/Ornith-1.0-35B-GGUF, 36,903,138,880 bytes, sha256 cbc992bca07901c1a51f33e65e6fc5d687de179c852a772dfd 15e4c3261dbf5c -- verified exact against the HF LFS oid post-download) as the primary screening quant; UD-Q4_K_XL and mmproj-F16 from unsloth/Ornith-1.0-35B- GGUF (22,324,804,000 bytes / 899,283,680 bytes, both sha256-verified exact against HF) as the comparability arm and vision file. Q8_0 is not in protocol.json quants.expected -- explicit PROTOCOL_OVERRIDE carried on every run (same rationale as qwen38-27b's cfg-0066: removes the quant confound, matches the candidate's own designated primary quant). VISION CLAIM IS UNCONFIRMED FROM THE PRIMARY SOURCE. The task brief that motivated this acquisition described Ornith-1.0-35B as "vision-capable," but the official README's own highlights section describes the family purely as agentic-coding models (9B-Dense / 31B-Dense / 35B-MoE / 397B-MoE) with zero mention of vision, multimodal input, or an mmproj artifact anywhere in the page -- and the OFFICIAL repo ships NO mmproj file at all (siblings: only the five weight quants + README + two logo assets). unsloth's mirror DOES ship mmproj-{BF16,F16,F32}.gguf, but unsloth mechanically derives mmproj files from whatever vision_config is present in the base checkpoint's config -- this is consistent with "post-trained on top of ... Qwen 3.5" (the README's own framing) carrying over Qwen3.5's vision tower structurally without the coding-focused fine-tune having trained or evaluated vision capability at all. Acquired per the job brief regardless (the file is small, 899 MiB, and useful library context either way) but NOT screened -- the NPU/GPU screening tiers do not currently define a vision screen (same caveat qwen38-27b's mmproj-F16 carries) -- and the "vision-capable" framing should not be repeated as a confirmed fact without an actual vision-path test against the unsloth mmproj. SCREEN METHOD NOTE: the first screen attempt (not ingested) omitted -np 1, defaulting to n_slots=4 / kv_unified=true -- a deviation from this job's stated house rule (--parallel 1 for every measured arm), inherited by copying the nemotron3-super-q4km-screen.sh template which has the same gap. Caught before ingestion; re-run clean with -np 1 explicit (n_slots=1, kv_unified=false confirmed in the server log) and BOTH runs scored identically (5/5, 1.000, tool-call counts 21 vs 25 -- both VALID, well within normal task-to-task variance for an agentic simulator), so the deviation did not appear to affect the result in this instance, but the ingested numbers (run-0330/run-0331) are from the corrected -np 1 run only. Energy: both runs carry energy: null -- run-meta.jsonl (aihydra) has the started_at/ended_at pair for a follow-up batch join before the ~10-day HA retention window closes (protocol §11.2); not joined within this job's bounded time budget. Not yet run: the throughput/depth matrix (qwen38-27b's queue-qwen38-screen.sh shape -- backend x depth x quant grid) and any MTP/speculative arm. The GGUF metadata was not inspected for MTP tensors in this job; if present, the qwen35moe family's own EOS-cliff precedent (clm-0055, on the architecture sibling qwen38-27b) is the first thing any MTP arm on this model should check for before trusting a spec-draft-n-max >= 4 result.
- evidence: run-0330 run-0331
clm-0069
- measured-here high ●●● volatility low · verified 2026-08-16
- github.com/julianmb/q38rocm (r/StrixHalo 1vpiwz0) is a genuine, substantial llama.cpp fork -- charlie12345/ROCmFPX, based on official llama.cpp b9438 (commit 22cadc194), pinned at e87d53e for this artifact -- carrying real custom ROCmFP4/ROCmFP4_FAST GGUF block-quant tensor formats and a real --spec-mtp-strict-qwen exact-verification mode for qwen35/qwen35moe MTP (implemented in tools/server/server-context.cpp + common/arg.cpp, gated on -np 1 and sufficient recurrent-rollback depth, dynamically caps draft length to stay inside one 256-token dense-attention KV block to avoid a ROCm floating-point rounding divergence they found at block boundaries). It is NOT vaporware. BUT: their own published/quickstart config does not use it. Two independent, concrete findings against the "mathematically lossless" framing the Reddit thread promotes: (1) run_server.sh (the repo's actual launcher) defaults to --spec-draft-n-max 6 with NO --spec-mtp-strict-qwen flag; the README's headline "Deep Spec" arm goes to n=7 -- both well past the n_max>=3 hard ceiling this lab's own screening independently established for the SAME architecture family (qwen38-27b, clm-0055: n_max>=4 corrupts generations via an EOS-cliff, target logits collapsing to ~92% on im_end after ~1.4k accumulated tokens). The fork's OWN source code self-documents the risk when strict mode is off: "Qwen MTP strict verification is disabled; greedy output may diverge from no-spec decoding" (server-context.cpp:866) -- i.e. their benchmarked 30.56-36.04 tok/s numbers were measured on a config their own fork admits is not verified-exact. (2) The FP4 artifact does not load on stock llama.cpp at all -- empirically confirmed, not just inferred from the README's claim: stock 3653e6d ROCm fails with a verbatim, precise error: gguf_init_from_reader: tensor 'output.weight' has invalid ggml type 101. should be in [0, 43) -- proving the ROCmFP4 tensor type (101) sits outside stock GGML's valid enum range and requires the ROCmFPX fork's extended type table. Per task protocol ("if the artifact only works on their fork, record that as the screening result and stop"), no GUARD or SMOKE was attempted.
- ARTIFACT: julianmb/Qwen-3.8-27B-ROCmFP4-FAST-GGUF, file Qwen3.8-27B-ROCmFP4- FAST.gguf, 14,562,236,384 bytes (13.56 GiB, 4.26 bpw), sha256 fb89c78d2be91cdb68eaaaa45b1270710bf34aa721dc1f0b9e3aa7b98d2e1da9 -- downloaded to aihydra ~/models/q38rocm-fp4-screen/ and sha256-verified exact against the HF LFS oid before the stock-load attempt. BASE + PATCHES (from build_engine.sh + ROCMFP4-UPSTREAM-INTEGRATION.md in the charlie12345/ROCmFPX clone): branch rocmfp4-upstream-b9438-integration, upstream baseline official llama.cpp b9438 / commit 22cadc194. Carried work: Q4_0_ROCMFP4 and Q4_0_ROCMFP4_FAST custom GGUF tensor formats + quantization tooling; CPU reference, HIP/ROCm and Vulkan runtime support for those types; ROCmFP4 KV-cache FlashAttention handling; an MTP host-path embedding-fetch cleanup; a Vulkan exact-scale-search pruning optimisation; plus unrelated StepFun Step 3.7 Flash conversion/runtime support carried in the same tree. q38rocm itself (the julianmb repo the Reddit thread links) ships NO llama.cpp source -- it clones charlie12345/ROCmFPX fresh at build time and pins commit e87d53e ("213") in its README/Limitations section; all of the above is charlie12345/ROCmFPX's own work, not julianmb's. --spec-mtp-strict-qwen MECHANISM (read from ROCmFPX source, not executed): gates on general.architecture == qwen35 or qwen35moe AND draft-mtp enabled; requires -np 1 ("Qwen strict MTP requires a single server slot/sequence"); requires llama_n_rs_seq(ctx_tgt) >= spec-draft-n-max ("bounded recurrent rollback covering the full draft"). When active: caps n_draft_max per step to stay within a 256-token dense-attention KV padding block ("A verification batch that straddles a block makes its earlier rows use a wider ROCm reduction than serial decoding and can change greedy output through rounding" -- server-context.cpp ~2469) and forces a full prompt-cache invalidation on a specific cache-hit path to keep MTP boundary state paired (server-context.cpp ~2929, "prompt cache cold fallback: reason=strict-qwen- exact-hit"). This targets a DIFFERENT failure mode than clm-0055's EOS-cliff (a ROCm floating-point rounding divergence at KV block boundaries vs. this lab's target-distribution collapse to im_end after accumulated session volume) -- it is NOT established, and should not be assumed, that enabling --spec-mtp-strict-qwen would prevent the EOS-cliff signature this lab characterised; that is an open question for anyone actually running their fork, out of scope here (task protocol: audit only, no third-party binaries executed). q38rocm's own npu_sidecar_drafter.py exposes --strict as an opt-in flag (default OFF, same as run_server.sh's silent omission) with the argparse help text "Enable strict lossless greedy equivalence" -- the marketing language ("mathematically lossless") describing a mode the fork COULD run in but that neither of the repo's two launch paths enables by default. STOCK-LOAD ATTEMPT (empirical, this job): `~/src/llama.cpp/build/bin/ llama-server -m Qwen3.8-27B-ROCmFP4-FAST.gguf -ngl 999 -fa on -c 4096 --load-mode none` on stock 3653e6d ROCm, rc=1, ~50ms to failure. Verbatim: "gguf_init_from_reader: tensor 'output.weight' has invalid ggml type 101. should be in [0, 43)" then "failed to load model from ...". This is the SAME charlie12345/ROCmFPX fork already carried as a blocker in this lab's records for the Nemotron-3.5-Lightning-30B community ROCmFP4 requant (nemotron35-lightning-30b.yaml note) and muse-glimmer-30b -- third confirmed instance of this exact fork-dependency pattern for a ROCmFP4-labelled community artifact. RECOMMENDED-CONFIG DRAFT DEPTH, for the public-reply / upstream-filing use case this audit feeds: run_server.sh DRAFT_N default = 6, README's own "Deep Spec" row = n=7. Both exceed this lab's independently-measured n_max=3 hard ceiling for stock draft-mtp on the same architecture family (qwen38-27b). Whether ROCmFPX's fork (a materially different draft-mtp implementation from stock, per the above) reproduces, avoids, or differently exhibits the EOS-cliff at n_max 6-7 is NOT established by this audit and would require actually running their fork, which is out of scope.
clm-0070
- measured-here med ●●○ volatility medium · verified 2026-08-16
- On gfx1151, Ornith-1.0-35B Q8_0 is measured under the per-model, per-phase backend rule (clm-0050), and Vulkan wins outright here -- unlike laguna-s-21 (ROCm-favoured, clm-0065) but matching nemotron3-super and deepseek-v4-flash's Vulkan-favoured pattern: Vulkan leads decode at every measured depth (55.70 vs 47.66 t/s at d0, +17%; 46.22 vs 39.75 at d32768, +16%; 32.37 vs 27.48 at d131072, +18%) and prefill at d0/ d32768 (1063.7 vs 865.3, +23%; 679.2 vs 525.2, +29%), with ROCm only marginally ahead on prefill at the deepest cell (244.8 vs 238.5, +2.6%). Most notably, Vulkan did NOT lose the GPU device at d131072 on this candidate -- a clean break from the laguna-s-21/deepseek-v4-flash device-loss precedent (clm-0054/clm-0059/clm-0065) that has held on every other hybrid/MoE model tested at this depth on this box to date. Vulkan is the backend this bench serves the tau2 capability arm on.
- METHOD -- ornith-35b-fullbench matrix: stock llama.cpp 3653e6d (ROCm) / 3653e6d6d (Vulkan, prefix-matches per house convention), default f16 KV (matches the 2026-08-16 screen's serving config, cfg-0093), fa on, -ngl 999, --load-mode none, pp1024/tg256, llama-bench defaults otherwise (-b 2048/-ub 512, not passed explicitly). Output sanity gate (chat-endpoint coherence check) passed for both backends before any throughput was recorded. N=3 fresh-process reps at d0/d32768 (max CoV 0.9% rocm-d0-pp, 0.2% vulkan-d32768-pp -- well inside the 3% scatter-flag threshold, a healthy code path per protocol §0), N=1 at d131072 per protocol Tier1 / deepseek-v4-flash convention. pp1024 / tg256 by depth (mean of N): | depth | ROCm | Vulkan | vk/rocm pp | vk/rocm tg | |---|---|---|---|---| | 0 | 865.28 / 47.66 | 1063.66 / 55.70 | 1.229 | 1.169 | | 32,768 | 525.19 / 39.75 | 679.22 / 46.22 | 1.293 | 1.163 | | 131,072 | 244.84 / 27.48 | 238.54 / 32.37 | 0.974 | 1.178 | Vulkan's d131072 cell (run-0343/run-0344, N=1) needed only its first attempt out of the 2-attempt device-loss cap and produced a clean, coherent JSON result -- no vk::DeviceLostError, no GPU device wedge, no dmesg evidence of a ring reset in the window (checked via `sudo dmesg -T`, per the house rule fixing the non-sudo capture bug noted on clm-0065: this job's own on-failure dmesg calls are sudo-prefixed, though they were never triggered since no failure occurred). Whether this holds at contexts deeper than 131,072 is untested; d131072 is this bench's own matrix ceiling. No MTP/speculation on either backend: this job's own from-scratch GGUF header parser (no numpy/gguf-py dependency available on-box, so a minimal pure-Python GGUF v3 reader was written) read all 40 blocks / 733 tensor names directly and found zero nextn/mtp/ eagle/medusa/draft tensors -- plain decode throughout, matching the tau2 arm.
- evidence: run-0332 run-0333 run-0334 run-0335 run-0336 run-0337 run-0339 run-0340 run-0341 run-0342 run-0343 run-0344
clm-0071
- measured-here high ●●● volatility low · verified 2026-08-16
- Ornith-1.0-35B's KV cache costs 20.0 KiB per token on llama.cpp at f16 (default KV, ROCm), measured from a two-point GTT load probe (c=4096 -> 36,750,315,520 bytes gtt_used, c=131,072 -> 39,350,784,000 bytes gtt_used; delta 2,600,468,480 bytes over 126,976 tokens). This is SMALLER than laguna-s-21's 48.0 KiB/token, nemotron3-super's 8.00 KiB/token is still smaller, and deepseek-v4-flash's 7.13 KiB/token smaller still -- consistent with this candidate's hybrid architecture, where only 10 of 40 blocks (full_attention_interval=4) carry standard GQA attention (head_count_kv=2, key/value length 256) and the other 30 hold SSM/linear-attention state that does not grow with context. At the matrix's deepest tested depth (c=131,072), projected GTT use is ~36.6 GiB against the 120 GiB (122,880 MiB) boot window -- comfortably inside, with ~83 GiB headroom, and nowhere near the fit ceiling.
- METHOD: two-point GTT probe on stock 3653e6d ROCm, Ornith-1.0-35B Q8_0, --load-mode none, default f16 KV, gtt_used sampled from /sys/class/drm/card0/device/mem_info_gtt_used immediately after the server reported healthy, server torn down between points. RAW (kvprobe-raw.tsv, aihydra ~/bench-results/ornith-35b-fullbench/): c=4096 -> gtt_used_after=36,947,447,808 bytes; c=131,072 -> gtt_used_after=39,547,916,288 bytes; delta 2,600,468,480 bytes = 2,539,520 KiB / 126,976 tok = 20.00 KiB/token exactly. INDEPENDENTLY REPRODUCED same session, second probe run with a settling convention (kvprobe-c4096-redo.log / kvprobe-c131072-redo.log): c=4096 -> gtt_used_after_settle=36,750,315,520 bytes; c=131,072 -> gtt_used_after_settle=39,350,784,000 bytes. Absolute readings differ from the first probe by ~197 MiB at BOTH points (idle-baseline drift between the two probe runs, ~3 minutes apart), but the DELTA is bit-for-bit identical: 2,600,468,480 bytes in both runs. Confidence raised to high on this exact reproduction across two independent fresh-process probe pairs -- stronger evidentiary standard than clm-0060/clm-0066's single-pair method, since the confound (idle-baseline drift) is visible in the raw readings and demonstrably does not touch the delta. CROSS-CHECK: fixed (weights + graph) component at c=4096 = 36,947,447,808 bytes - (4,096 tok * 20.00 KiB/tok * 1024) = 36,947,447,808 - 83,886,080 = 36,863,561,728 bytes = 34.331 GiB, against the GGUF's own model_size (36,903,138,880 bytes = 34.371 GiB per the candidate record's acquisition sha256 verification) -- agrees to within 0.040 GiB (compute-graph + SSM-state overhead), tighter agreement than clm-0060's laguna cross-check. ARCHITECTURE BASIS for the small per-token cost: this job's own from-scratch GGUF header parser found qwen35moe.full_attention_interval=4 (10 of 40 blocks are full-attention: block ids 3,7,11,15,19,23,27,31,35,39; the other 30 carry ssm_a/ssm_alpha/ssm_beta/ssm_conv1d/ssm_dt.bias/ssm_norm/ssm_out tensors instead of attn_q/k/v/output). KV per token per full-attention layer at f16 = head_count_kv(2) * (key_length(256) + value_length(256)) * 2 bytes = 2048 bytes; x10 layers = 20,480 bytes/token = 20.00 KiB/token -- matches the measured delta exactly, a full first-principles reconciliation rather than just an order-of-magnitude check.
clm-0072
- measured-here med ●●○ volatility medium · verified 2026-08-16
- Ornith-1.0-35B Q8_0 completed the FULL standard 26-task tau2 airline set cleanly (rc=0, not a wall-bound cut) at 0.8846 mean reward (23/26), 202 tool-call messages, 0 empty assistant turns, 0 infrastructure errors -- VALID under the SMOKE gate. This is the NEW LEADER among full standard-26-task tau2 comparisons run on this box, ahead of deepseek-v4-flash's 0.846, nemotron3-super's 0.769, laguna-s-21's 0.6923 and qwen38-27b's 0.577. Wall time 4,193s (69.9 min) against a 28,800s (8h) safety ceiling never approached. Energy: 149.12 Wh across the run window (mean 128.0 W, delta 117.9 W over the 10.1 W idle floor), whole-session, joined same-session from HA history -- 6.4834 Wh per correct answer (0.1964 pence at 30.3 p/kWh), ALSO the lowest Wh-per-correct-answer of any full standard-26-task tau2 comparison run on this box: cheaper than laguna-s-21's 7.2588 Wh (clm-0067), deepseek-v4-flash's 24.29 Wh (clm-0061), nemotron3-super's 32.43 Wh, and qwen38-27b's 38.29 Wh (clm-0057) -- achieved on Vulkan (this bench's own throughput-matrix pick, clm-0070) at plain decode (no MTP/speculation exists in this GGUF, clm-0070's own tensor scan) against a raw ~46.2 t/s decode floor at the serving depth (d32768) -- both the highest capability score AND the lowest energy-per-correct-answer of any full-26-task comparison on this box, a joint win this programme has not seen before (every prior leader on one metric has trailed on the other).
- METHOD -- tau2-bench airline, harness 668d3bcd135c02aa3438f987ef45735b7c163ee3 (same version as qwen38-27b/deepseek-v4-flash/nemotron3-super/laguna-s-21's full-bench arms), --num-trials 1 --seed 42 --max-steps 200 --max-concurrency 1, user simulator pinned per protocol (openrouter/anthropic/claude-haiku-4.5, temperature 0). Agent: local llama-server endpoint on cfg-0095 (Vulkan, default f16 KV, -c 32768, --parallel 1, --load-mode none, no speculation), agent sampling temperature 0.0 / max_tokens 4096. Task-ids 0-25 explicit (the standard 26-task set). Guarded by run-0338 (house 4-item guard, 4/4, against this exact serving config -- Vulkan, not the ROCm config the screen's run-0330 guard covered -- run fresh immediately before this arm, since backend is itself the `binary` lever this guard exists to catch on a config swap). All numbers computed directly from raw artifacts (results.json, run-meta.jsonl, tier3-armcompare.py's own arm summary), no delegate/draft to reconcile against for this job. FAILED TASKS (3 of 26, reward 0.0): task 7 (11 tc, 156.2s), task 14 (8 tc, 294.7s), task 23 (11 tc, 332.4s -- the longest task in the run, and it succeeded on retry 1 per the harness's own retry log before landing at reward 0.0 on the scored attempt). Every failure resolved via user_stop -- the simulator ending the conversation judging the task unresolved, not a harness timeout or agent error. cut_at_max_steps: 0 across all 26 tasks. EFFICIENCY-LEADER CLAIM SCREENED AGAINST EVERY OTHER MODEL PAGE'S wh_correct_energy BEFORE PUBLISHING, per house policy: laguna-s-21 eng-0124 = 7.2588 Wh/correct, deepseek-v4-flash eng-0107 = 24.29 Wh/correct, nemotron3-super eng-0116 = 32.43 Wh/correct, qwen38-27b eng-0081 = 38.29 Wh/correct -- ornith-35b leads all four on the SAME protocol (full standard 26-task set, pinned haiku-4.5 simulator, same energy-join method, same idle baseline). ONE LOWER NUMBER EXISTS ON THE SITE: qwen35-122b's eng-0034 records 5.713 Wh/correct, but that run is a 5-task SMOKE window predating the standard-26-task convention -- a different scale on a different protocol, not a like-for-like comparison, and this claim does not assert an unqualified site-wide superlative on the strength of it (same scoping laguna-s-21's clm-0067 already established). Speculation/MTP: not applicable -- this job's own from-scratch GGUF header parser found zero nextn/mtp/eagle/medusa/draft tensors among all 733 tensors in the Q8_0 GGUF (clm-0070). Ornith-1.0-35B has no speculative-decode path exercised in this bench -- this candidate's dual lead (capability AND energy) is achieved at a raw plain-decode floor (46.22 t/s at the serving depth, Vulkan), the highest of any full-bench candidate's serving-depth decode figure in this programme to date. CAVEAT ON EVERY CROSS-MODEL COMPARISON ABOVE -- none of these are controlled A/B: different models, different architectures, quant tiers (Q8_0 here vs Q4_K_M/IQ3_XXS/ IQ4_XS/UD-Q4_K_XL for the others -- a HIGHER-precision quant than every other full-bench candidate in this programme, per protocol §9's "higher quant does not fit, so test lower" grey-zone logic run in reverse: this candidate specifically fits Q8_0 comfortably), and different backends (Vulkan here and for nemotron3-super; ROCm for deepseek-v4-flash and laguna-s-21). Reported because it is the only same-shape (26-task, same simulator, same protocol, same energy-join method) set of data points available in this programme, not because the confound list has been controlled for -- and the quant-tier gap in particular means this is NOT presented as evidence that Ornith-1.0-35B is intrinsically more capable or efficient than the other candidates at matched precision.
- evidence: run-0345
clm-0073
- measured-here high ●●● volatility low · verified 2026-08-16
- The staged qwen36-27b-mtp GGUF (unsloth/Qwen3.6-27B-GGUF, Q4_K_M) is NOT the uniform-dense-attention, MTP-capable model this candidate's own record assumed. A from-scratch GGUF header parse (851 tensors, every metadata KV) found general.architecture=qwen35, qwen35.full_attention_interval=4: 16 of 64 blocks carry standard full-attention tensors (head_count_kv=4, key/value length 256), the other 48 carry ssm_a/ssm_alpha/ssm_beta/ssm_conv1d/ssm_dt.bias/ssm_norm/ssm_out — a HYBRID architecture (same shape class as ornith-35b and the production 122B family), not the uniform dense-attention case the candidate's why_listed rationale argued MTP's larger speed-up would apply to. FFN carries no "_exps" tensors, so "27B dense" (26.9B total params, confirmed by llama-bench's own model_n_params field) remains correct about MoE routing — the correction is about attention uniformity only. Separately: zero MTP/speculation tensors exist in this artifact. An exhaustive needle scan of all 851 tensor names found no nextn/mtp/eagle/medusa/draft tensor anywhere, and this was confirmed empirically, not just by absence: a live llama-server load with --spec-type draft-mtp --spec-draft-n-max 3 refuses to start on BOTH the ROCm and Vulkan binaries with the identical error — "context type MTP requested but model doesn't contain MTP layers" / "failed to create MTP context". The gguf's own general.base_model.0.repo_url metadata field is https://huggingface.co/Qwen/Qwen3.6-27B (no "-MTP" suffix) — this staged artifact is the BASE Qwen3.6-27B checkpoint, not the MTP-head variant the candidate's display name and listing rationale assumed. No spec/MTP arms exist in this bench because there is no speculation path to measure.
- METHOD: gguf_probe2.py, a from-scratch GGUF metadata + tensor-name parser with no numpy/gguf-py dependency, run against ~/models/qwen36-27b/Qwen3.6-27B-Q4_K_M.gguf on aihydra. Every metadata KV dumped unfiltered; all 851 tensor names scanned for ssm/mamba/conv1d/hybrid/linear-attention AND nextn/mtp/draft/eagle/medusa/a_log/dt_bias/x_proj/in_proj/out_proj markers. Per-block tensor schema grouped into exactly 2 classes across all 64 blocks, matching qwen35.full_attention_interval=4 exactly (blocks 3,7,11,15,19,23,27,31,35,39,43,47, 51,55,59,63 = 16 of 64 are full-attention). LIVE CONFIRMATION: a direct llama-server invocation on both binaries (~/src/llama.cpp/build/bin, ROCm, and ~/src/llama.cpp-vk3653e6d/build/bin, Vulkan), identical flags except -dev, with --spec-type draft-mtp --spec-draft-n-max 3 added. Both processes fail to load with the same log line: "llama_init_from_model: context type MTP requested but model doesn't contain MTP layers" / "common_speculative_init_result: failed to create MTP context" / "load_model: failed to create MTP context" — server exits, never becomes healthy. Raw logs preserved: aihydra ~/bench-results/qwen36-27b-mtp-fullbench/ mtp-attempt-rocm.log, mtp-attempt-vulkan.log. This closes a real gap in the candidate's own record: its 2026-06-20 "listed" and 2026-06-27/2026-08-13 "screened" entries all assumed dense-uniform attention plus an MTP path, neither of which the staged artifact actually has. The 2026-08-13 screen's own SMOKE result (0.80 mean, joint-strongest on this box at the time) is unaffected by this correction — it was measured against the real artifact regardless of what the candidate record believed about its architecture — but the "dense architecture — the case where MTP's larger dense speed-up applies" why_listed rationale is now known to not describe what was actually screened.
clm-0074
- measured-here high ●●● volatility low · verified 2026-08-16
- On qwen36-27b-mtp Q4_K_M, ROCm wins prefill decisively at every tested depth over Vulkan (357.8 vs 302.7 t/s at d0, +18%; 210.5 vs 95.7 t/s at d32768, +120%) and is the ONLY backend that completed the deepest matrix cell (d131072) without a device-loss event — Vulkan device-lost on BOTH allowed attempts at that depth (vk::DeviceLostError, GPU wedged, recovered via kernel amdgpu ring reset each time; ROCm completed cleanly on its first attempt at 95.4 pp / 8.44 tg t/s). Decode is roughly matched between backends, Vulkan marginally ahead (12.83 vs 12.08 t/s at d0, +6%; 11.33 vs 10.86 t/s at d32768, +4%). ROCm is this bench's serving-config pick on both grounds: it is faster where it matters most (prefill, and by a wide margin at depth) and it is the only backend that reaches the full depth range this candidate's 262K native context makes plausible to serve at.
- METHOD: llama-bench, stock 3653e6d (ROCm) / 3653e6d6d (Vulkan), explicit f16 KV, -fa 1, --load-mode none, pp1024/tg256. N=3 fresh-process reps at d0/d32768 (max CV 1.6% on ROCm prefill, 3.0% on Vulkan prefill at d32768 — at the house 3% flag threshold, but decode at the same depth was clean at 0.03%, read as ordinary prefill scatter rather than a defective-path signature), N=1 at d131072 per house convention. DEVICE-LOSS EVIDENCE: dmesg -T (sudo) captured for both Vulkan d131072 attempts — "amdgpu 0000:c5:00.0: ring comp_1.2.0 timeout" / "Starting comp_1.2.0 ring reset" / "Ring comp_1.2.0 reset succeeded" / "device wedged, but recovered through reset" — both attempts hit the identical failure signature and both recovered cleanly (no reboot needed). This matches the laguna-s-21/deepseek-v4-flash device-loss precedent at comparable depth on this silicon; ornith-35b broke that streak once (its own Vulkan arm completed d131072 cleanly), this candidate does not — the streak-break was not a general fix, it was ornith-specific. No MTP/speculation confound: this candidate's staged GGUF carries no MTP tensors at all (clm-0073), so every cell in this matrix is plain decode on both backends — the backend comparison is not entangled with a speculation lever.
- evidence: run-0346 run-0347 run-0348 run-0349 run-0350 run-0351 run-0352 run-0353 run-0354 run-0355 run-0356
clm-0075
- measured-here high ●●● volatility low · verified 2026-08-16
- qwen36-27b-mtp's KV cache costs 64.00 KiB per token on llama.cpp at f16 (explicit KV, ROCm), measured from a two-point GTT load probe (c=4096 -> 16,821,088,256 bytes gtt_used, c=131,072 -> 25,142,587,392 bytes gtt_used; delta 8,321,499,136 bytes over 126,976 tokens = exactly 64.00 KiB/token). This is LARGER than every other hybrid-architecture full-bench candidate on this board (ornith-35b's 20.00 KiB/token, nemotron3-super's 8.00, deepseek-v4-flash's 7.13), consistent with this candidate carrying MORE full-attention layers than any of them: 16 of 64 blocks (25%, full_attention_interval=4), each with head_count_kv=4 (double ornith-35b's 2), key/value length 256. At the matrix's deepest tested depth (c=131,072), projected GTT use is ~23.4 GiB against the 120 GiB (122,880 MiB) boot window — comfortably inside, with ~97 GiB headroom to spare; fit was never this candidate's constraint.
- METHOD: two-point GTT probe on stock 3653e6d ROCm, qwen36-27b-mtp Q4_K_M, --load-mode none, explicit f16 KV (-ctk f16 -ctv f16), gtt_used sampled from /sys/class/drm/card0/device/mem_info_gtt_used 10s after the server reported healthy (settling convention), server torn down between points. RAW (kvprobe-c4096.log / kvprobe-c131072.log, aihydra ~/bench-results/qwen36-27b-mtp-fullbench/): c=4096 -> gtt_used_after_settle=16,821,088,256 bytes (15.6621 GiB); c=131,072 -> gtt_used_after_settle=25,142,587,392 bytes (23.4159 GiB); delta 8,321,499,136 bytes = 8,126,460.875 KiB / 126,976 tok = 64.0006 KiB/token, i.e. exactly 64 KiB/token to four significant figures. FIRST-PRINCIPLES RECONCILIATION, TWO WAYS. (1) Fixed component: the c=4096 reading (15.6621 GiB) sits within 3.8 MB of the GGUF's own on-disk file size (16,817,244,384 bytes = 15.6584 GiB, llama-bench's own model_size field) — the tightest base-weight agreement recorded in this programme to date (ornith-35b's equivalent check agreed to within 0.040 GiB; this one to within 0.0037 GiB), consistent with this candidate having very little fixed compute-graph or SSM-state overhead relative to raw weight size. (2) Per-token slope from architecture math: KV per token per full-attention layer at f16 = head_count_kv(4) * (key_length(256) + value_length(256)) * 2 bytes = 4,096 bytes; x16 full-attention layers (qwen35.full_attention_interval=4 = 16 of 64 blocks, clm-0073) = 65,536 bytes/token = 64.00 KiB/token — matches the measured GTT delta EXACTLY, not just to an order of magnitude. Not independently re-probed a second time this session (unlike ornith-35b's bit-for-bit reproduction across two probe pairs) — confidence is rated high on the strength of the double first-principles reconciliation (file-size match AND architecture-math match) rather than repro alone; a second probe pair would raise it further but was not run given this candidate's much larger time cost elsewhere in the session (tau2 alone ran 11.3h).
clm-0076
- measured-here high ●●● volatility low · verified 2026-08-17
- qwen36-27b-mtp is the WORST full-bench candidate on this board by energy efficiency, by a wide margin, despite passing the cheap house guard twice (4/4, no cliff) and a directionally-consistent SMOKE screen (0.80 mean, 2026-08-13). The standard 26-task tau2 airline set could NOT be completed in one 8h session — rc=124 at 21/26 task attempts (20 scored, one an unrecoverable infrastructure_error with zero messages) — and needed a same-build/same-config CONTINUATION session to finish the remaining 5 tasks. COMBINED across both sessions: 25 of 26 tasks scored (task 15 excluded as an infrastructure_error, not a 0), 16 passed, mean 0.6400 — above qwen38-27b's 0.577 but well below deepseek-v4-flash (0.846), nemotron3-super (0.769), laguna-s-21 (0.6923) and ornith-35b (0.8846). This is NOT a clean full-26-task completion like every other benched candidate on this board and that caveat travels with every comparison above. The energy figure is the decisive finding. Combined tau2 wall time across both sessions was 40,836s (11.34h) for the SAME 26-task set ornith-35b completed in 4,193s (70 min) — roughly 10x longer — driven by this candidate's raw ~11-12 t/s decode floor (a third of ornith-35b's ~46 t/s, clm-0074) COMPOUNDED by repeated harness-level retries from a malformed-output failure class this session observed directly ("AssistantMessage must have either content or tool_calls", zero content and zero tool_calls in one turn). Combined energy: 2055.19 Wh across both windows (whole-session, same wall-meter method and idle baseline as every other board entry) = 128.45 Wh per correct answer. That is more than 3x qwen38-27b's previous-worst 38.29 Wh/correct, and roughly 20x ornith-35b's leading 6.48 Wh/correct — the widest efficiency gap this programme has measured between two full-bench candidates on the same protocol.
- METHOD — tau2-bench airline, harness 668d3bcd135c02aa3438f987ef45735b7c163ee3 (same version as every other full-bench arm on this board), --num-trials 1 --seed 42 --max-steps 200 --max-concurrency 1, user simulator pinned per protocol (openrouter/anthropic/claude-haiku-4.5, temperature 0). Agent: local llama-server endpoint on cfg-0099 (ROCm, explicit f16 KV, -c 32768, --parallel 1, --load-mode none, no speculation — none exists in this GGUF, clm-0073), agent sampling temperature 0.0 / max_tokens 4096. Guarded twice, fresh each session (run-0357, run-0359), both 4/4 — no cliff detected on either check. ARM 1 (run-0358): task-ids 0-25 requested, 21 attempted before the 8h TAU2_TIMEOUT safety ceiling (rc=124). Task 15's final attempt returned reward=null, termination_reason=infrastructure_error, 0 messages — excluded from tasks_total per the SMOKE validity gate's own logic (a zero-message result cannot be scored either way), not counted as a pass or a fail. 20 scored, 14 passed, mean 0.7000. ARM 2 (run-0360): task-ids 21-25, same build/config/seed as arm 1, run in a fresh session after arm 1's timeout, house guard re-run fresh first (protocol §9's "split across sessions" safe lever — same config, bookkeeping only). rc=0, all 5 completed. 2 passed, 3 failed, mean 0.4000. COMBINED: (14+2)/(20+5) = 16/25 = 0.6400 mean reward across the 25 scored tasks. RETRY PATTERN: at least 3 distinct tasks (6, 7, 8, 10, 14, 21, 23, 25 across both arms — 8 of the 25 scored/attempted tasks, 32%) needed one or more internal harness retries before landing on a scored attempt, each retry re-running a multi-turn conversation from the start and multiplying wall time. Task 14 needed 3 retries before scoring 0.0; task 15's final (unretryable) attempt failed outright; task 23 (arm 2) accumulated roughly 90 minutes of cumulative wall time across attempts before scoring 0.0 — the single most expensive task this bench measured across every candidate to date. The malformed-output signature ("AssistantMessage must have either content or tool_calls. Got AssistantMessage") indicates the model intermittently emits a fully empty turn (no text, no tool call) at this depth/config — a genuine failure mode distinct from, and additional to, the architecture/MTP corrections in clm-0073. ENERGY, computed the same way as every other board entry: wall-meter counter-difference on sensor.hardware_ai_hydra_energy (aihydra, Home Assistant), raw ~10s-resolution history pulled fresh this session and linearly interpolated to each run-meta window, against the 10.1 W box-idle-all-empty baseline. Arm 1 (eng-0141): 1446.5888 Wh over 28,800s (mean 180.82 W, delta 170.72 W), 103.3278 Wh/correct on its own 14-passed denominator. Arm 2 (eng-0143): 608.6042 Wh over 12,036s (mean 182.04 W, delta 171.94 W), 304.3021 Wh/correct on its own thin 2-passed denominator (not representative alone — the COMBINED figure is the one that belongs on the model page). Combined: 1446.5888 + 608.6042 = 2055.1930 Wh / 16 correct = 128.4496 Wh/correct, 3.8920 pence at the standing 30.3 p/kWh tariff. EFFICIENCY-COMPARISON SCREENED AGAINST EVERY OTHER BENCHED MODEL PAGE'S wh_correct_energy BEFORE PUBLISHING, per house policy: ornith-35b eng-0134 = 6.4834 Wh/correct, laguna-s-21 eng-0124 = 7.2588, deepseek-v4-flash eng-0107 = 24.29, nemotron3-super eng-0116 = 32.43, qwen38-27b eng-0081 = 38.29 — qwen36-27b-mtp's 128.45 Wh/correct trails all five, and by a materially wider margin than separates any adjacent pair among them (the next-widest gap, qwen38-27b to nemotron3-super, is ~1.2x; this candidate to qwen38-27b is >3.3x). CAVEAT ON THE CROSS-MODEL COMPARISON: not a controlled A/B — different models, different quant tiers, different backends (ROCm here; Vulkan for ornith-35b and nemotron3-super, ROCm for deepseek-v4-flash and laguna-s-21). Reported because it is the only same-shape (26-task set attempted, same simulator, same protocol, same energy-join method) set of data points available in this programme. The BOUND-CUT status of this run (25/26 scored rather than a clean 26/26) is the one confound specific to this candidate among the six, and is why every number above states its own denominator explicitly rather than assuming "26" the way the other five pages can.
- evidence: run-0357 run-0358 run-0359 run-0360 eng-0141 eng-0143 eng-0144
clm-0077
- measured-here high ●●● volatility low · verified 2026-08-17
- Deep-Thought-Posttrain's KV cache costs a measured 40.00 KiB/token (f16 KV, ROCm and Vulkan identical — both default to the same type_k/type_v) — reproduced exactly from first principles against the GGUF's own attention metadata (2 x 32 layers x 5 KV heads x 64 head_dim x 2 bytes) and cross-checked against a real two-point GTT probe (c=2048 vs c=8192, /sys/class/drm/card0/device/mem_info_gtt_used readings before and after each load): (1,258,139,648 - 1,006,481,408) bytes / (8192-2048) tokens = 40,960 bytes/token exactly. At this candidate's full native context (8192 tokens, the ceiling — not a chosen serving depth), total KV footprint is 0.3125 GiB against a 0.6740 GiB weights footprint on a 120 GiB pool: this candidate poses no realistic fit or OOM risk on this hardware at any context it can actually use, which is a genuinely uninteresting finding in isolation, but sets the honest baseline the site's larger candidates are compared against.
- Two independent instruments agree exactly (empirical sysfs delta vs first-principles GGUF-metadata arithmetic), which is the highest-confidence pattern this lab's KV-cost claims can reach — same standard as ornith-35b's clm-0071.
clm-0078
- measured-here high ●●● volatility low · verified 2026-08-17
- Deep-Thought-Posttrain's full 26-task tau2 airline run (run-0372) recorded tool_call_messages = 0 across all 26 simulations — the same SMOKE validity gate (protocol.json comparison_rules, added 2026-08-14) that flagged npu-lfm2's vacuous 1.00 mean reward as invalid on this site's first NPU screen.
- Digit-backing claim for the model page's compare caveat (site-design-v2 §7: no number appears in rendered prose without a cited claim containing it verbatim).
- evidence: run-0372
clm-0079
- measured-here high ●●● volatility low · verified 2026-08-17
- Qwen3-Coder-Next 80B (qwen3next hybrid, 512 experts/10 active + 1 shared, ~3B active) Q8_0's own throughput matrix has Vulkan leading decode at EVERY depth measured (44.13/37.0/26.4/21.73/19.15 t/s at d0/d32768/d131072/d204800/d262144 vs ROCm's 37.66/31.06/21.29/17.51/15.46) and prefill at the two shallowest cells (614.0/450.4 vs 476.6/346.8 t/s at d0/d32768) -- but the pattern INVERTS at the two deepest cells: ROCm overtakes prefill at d204800 (124.6 vs 98.75 t/s, Vulkan -20.7%) and by a wider margin at d262144, this candidate's model-max (104.56 vs 68.77 t/s, Vulkan -34.2%). Vulkan lost the GPU device zero times across all 10 matrix cells (both backends, all 5 depths, 2-attempt cap never triggered) -- a clean streak matching ornith-35b's precedent-breaking result and extending it to this programme's largest full-bench candidate so far (80B total, vs ornith's 35B). Decode remains the dominant real-world driver of interactive throughput, so Vulkan is the backend this bench serves the capability arm on, same choice as ornith-35b/nemotron3-super/deepseek-v4-flash despite the deep-prefill inversion this candidate newly shows.
- Full matrix: cfg-0112 (ROCm, stock 3653e6d) and cfg-0113 (Vulkan, stock 3653e6d6d), both Q8_0, default f16 KV, -fa 1, --load-mode none, N=3 fresh-process reps at d0/d32768, N=1 at d131072/d204800/d262144 per protocol.json's throughput_ladder. Guard chain wired from birth: every performance run in both series carries `guard:` pointing at its backend's matching 4/4 house guard (run-0394 ROCm, run-0395 Vulkan -- the tau2-serving-config guard, since a backend swap is the `binary` lever the guard exists to catch). Scatter checked at every cell via sweep.sh's own CV gate; no cell exceeded ~1% run-to-run variance where N=3 applied. The deep-prefill inversion (ROCm ahead past d204800) is a genuine new finding for this programme -- every prior full-bench candidate that survived Vulkan to comparable depth (ornith-35b) kept Vulkan ahead on prefill too, or lost the device entirely (laguna-s-21, deepseek-v4-flash). This candidate is the first to show Vulkan surviving cleanly AND losing the prefill lead at depth, a distinct outcome from either precedent.
- evidence: run-0396 run-0397 run-0406 run-0407 run-0414 run-0415 run-0424 run-0425
clm-0080
- measured-here high ●●● volatility low · verified 2026-08-17
- Qwen3-Coder-Next 80B's KV cache costs 24 KiB/token (f16 KV), computed from the GGUF's own architecture metadata rather than a live GTT probe: qwen3next. full_attention_interval=4 over 48 blocks means 12 full-attention layers (GQA, attention.head_count_kv=2, attention.key_length=attention.value_length=256), each costing 2 kv_heads x 256 head_dim x 2(K+V) x 2 bytes(f16) = 2,048 bytes/token, and the other 36 blocks are gated-DeltaNet-class SSM/linear- attention with NO growing KV cache at all. This job's own live FIT probe at c=32768 measured 81,662 MiB GTT (delta from empty ~81,644 MiB against the 84,812,055,968-byte / 78.99 GiB Q8_0 weight footprint), consistent with the metadata-derived KV rate to within measurement noise (weights + ~0.75 GiB KV at c=32768 against the observed delta). Projected GTT use at this candidate's own model-max (c=262,144) is ~85 GiB against the 120 GiB boot window -- comfortably inside, ~35 GiB of headroom to spare, never a fit risk at any depth this bench measured.
- Unlike ornith-35b's two-point live-probe KV measurement (kvprobe-ornith-redo.sh, two fresh-process FIT-only loads at different contexts), this candidate's KV cost was derived from the GGUF's own metadata (full_attention_interval, head_count_kv, key/value length) and cross-checked against the single FIT probe this job's screen already captured (c=32768, 24s load, 81,662 MiB GTT) -- a live two-point GTT delta probe (e.g. c=4096 vs c=131072) was not run separately this session; the metadata-derived rate is the primary basis for this claim, with the single live point as a consistency check rather than an independent confirmation. Same architecture family and per-layer KV shape (head_dim 256, 2 kv_heads) as ornith-35b (20 KiB/token, 10 full-attention layers) and the production 122B -- this candidate is 48 blocks / 12 full-attention vs ornith's 40 blocks / 10, so the per-token rate scales exactly with the extra 2 full-attention layers (12 x 2048 = 24,576 bytes vs ornith's 10 x 2048 = 20,480 bytes).
- evidence: run-0394
clm-0081
- measured-here high ●●● volatility low · verified 2026-08-17
- Qwen3-Coder-Next 80B ships NO speculative-decode or MTP head in its official Q8_0 GGUF. This job's own from-scratch GGUF header parser (gguf_probe2.py, no numpy/gguf-py dependency) read all 807 tensors across all 4 shards directly and ran an exhaustive needle scan for nextn/mtp/eagle/medusa/draft on every tensor name: ZERO hits on every needle, on every shard. The base model's own config.json independently corroborates this before any download: standard Qwen3NextForCausalLM causal-LM config, no nextn/draft-head keys of any kind. Plain decode throughout this bench -- no spec-decode arm was attempted or is possible against this artifact.
- Checked BEFORE claiming anything MTP-related, per this lab's own standing caution (burned once by a mislabeled non-MTP artifact): the acquired-gate candidate record (content/candidates/qwen3-coder-next-80b.yaml) ran this exact scan on 2026-08-17 immediately post-download, across all 4 shards independently (not just the first), before any bench work began. Matches the same house-standard verification ornith-35b (clm-0070) and qwen36-27b-mtp's candidate record both carry. No EOS-cliff watch applies here (that hazard class is specific to models WITH an MTP head run past spec-draft-n-max >= 4; this candidate has no MTP head to watch).
clm-0082
- measured-here med ●●○ volatility medium · verified 2026-08-17
- Qwen3-Coder-Next 80B Q8_0 completed the FULL standard 26-task tau2 AIRLINE set cleanly (rc=0, not a wall-bound cut) at 0.5385 mean reward (14/26), 192 tool-call messages, 390 assistant messages, 0 empty assistant turns, 0 infrastructure errors -- VALID under the SMOKE gate, served on Vulkan (this bench's own throughput-matrix pick, clm-0079) at plain decode (no MTP/ speculation exists in this GGUF, clm-0081). This is the LOWEST capability score of any full standard-26-task tau2 comparison run on this box, below qwen38-27b's 0.577 -- expected and unsurprising: this is a CODER model, purpose-built and marketed for agentic coding workflows (Terminal-Bench, SWE- bench-style tasks), being scored on tau2's AIRLINE customer-service domain, not its home domain. The airline benchmark is this programme's cross-model comparability instrument -- the same 26 tasks, same pinned simulator, same protocol every other full-bench candidate on this board was scored against -- and that comparability is exactly why the number is worth publishing, but it should be read as "how a coding specialist performs on an out-of-domain agentic task," not as this model's ceiling on the work it is actually built for. Energy tells the sharper story: 67.054 Wh across the run window (mean 103.69 W, delta 93.59 W over the 10.1 W idle floor), whole-session -- 4.7896 Wh per correct answer (0.1451 pence at 30.3 p/kWh), the NEW LOWEST Wh-per-correct- answer of any full standard-26-task tau2 comparison on this board, undercutting ornith-35b's previous-leading 6.4834 Wh by 26%. This candidate therefore holds the LOWEST capability score AND the LOWEST energy-per-correct-answer of this programme's full-bench series SIMULTANEOUSLY -- the inverse of ornith-35b's joint-highest result, and a genuinely new combination this board has not shown before. The efficiency figure is real and measured on the identical protocol as every other comparator, but it is not a capability win, and the verdict does not present it as one.
- METHOD -- tau2-bench airline, harness 668d3bcd135c02aa3438f987ef45735b7c163ee3 (same version as every prior full-bench arm on this board), --num-trials 1 --seed 42 --max-steps 200 --max-concurrency 1, user simulator pinned per protocol (openrouter/anthropic/claude-haiku-4.5, temperature 0). Agent: local llama-server endpoint on cfg-0111 (Vulkan, default f16 KV, -c 32768, --parallel 1, --load-mode none, no speculation), agent sampling temperature 0.0 / max_tokens 4096. Task-ids 0-25 explicit (the standard 26-task set). Guarded by run-0395 (house 4-item guard, 4/4, against this exact serving config, run fresh immediately before this arm since backend is the `binary` lever the guard exists to catch on a config swap). A 5-task SMOKE screen (run-0432, ROCm, cfg-0110) preceded this arm and scored 0.800 (4/5) -- directionally consistent (below the board's higher scorers, not a screen/full-bench reversal), 2.4205 Wh/correct on its own 5-task window (eng-0167, NOT a like-for-like comparator against the standard-26-task figure above, same scoping every prior candidate's smoke window carries). FAILED TASKS (12 of 26, reward 0.0): tasks 0, 7, 8, 10, 11, 14, 15, 18, 20, 21, 23, 24. No single shared pattern in tool-call count or duration is visible from this bench's own arm summary alone (durations range 44s-152s across both passed and failed tasks); a transcript read would be needed to characterise the failure mode further, and this claim does not speculate beyond the aggregate. Every task resolved via user_stop (the simulator ending the conversation, judging it unresolved) -- 0 tasks cut at max_steps across all 26. EFFICIENCY-LEADER CLAIM SCREENED AGAINST EVERY OTHER MODEL PAGE'S wh_correct_energy BEFORE PUBLISHING, per house policy: ornith-35b eng-0134 = 6.4834 Wh/correct, laguna-s-21 eng-0124 = 7.2588 Wh/correct, deepseek-v4-flash eng-0107 = 24.29 Wh/correct, nemotron3-super eng-0116 = 32.43 Wh/correct, qwen38-27b eng-0081 = 38.29 Wh/correct -- this candidate leads all five on the SAME protocol (full standard 26-task set, pinned haiku-4.5 simulator, same energy-join method, same idle baseline). ONE LOWER NUMBER EXISTS ON THE SITE: qwen35-122b's eng-0034 records 5.713 Wh/correct, but that run is a 5-task SMOKE window predating the standard-26-task convention -- a different scale on a different protocol, not a like-for-like comparison, and this claim does not assert an unqualified site-wide superlative on the strength of it (same scoping every prior leader's claim already established). CAVEAT ON EVERY CROSS-MODEL COMPARISON ABOVE -- none of these are controlled A/B: different models, different architectures, quant tiers (Q8_0 here vs Q4_K_M/IQ3_XXS/IQ4_XS/UD-Q4_K_XL for the others), and different backends (Vulkan here and for nemotron3-super/ornith-35b; ROCm for deepseek-v4-flash and laguna-s-21). The lower Wh/correct here is plausibly explained in large part by the much lower absolute token cost of a lower-reward run (fewer turns needed to reach a user_stop on a task the model does not solve is not the same as efficient problem-solving) -- this claim does NOT assert this candidate is intrinsically more energy-efficient PER UNIT OF CAPABILITY than ornith-35b or any other comparator; it reports the measured Wh-per-correct-answer figure on the standard protocol, with the capability-score caveat stated in the same breath, exactly as required.
- evidence: run-0433 run-0432
clm-0083
- measured-here med ●●○ volatility medium · verified 2026-08-18
- Ling-3.0-flash's corrected post-PR-26608 Q4_K_M GGUF completed a valid full 26-task tau2 airline run on aihydra at 0.500 mean reward (13/26), with 170 tool-call messages, 0 empty assistant messages, 0 infrastructure errors and 0 max-step cuts. The same ROCm/plain-decode config passed the house 4-item guard and measured 36.08 t/s decode at d0, 32.89 t/s at d32768. Energy was 69.218 Wh for the whole tau2 window, or 5.32 Wh per correct answer at 30.3 p/kWh.
- METHOD -- Ling-3.0-flash Q4_K_M from bloomer010/Ling-3.0-flash-GGUF, sha256 bcce6e32799749e8e52a52c161127b989db52430e00fbebc77f4d61ef94e754d, on llama.cpp post-BailingMoE3 build 7077abb. ROCm only, f16 KV, --load-mode none, -fa on, --parallel 1, --jinja, temperature 0, max_steps 200, seed 42, pinned user simulator openrouter/anthropic/claude-haiku-4.5. No speculative decoding or MTP path was exercised; the candidate's shipped MTP activation state remains a separate open question. Energy joined same session from raw HA sensor.hardware_ai_hydra_energy history, interpolated to run-meta windows.
- evidence: run-0435 run-0436 run-0437 run-0438 run-0439 eng-0173
clm-0084
- measured-here med ●●○ volatility medium · verified 2026-08-18
- On gfx1151 with the same DeepSeek-V4-Flash-0731 UD-IQ3_XXS artifact and f16 KV, the carried Strix Halo Vulkan fork at baf6360be passed the house 4-item guard and completed every bounded llama-bench cell through d262144 without a device loss. Mean prefill/decode was 219.36/19.54 t/s at d0, 181.55/17.83 at d32768, 142.56/15.41 at d131072, 126.50/14.24 at d204800 and 115.04/13.32 at d262144. This directly supersedes the stock-build-only conclusion in clm-0059 that Vulkan could not survive beyond d0, but it does not prove Vulkan generally safe: the fork bundles multiple kernel, prefill, DeepSeek and newer-upstream changes, and only a binary guard—not the full capability suite—ran on this configuration.
- METHOD -- Nathanw1014/strix-halo-llamacpp portable Vulkan payload baf6360be (build 10565), staged with its pinned Mesa 26.3.0-devel RADV driver and documented GGML_VK tuning environment. Same 97 GiB DeepSeek artifact as the earlier stock matrix; f16 K/V, flash attention, --load-mode none, pp1024, tg256, batch 2048, ubatch 512 and 16 threads. N=3 fresh processes at d0/d32768/d131072; N=1 at d204800/d262144. All artifact rows report rc=0 and build baf6360be. Throughput windows carry wall-meter Wh totals in eng-0179 through eng-0183, never Wh per correct answer. Artifact provenance: screen timestamps.jsonl sha256 804f0f4bb2f6cdfdeb8f69411e5e1abce8d1a5026c347c68bfebfc490392c31c, status.tsv b581498c41398cf7da92c4b36d6828e26942518a42a1043702aa4490099a13a4; deep timestamps.jsonl 20f34534fbb77f26e2ead95c856e5e8b251ce083d49631ee50cb9aa380281c30, status.tsv 55bdafeb5651d9170536099d712ac4711a8544b862b599531e6712107b72a836.
- evidence: run-0440 run-0441 run-0442 run-0443 run-0444 run-0445 run-0446 run-0447 run-0448 run-0449 run-0450
clm-0085
- measured-here med ●●○ volatility medium · verified 2026-08-18
- Gemma 4 26B-A4B UD-Q4_K_M did not earn a reflex-tier promotion on HG-001: the corrected ROCm stock run passed coherence, native tool-call arguments and the parallel-1 isolation skip, but failed the house needle retrieval at the first declared depth, 2048 tokens, for a 3/4 guard result. The predeclared stop gate fired before tau2 smoke, throughput, Vulkan testing or energy measurement, so no performance number or aihydra capability claim is admissible for this candidate. Existing evidence does not isolate a runtime crash or backend/device fault: the model loaded and answered the other guard probes, and the failing artifact is a retrieval/probe check returning tool-call text instead of the needle. Treat the rejection as a measured retrieval-envelope/probe failure, not as a general model-quality verdict.
- Evidence source: HG-001 job card and aihydra ~/bench-results/gemma4-26b-promotion-20260818T220904Z/{status.tsv,queue.log, guard-rocm-d2048.txt}. Status shows artifact/build preflight OK (bytes=16947541728, sha256=f2c28b3dc4776931ac6f879e11f203dec637ea0f14267a86ec8f6165f63f293f, ROCm 3653e6d, Vulkan 3653e6d6d), ROCm serve OK, guard rc=1, ALL FAIL safe_depth=none. The earlier 20260818T220644Z attempt stopped during preflight before model load and is not benchmark evidence.
- evidence: run-0451
clm-0086
- measured-here high ●●● volatility medium · verified 2026-08-19
- Ling-3.0-flash's shipped NextN head is genuinely activatable through llama.cpp 7077abb's draft-MTP path on ROCm: each bounded n_max=1, 2 and 3 arm created an MTP draft context, emitted nonzero draft accounting, passed all 15 varied output-sanity generations and passed the house guard 4/4. Acceptance was 405/490 (82.65%), 464/693 (66.96%) and 456/899 (50.72%) as depth increased. Activation did not translate into a promotion-worthy speedup: against the same-session 42.2749 tok/s plain median, n1 reached 42.5686 tok/s (1.007x), n2 40.9251 (0.968x), and n3 35.9127 (0.850x), so none met HO-002's predeclared 1.10x threshold. Whole-window wall energy was 1.9689 Wh plain, 2.2954 Wh n1, 2.3259 Wh n2 and 2.6385 Wh n3, but those windows include unequal startup, probe and guard durations and therefore do not establish per-token or per-correct energy. Keep plain decode as the production recommendation; MTP is supported and correctness-preserving within this bounded guard, but provides no measured production benefit here.
- Successful root: aihydra /home/aihydra/bench-results/ling-30-flash-mtp-activation-20260819T041300Z. Runner sha256 cedd3027ac8db3f0ab17ba25b85ebc5d9ac95725570295cf806d23c78be9e200; probe sha256 051c82cf61406078116e6f2ddc867417c8f62e1c1e46401b2af2d0a8033c467e. Foreground rc 0; run-meta records contention=false. The earlier roots at 20260818T230514Z (rc=141 supervisor exit) and 20260819T031245Z (rc=32 gate regex bug) stopped after the plain guard and before any probe or MTP arm; they are operational history, not capability evidence.
- evidence: run-0452 run-0453 run-0454 run-0455 run-0456 run-0457 run-0458 run-0459 eng-0184 eng-0185 eng-0186 eng-0187
clm-0087
- measured-here high ●●● volatility medium · verified 2026-08-19
- NVIDIA Nemotron 3.5 Lightning 30B-A3B's native NextN head is genuinely active through llama.cpp 7077abb's draft-MTP path on ROCm. All n_max=1, 2 and 3 arms created an explicit MTP draft context, produced exact nonzero draft counters, kept all 15 varied sanity prompts clean and passed the house guard 4/4. Against the same-session 65.3310 tok/s plain fixed-workload mean, n1 reached 75.5813 tok/s (1.1569x), n2 69.9942 (1.0714x), and n3 61.6298 (0.9433x), with population CVs of 1.0331%, 0.0732%, and 1.0110%. Fixed-run draft acceptance declined from 118/164 (71.95%) to 147/274 (53.65%) and 155/381 (40.68%) as depth increased. N1 is the best bounded configuration and earns only the card's conditional five-task smoke; n3 is rejected. Do not promote the model or run full tau2 from this performance evidence alone. Gross whole-window energy was 1.2437 Wh plain, 1.3166 Wh n1, 1.2939 Wh n2, and 1.5277 Wh n3; these startup/probe/guard windows establish neither Wh/token nor Wh/correct.
- Successful scientific root: aihydra /home/aihydra/bench-results/nemotron35-lightning-mtp-ab-20260819T112819Z. Exact commands are preserved there as command-plain.sh and command-mtp-n1..3.sh; runner sha256 004b50885c6fdc258130a444a087253bb5959b6f89230c7b273988a5cd54e51b; probe sha256 27751e8c83918417d331221bd5b30222ea550e681598eab77e3da64a6ad8f284; model sha256 6110e2e2e6cd324e6ee69ddced5a6b34fad6c94ca9827222a1e420fb92e3c90b. The earlier HG-002 attempts failed on guard integration, exact-marker matching, and supervisor source-grep logic. They are infrastructure incidents, not scientific evidence, and are not included here.
- evidence: run-0460 run-0461 run-0462 run-0463 run-0464 run-0465 run-0466 run-0467 eng-0189 eng-0190 eng-0191 eng-0192
clm-0088
- measured-here high ●●● volatility medium · verified 2026-08-19
- HG-002 Phase C's corrected matched plain control completed the five-task tau2 airline smoke at 4/5, mean 0.80, with six valid nonempty tool calls, zero empty turns, zero infrastructure errors and zero max-step cuts. MTP n1 scored 1.0 on each of tasks 0-3 with 15 valid calls, but task 4 returned an empty AssistantMessage and tau2 closed it as an infrastructure error; the arm is invalid and has no aggregate mean. Therefore the Phase A 1.1569x bounded throughput lift does not earn full tau2 promotion. Plain remains the recommendation and HG-002 closes without promoting n1.
- The first plain root independently produced the same 0.80 scientific result, but its parser false-negative closed rc 1 and is not represented as a formal pass. Exact artifact roots and hashes are preserved on the run records.
- evidence: run-0468 run-0469 run-0470 eng-0193 eng-0194 eng-0195 inc-0006
clm-0089
- measured-here high ●●● volatility medium · verified 2026-08-19
- Under HG-003's corrected tokenizer-measured protocol, LFM2-24B-A2B Q4_K_M failed retrieval at the minimum tested depth of 2048 tokens with the needle planted early at token position 512. The model returned plain placeholder content with non-empty token and position fields, but neither made a native tool call nor reproduced the planted token. The pre-registered bounded stop therefore fired before late placement and all Stage B guard or performance work. Safe depth is below the practical 4096-token promotion threshold, so this candidate earns no promotion and no performance claim. This is narrow retrieval-capability evidence, not a broad model-quality judgment or an infrastructure fault.
- Evidence root: /home/aihydra/bench-results/lfm2-24b-retrieval-promotion-scientific-20260819T141200Z-supervisor-2edef59c. Exact model sha256 eb4d2d4d4e61b795726c2f526c4434ca6bc725ad7a783691b58681f025cf58f2; ROCm binary build commit 3653e6d (binary sha256 17f55ef427a56ebd98aabf0d3c48e43a93e384e1db5075c5c7bed5ed6a0aad01), with clean source worktree at ce7689f; prompt sha256 d68e71389f75666f2bcc63f55224e4de85a4dafe5a946d99e0a57f81e790ac8c.
- evidence: run-0473 eng-0197
clm-0090
- community low ●○○ volatility high · verified 2026-08-19
- llama.cpp PR #27342's DFlash2 community evidence does not yet justify replacing Qwen3.8-27B's measured serving recommendation, but it does justify a future matched lab check once the implementation is stable: on one Strix Halo Vulkan report DFlash2 Q8_0 slightly beat MTP on prose within noise (18.96 vs 18.36 tok/s) and more clearly on code (22.68 vs 20.15 tok/s), while adjacent reports show hardware- and concurrency-specific hazards including Intel B70 multi-agent collapse, a V100 vision/M-RoPE failure in the draft context, V100 n_max=7 regression, RTX 3090 near-parity/modest gain, and Blackwell better scaling at higher parallelism.
- Source checked live via `gh pr view 27342 --repo ggml-org/llama.cpp --comments` on 2026-08-19. The PR is OPEN and the numbers are self-reported in GitHub comments, not measured here. Treat this as triage context only. If run on aihydra later, the comparison must be matched in one card against this lab's plain floor and draft-MTP control on the same model, quant, backend, build, prompt classes and sampling; use varied prompts plus the house guard, keep `--parallel 1` unless a separate isolation study is explicitly the experiment, and record DFlash2 as a distinct speculation path rather than inheriting MTP's correctness or energy evidence. The Strix Halo block_size observation is also only a future diagnostic: block_size above the drafter's n_extract appeared to be a no-op, but that came from patching GGUF metadata by hand and is not a publishable setting without a controlled run.
- evidence: con-0003
clm-0091
- measured-here med ●●○ volatility medium · verified 2026-08-20
- Nemotron-3-Super-120B-A12B now has a single record-backed Vulkan d204800 throughput cell under the existing cfg-0108 fingerprint: 130.8127 t/s prefill (run-0492) and 15.8832 t/s decode (run-0493), rc=0, contention=false, with the same stock llama.cpp 3653e6d6d Vulkan / UD-Q4_K_M / f16 KV / --load-mode none boundary as the prior d131072 Vulkan rows. This is a narrow backend-depth extension, not a production/capability verdict: compared only to the existing matched records, Vulkan d204800 is 20.05% slower than ROCm d204800 on prefill and 5.81% faster on decode, while versus Vulkan d131072 it is 11.98% slower on prefill and 4.45% slower on decode.
- METHOD — HO-006 admitted evidence from /home/aihydra/bench-results/ho006-nemotron3-super-vulkan-d204800-crossover-r1, variant vulkan-f16-d204800-rep1. Command: llama-bench on /home/aihydra/src/llama.cpp-vk3653e6d/build/bin/llama-bench, model NVIDIA-Nemotron-3-Super-120B-A12B-UD-Q4_K_M-00001-of-00003.gguf, -ngl 999 -fa 1 -b 2048 -ub 512 --load-mode none -dev Vulkan0 -ctk f16 -ctv f16 -t 16 -p 1024 -n 256 -d 204800 -o json. Run-meta anchor: 2026-08-19T23:34:06Z to 2026-08-19T23:58:15Z, 1449 s, rc=0, contention=false, host aihydra. Postrun validation reported no fault markers, no residual llama processes, and released box claim. ENERGY STATUS — wall-meter energy is now joined as eng-0198 from the Home Assistant counter-difference over the preserved run-meta window: 69.1300 Wh gross, 171.75 W mean. This is throughput-window energy only, not a correct-answer denominator.
- evidence: run-0387 run-0388 run-0389 run-0390 run-0492 run-0493
clm-0092
- measured-here med ●●○ volatility medium · verified 2026-08-20
- Ornith-1.0-35B UD-Q4_K_XL has a bounded HO-005 r3 preflight record on the stock 3653e6d6d Vulkan/f16-KV boundary: the Q4 artifact identity was sidecar/local-sha matched, FIT/header/load at c32768 was healthy, the schema-visible guard passed 4/4 (run-0494), and the 5-task tau2 airline smoke passed 5/5 with mean_reward 1.000 (run-0495). The same admitted r3 session also completed one planned fresh-process d262144 throughput cell: 86.1589 t/s prefill (run-0496) and 24.3129 t/s decode (run-0497), with no reviewed DeviceLost/OOM marker. This is only a bounded preflight/smoke plus one model-max throughput cell: it is not a full-26 Q4 capability claim, not a Q4/Q8 equivalence claim, and not an inheritance of Q8_0 capability across a lossy weight-quant change.
- METHOD — HO-005 reviewer-admitted r3 evidence from /home/aihydra/bench-results/ho005-ornith-q4kxl-vulkan-preflight-r3 only. Artifact: /home/aihydra/models/ornith-35b/Ornith-1.0-35B-UD-Q4_K_XL.gguf, 22,324,804,000 bytes, sha256 67081ae4a1a291bd6c72834094ea056332cb3cb5fa15e88536ec7f233a475b71, source unsloth/Ornith-1.0-35B-GGUF per the candidate record. Build/backend: /home/aihydra/src/llama.cpp-vk3653e6d HEAD 3653e6d6d547ec763317d9ecd0ace334a7e21359, build_commit 3653e6d6d, Vulkan0 Radeon 8060S, -ngl 999 -fa 1 -b 2048 -ub 512 --load-mode none, f16 KV, --parallel 1/-np 1; tau2 head 668d3bcd135c02aa3438f987ef45735b7c163ee3. FIT proof: c32768 healthy 2026-08-20T00:17:59Z..00:18:15Z, rc=0, gtt_used_bytes 22444122112. Guard: 2026-08-20T00:18:15Z..00:18:23Z, rc=0, 4/4, native tool call non-empty, isolation skipped only because --parallel 1. Smoke: 2026-08-20T00:18:23Z..00:26:35Z, task ids 0..4, tool_call_messages 23, empty_argument_tool_calls 0, empty_assistant_turns 0, infrastructure_errors 0, rc=0. Non-fatal tau2 log cost-map ERROR lines were present and are not treated as infrastructure errors. Throughput: 2026-08-20T00:26:37Z..00:49:37Z, pp1024/tg256 N=1, parse ok, stddev/sample vectors preserved on run-0496 and run-0497. Failed r1/r2 custody/wrapper attempts are not ingested as metrics. ENERGY STATUS — wall-meter energy is now joined from Home Assistant counter-differences over the preserved run-meta windows: guard eng-0199 (0.1788 Wh), 5-task smoke eng-0200 (16.0646 Wh; 3.2129 Wh/correct for this bounded smoke only), and d262144 throughput eng-0201 (61.9776 Wh gross). These joins do not broaden the claim into full-26 Q4 capability or Q4/Q8 equivalence.
- evidence: run-0494 run-0495 run-0496 run-0497
clm-0093
- measured-here med ●●○ volatility medium · verified 2026-08-20
- Ornith-1.0-35B UD-Q4_K_XL has its own measured HO-005 full-26 tau2 airline capability record on the stock 3653e6d6d Vulkan/f16-KV boundary: after a fresh schema-visible guard passed 4/4 (run-0498), the full task-id 0..25 tau2 airline run completed rc=0 at 22/26 with mean_reward 0.8461538461538461 (run-0499), 185 tool-call messages, 0 empty-argument tool calls, 0 empty assistant turns, 0 infrastructure errors, and 0 max-step cuts. This is measured Q4_K_XL evidence, not inherited Q8_0 capability and not a Q4/Q8 equivalence claim; the earlier bounded preflight/smoke claim clm-0092 remains bounded historical evidence.
- METHOD — HO-005 reviewer-admitted full-26 r1 evidence from /home/aihydra/bench-results/ho005-ornith-q4kxl-vulkan-full26-r1. Artifact: /home/aihydra/models/ornith-35b/Ornith-1.0-35B-UD-Q4_K_XL.gguf, 22,324,804,000 bytes, sha256 67081ae4a1a291bd6c72834094ea056332cb3cb5fa15e88536ec7f233a475b71, sidecar/local sha match, source unsloth/Ornith-1.0-35B-GGUF per the candidate record. Build/backend: /home/aihydra/src/llama.cpp-vk3653e6d HEAD 3653e6d6d547ec763317d9ecd0ace334a7e21359, build_commit 3653e6d6d, Vulkan0 Radeon 8060S, f16/f16 KV, command flags exactly -ngl 999 -fa 1 -b 2048 -ub 512 -c 32768 -dev Vulkan0 -ctk f16 -ctv f16 -t 16 --load-mode none --parallel 1 --jinja --reasoning-format deepseek -n 4096 --slots. Guard: guard-c32768-depth8000 immediately before tau2, 2026-08-20T03:21:15Z..2026-08-20T03:21:24Z, rc=0, 4/4; guard-output.txt shows coherent generation, native get_booking tool call with non-empty arguments, retrieval at depth 8000, and isolation skipped only because --parallel 1/-np 1. Tau2: suite tau2-bench-airline@668d3bc / tau2 head 668d3bcd135c02aa3438f987ef45735b7c163ee3, domain airline, task ids exactly 0..25, seed 42, num_trials 1, max_concurrency 1, max_steps 200, agent temperature 0 and max_tokens 4096, user simulator openrouter/anthropic/claude-haiku-4.5 temperature 0. Deterministic scoring and judge null match the admitted Q8 baseline/run-0345 practice; tau2-judge-note.txt records the CLI/judge rationale. Raw tau2-full26-results.json and tau2-full26-summary.json report failed task ids 7, 21, 23, 24, all termination_reason user_stop. ENERGY STATUS — hborchestrator joined the wall-meter energy while raw Home Assistant history was still available. run-0498 is wired to eng-0202: 0.1888 Wh over the 9 s guard window, with no correct-answer denominator. run-0499 is wired to eng-0203: counter interpolated from 27.780999128 to 27.927536928 kWh across 2026-08-20T03:21:24Z..2026-08-20T04:31:09Z, 146.5378 Wh total, 6.6608 Wh/correct over 22 correct tasks, 2.0182 pence/correct at the standing 30.3 p/kWh tariff. This claim does not make a Q4/Q8 equivalence or energy-ranking claim. CUSTODY/HEALTH — cleanup.log claim string was "claimed by HO-005 Ornith Q4 full-26 tau2 t_210dd8e6, started 2026-08-20T03:19:50Z, no fixed deadline"; cleanup left BOX_AFTER=ABSENT and no residual llama. postrun-validation.json reports final_state success, box_after ABSENT, residual_llama empty, and operational_fault_marker_hits empty. supervise-preclaim-proof showed GPU use 0% and llama-server ABSENT before claim; no reviewed DeviceLost/OOM/SVM/resident- limit/kernel-fault marker appears in operational logs. COMPARISON BOUNDARY — run-0499 may be compared to the Q8_0 full-26 baseline run-0345 only as context: Q8_0 measured 23/26 and mean 0.8846 there, while this lossy UD-Q4_K_XL arm measured 22/26 and mean 0.8461538461538461 here. The two measurements do not establish quant equivalence or permit inheriting Q8_0 capability/energy claims onto Q4_K_XL.
- evidence: run-0498 run-0499
clm-0094
- community med ●●○ volatility high · verified 2026-08-20
- The llama.cpp #25618 F16-V result is a prompt-sensitive exactness result, not a general proof that MTP preserves the target trajectory on Qwen3.8-27B. The independent Strix Halo Vulkan report reproduced the original prose prompt as a byte-exact PASS, but the same target/draft artifacts, commit, f16 K/V cache, server flags, apply-template-to-completion path and greedy reporter sampling produced 0/5 exact-token parity on five other prompts; replaying those five prompts under a separate fixed sampler also produced 0/5 with unchanged first mismatch locations. The bounded conclusion is that one exact MTP trajectory can be real evidence for that trajectory, while varied prompts remain required before claiming baseline/MTP invariance for the runtime.
- Source is a GitHub issue comment, not a local HaloBench run. The artifacts and hashes are recorded in con-0004's upstream report rather than as content/runs because no aihydra run record exists for this diagnostic. Do not promote its timing fields as performance metrics: the report itself states the five-prompt timing totals were diagnostic only because token totals differed. Protocol impact: any future MTP-losslessness or exact-invariance dossier needs a varied prompt set and must preserve the failing prompts or raw token arrays; an exact single-prompt control can show a narrow PASS but cannot clear the broader invariance claim.
- evidence: con-0004
clm-0095
- measured-here med ●●○ volatility medium · verified 2026-08-20
- HO-005 measured a reviewer-admitted Ornith-1.0-35B UD-Q4_K_XL production-depth performance frontier for three exact guarded fingerprints: Vulkan f16/f16 KV, ROCm f16/f16 KV, and Vulkan q8_0/q8_0 KV. The admitted llama-bench cells are schema-visible as run-0503 through run-0522, with guards run-0500 through run-0502, all preserving the exact artifact/runtime/KV/batch/load-mode boundaries and the reviewer caveats.
- Scope boundary: this is guarded performance and energy-ingestion bookkeeping for the exact HO-005 r2 matrix at /home/aihydra/bench-results/ho005-ornith-q4q8-tuning-matrix-r2. It does not claim Q4/Q8 capability equivalence, Q5_K_M behavior, q4_0 KV behavior, unmeasured lossy-KV quality, or a broad production recommendation. Q5_K_M and q4_0/q4_0 KV remain explicitly excluded from this matrix. The r1 failed attempt and r2 intermediate rc=1 wrapper/supervision event are preserved as runner-mechanics deviations, not admitted metrics. Energy remains structured unjoined on these records because this hbrunner profile lacked Home Assistant history credentials at ingestion time; every run carries the exact window(s) needed for later wall-meter Wh-total joins.
- evidence: run-0500 run-0501 run-0502 run-0503 run-0504 run-0505 run-0506 run-0507 run-0508 run-0509 run-0510 run-0511 run-0512 run-0513 run-0514 run-0515 run-0516 run-0517 run-0518 run-0519 run-0520 run-0521 run-0522
clm-0096
- measured-here med ●●○ volatility medium · verified 2026-08-20
- HO-001 v0.6.6 admitted fork-specific DeepSeek-V4-Flash sparse-path and performance evidence for Nathanw1014 portable Vulkan payload source 7b6c6133 build 10569, with matching f16/f16, q8_0/q8_0, and q4_0/q4_0 K/V cache guard plus throughput cells through d262144 and sparse-path proof for q8_0 and q4_0.
- Compatibility claim restored for pre-existing DeepSeek references while HO-009 uses clm-0097. Boundary: fork-specific sparse-path/performance evidence only; no full capability result, no q8_0/q4_0 quality equivalence, no production recommendation, no stock llama.cpp attribution, and no wide speculation/verify claim.
clm-0097
- measured-here med ●●○ volatility medium · verified 2026-08-20
- HO-009 measured reviewer-admitted MTP activation/safety/performance evidence for the genuine Qwen3.6-35B-A3B-MTP UD-Q4_K_M artifact on aihydra. Native MTP draft counters activated for n_max 1/2/3 in Stage A; n_max=2 was the best clean admitted arm and improved server /completion decode throughput versus the plain f16/f16 baseline through d204800 under the exact recorded runtime/config boundary.
- Scope boundary: MTP activation/safety/performance only for /home/aihydra/bench-results/ho009-qwen36-35b-a3b-mtp-refinement-r2. This is not a full capability result, production recommendation, quality equivalence claim, KV-quant conclusion, or inheritance from invalid 27B-MTP evidence. No n_max>=4 safety claim is admitted. n2 d262144 server-perf is preserved as a dropped HTTP 400 context-fit/refusal cell with no substitution. Energy is joined to Home Assistant wall-meter records as Wh total only for every reviewed window.
- evidence: run-0523 run-0524 run-0525 run-0526 run-0527 run-0528 run-0529 run-0530 run-0531 run-0532 run-0533 run-0534 run-0535 run-0536 run-0537 run-0538 run-0539
clm-0098
- measured-here med ●●○ volatility medium · verified 2026-08-20
- HO-004 measured a negative/speculation-path diagnostic for Qwen3.8-27B Q8_0 on the reviewed v0.6.5 portable Vulkan runtime: stock MTP n_max=3 and DFlash2 Q8_0 width 4 both showed log-backed draft activation, but plain, stock-MTP, and DFlash2 arms all failed Stage A on code/toolish sentinel-empty-content checks, so no guard, Stage B performance, throughput, or energy-efficiency cell is admitted.
- Scope boundary: negative HO-004 Stage A evidence from /home/aihydra/bench-results/ho004-qwen38-dflash2-matched-r1 only. This does not change the production recommendation, does not establish full capability, does not claim DFlash2 quality equivalence or throughput improvement, does not inherit community DFlash2 numbers, and does not create any stock MTP n_max>=4 safety claim. Math/creative/reason responses existed in Stage A but are preserved only as non-promoted evidence because code/toolish sentinel failures tripped the stop gate. Home Assistant wall-meter energy was joined for the three Stage A windows as Wh total only, not Wh/correct and not throughput-energy efficiency.
- evidence: run-0541 run-0542 run-0543 eng-0222 eng-0223 eng-0224
clm-0099
- measured-here med ●●○ volatility medium · verified 2026-08-21
- HG-004 Phase-1 captured one real GLM-4.7-Flash Q4_K_M raw llama-server HTTP 200 response through the pinned installed Tau2/native-template task-1 boundary on aihydra; the normalized response was reasoning-only with empty assistant content, no tool_calls, non-empty reasoning_content, and finish_reason length, and Tau2 Phase-2 was not launched.
- Narrow diagnostic boundary only, admitted by hbreviewer card t_f026a804 from /home/aihydra/bench-results/hg004-glm47-phase1-boundary-r1/hg004-glm47-phase1-boundary-r1-20260820T231235Z and /Users/wardenmac/.openclaw/workspace/halobench/reconcile/HG-004-phase1-boundary-review.md. This resolves the immediate boundary toward a served-path response capture rather than pre-simulation/server-start failure. It does not admit a tau2 score, tasks-passed/tasks-total metric, throughput or tok/s metric, Wh/correct or energy ranking, production recommendation, GLM capability acceptance/rejection, or a retroactive full explanation of the historical 2026-08-10 simulations: [] artifact.
- evidence: run-0544
clm-0100
- measured-here med ●●○ volatility medium · verified 2026-08-21
- HO-011 measured Ornith-1.5-35B Q4_K_M screening and full-26 tau2 capability on the stock ROCm0 3653e6d runtime at c32768 f16 KV. Screening passed guard 4/4, smoke 4/5 (mean 0.800, 74 tool-call messages, 0 empty turns, 0 infra errors), and throughput at d262144 of 142.38 t/s prefill and 20.99 t/s decode (N=5 fresh-process reps). The full-26 tau2 airline run scored 8/25 evaluated with mean_reward 0.320, 1 infra error (task 20 JSON parse, 4 retries), 5 TOO_MANY_ERRORS terminations, 2 MAX_STEPS terminations, and 8 context-overflow errors from conversation buildup exceeding c=32768. This is a sharp regression from Ornith-1.0-35B UD-Q4_K_XL (22/26, 0.846) — a version-replacement step-back on the standard suite.
- Scope boundary: HO-011 reviewer-admitted from /home/aihydra/bench-results/ho011-ornith15-q4km-preflight-r6 (screening) and /home/aihydra/bench-results/ho011-ornith15-q4km-full26-r1 (full-26 tau2). This is Q4_K_M evidence on its own boundary; it does not inherit Ornith-1.0 Q8_0 or UD-Q4_K_XL capability, does not establish 1.0/1.5 equivalence, and does not create an MTP/speculation, energy-ranking, or production-recommendation claim. The screening result is a sharp regression and does not justify promotion to full bench. Energy windows are recorded but not joined — this hbreviewer profile lacks Home Assistant credentials; run-meta timestamps are preserved for later batch join. Failed preflight attempts r1 through r5 and their bak001-bak005 backups are operational history only, not admitted metrics. Compare: Ornith-1.0 UD-Q4_K_XL (run-0499, clm-0093) measured 22/26 mean_reward 0.846 on the same tau2 airline suite.
- evidence: run-0545 run-0546 run-0547 run-0548 run-0549 run-0550
clm-0101
- measured-here med ●●○ volatility medium · verified 2026-08-21
- HU-004 reviewed the raw llama-server long-prompt guard for the genuine Qwen3.6-35B-A3B-MTP UD-Q4_K_M artifact on aihydra (runtime commit 7077abb, ROCm0/gfx1151, f16/f16 KV, context 204800). The guard passed ALL 20 cells across the plain (cfg-0141) and native-MTP n2 (cfg-0142) arms, at the sg (target 8K: 7881/7879 prompt tokens) and lp16k (15869/15867), lp175k (17353/17351), lp20k (19849/19847) and lp40k (39786/39784) brackets, each with a natural-language and a native tools fixture. No silent-empty-success, no empty-tool-arguments, no explicit context refusal, no HTTP error, no timeout, and no server crash was observed. ggml-org/llama.cpp#27442 was NOT reproducible on this ROCm0/gfx1151 f16-KV fingerprint.
- Scope boundary: guard evidence only, from /home/aihydra/bench-results/hu004-qwen36-mtp-n2-longprompt-guard-r1 (kanban task t_6b1eba47). Passing this guard clears the silent-empty- success and empty-tool-argument terminal classes on THIS exact ROCm0/gfx1151 f16-KV n2 fingerprint up to the 40K prompt bracket. It does not admit a capability, tau2, throughput, or production result; it does not qualify any other backend, build, KV policy, or context; it does not establish n2 d262144; and it claims nothing about n_max>=4. One behavior recorded honestly in run-0552: the n2 lp20k-tool cell returned 79 repeated identical get_booking tool calls (finish_reason length, 57.1 wall s vs ~8 s plain) carrying the correct non-empty required booking_id, one call with a truncated argument string. No required tool argument was empty and no empty assistant turn occurred, so no guard terminal class fired and the cell remains guard-pass. Energy windows are preserved in run-meta timestamps but not joined — this hbreviewer profile lacks Home Assistant credentials.
- evidence: run-0551 run-0552 cfg-0141 cfg-0142
clm-0102
- measured-here med ●●○ volatility medium · verified 2026-08-22
- LLaDA2.2-flash (100B-A13B block-diffusion MoE, Q4_K_S) solved 5/5 tau2-bench airline screening tasks (task_ids 0-4), mean reward 1.000, with real multi-turn tool use, running GPU-accelerated on aihydra (Strix Halo, gfx1151) via the headbouyJB/diffuse-cpp fork with the GPU-resident KV cache and full Levenshtein editing enabled. This is a 5-task SCREEN, not the full 26-task bench. Energy-joined (eng-0226): 62.10 Wh total, 12.42 Wh per correct answer.
- evidence: run-0553 eng-0226
clm-0103
- measured-here high ●●● volatility low · verified 2026-08-22
- The headbouyJB/diffuse-cpp fork makes a scaled block-diffusion LLM benchmarkable on portable AMD hardware: GPU offload of the MoE forward (gfx1151), OpenAI tool-calling, and a GPU-resident inter-step KV cache with cross-turn prompt reuse together take decode-at-long-context from ~1 tok/s (host-array cache) to ~5-12 tok/s, turning a multi-turn diffusion agent from impractical (~tens of seconds to minutes per turn) into a runnable tau2 screen (~1-3 min/task).
- evidence: run-0553 con-0005
clm-0104
- measured-here med ●●○ volatility medium · verified 2026-08-22
- ggml_flash_attn_ext on this HIP/gfx1151 build does not fully honour an additive F16 block-causal mask: a committed prefix's per-layer K/V changed by several logits when a later masked block entered the attention window (compounding through layers), while the manual ggml_soft_max_ext path with the identical mask did not. The boundary tracks 64 keys (= 2x the Wave Size 32), pointing at a tiling/reduction issue in the flash kernel's handling of masked lanes. Inferred so far via a downstream KV-cache oracle; an independent generation-free minimal repro is drafted but NOT yet run, so this is not yet an upstream-filed bug.
- evidence: run-0553
clm-0105 superseded
- community med ●●○ volatility medium · verified 2026-08-22
- SUPERSEDED (operative claim) as of strix-halo-llamacpp v0.6.10. As written against v0.6.9 (2026-08-22): on hybrid GDN (qwen35moe) targets the then-stable Strix Halo Vulkan fork did NOT guarantee token-exact MTP rollback after a rejected draft. v0.6.8 had introduced a one-line change forcing MTP rollback through full sequence-state checkpoints; v0.6.9 (2026-08-22T04:09:43Z) reverted that line because the full-checkpoint save/restore path deadlocked deterministically on Vulkan with a hybrid target (Qwen3.6-35B-A3B stalled a few hundred tokens into a long response, all threads parked in futex_do_wait, GPU idle, GTT flat). With the revert, rollback returned to the v0.6.4-v0.6.7 fast snapshot-plane restore, which the vendor stated "can diverge slightly from a no-draft run after a rejected draft" - a subtle distribution drift after rejection, not garbled output. That was a deliberate availability / exactness tradeoff by the fork, pending a state-save-path fix. v0.6.10 (clm-0107) re-lands token-exact full-checkpoint rollback; the operative tradeoff claim above no longer holds on the current stable fork.
- Historical fourth arm of the MTP-correctness dossier (with con-0009/clm-0107). Fork-vendor release-note context only: single-payload verification on gfx1151 Vulkan, speed figures explicit single runs NOT the BENCHMARKS.md protocol, so nothing here inherits into a HaloBench number. The v0.6.9-era exactness tradeoff is superseded by v0.6.10 (token-exact rollback restored via 9c5d899 + f25eefe). Reassessment outcome (see clm-0107): the house varied-prompt/raw-token MTP-invariance requirement does NOT relax - the v0.6.9 tradeoff was a fork-rollback-path issue, whereas the #25618 quantized-target divergence (clm-0094/clm-0104/clm-0106) and #26750 spec-path collapse (clm-0102) are separate upstream verify/target-side issues that still stand. And HO-009 plain-vs-n2 results (clm-0097, clm-0101) were measured on UPSTREAM ggml-org/llama.cpp 7077abb on ROCm0/gfx1151 - which never used the fork's snapshot-plane rollback path - so this tradeoff never directly applied to those numbers and v0.6.10 changes no comparison boundary there.
- evidence: con-0007 con-0009
clm-0106
- community med ●●○ volatility high · verified 2026-08-22
- The quantized-target divergence in the llama.cpp MTP/DSpark speculative-verify path corroborates across a NEW model, OS and GPU family: on NVIDIA Vulkan/Windows 11 (llama.cpp b10566 / bb4caa754, GTX 1660 SUPER) the official LiquidAI LFM2.5-2.6B pair diverges from vanilla on a Q8_0 target, while the SAME Q8_0 draft against an F16 target produces byte-identical greedy output to the F16 vanilla target (40/64 draft tokens accepted). The divergence is a property of the quantized-target x speculative-verify interaction, not the drafter: with the Q8_0 target the DSpark output consistently and deterministically differs from vanilla, but loading the draft model with --spec-draft-p-min 1 (zero draft tokens) restores byte-exact match to vanilla Q8_0, and the Q8_0 mismatch persists with --spec-draft-n-max 1 and target KV held in F16 (so it is neither a larger speculative block nor quantized target KV). It also diverges at n_max=1, unlike the Qwen3 boundary earlier in #25618 where n_max=1 was reported lossless.
- Upstream report only, the fifth arm of the #25618 MTP-correctness dossier (with con-0004/clm-0094, con-0005/clm-0102, con-0006/clm-0104, con-0007/clm-0105). This is INDEPENDENT corroboration of the quantized-target-divergence axis, deliberately on a non-Qwen3.8, non-ROCm/AMD family: LiquidAI LFM2.5-2.6B, NVIDIA GTX 1660 SUPER, Windows 11, Vulkan, Q8_0 target (the original report was Qwen3.8 Q6_K/mixed Q4 and ROCm/RDNA4/Vulkan/Metal). The strong differential control (same Q8_0 drafter, only the target quant differs, F16 target byte-identical to vanilla) isolates target quantization as the causal variable and broadens the house warning that the drift is on the target/verify side, not the drafter. Reported surface is NVIDIA Vulkan/Windows, NOT our ROCm/Vulkan Strix Halo gfx1151 aihydra box, so none of these figures are a HaloBench benchmark gain. It does NOT change the house varied-prompt/raw-token MTP-invariance requirement -- it STRENGTHENS it: quantized-target drift is now evidenced on multiple unrelated families, so a local varied-prompt / raw-token control remains required before any baseline/MTP invariance claim on our surface, and the target quantization axis must be treated as a live loss source regardless of drafter or model family. The n_max=1 divergence on LFM2 (vs lossless on Qwen3.8 at n_max=1) also warns that the narrow n_max=1 boundary is not a general safety property across targets.
- evidence: con-0008
clm-0107
- inferred med ●●○ volatility medium · verified 2026-08-22
- The strix-halo-llamacpp fork v0.6.9 MTP rollback-exactness tradeoff is REVERSED in v0.6.10: with the root-cause fix (server no longer re-verifies replayed draft tokens after a checkpoint restore, 9c5d899) and full-checkpoint MTP rollback re-applied (f25eefe), MTP rollback on hybrid GDN (qwen35moe) targets is token-exact again on the current stable fork, and the long-run stall is gone. The prior operative claim in clm-0105 - that the current stable fork does NOT guarantee token-exact MTP rollback after a rejected draft - is superseded as of v0.6.10 (2026-08-22T12:03:16Z). This does NOT relax the house varied-prompt/raw-token MTP-invariance requirement: the v0.6.9 tradeoff was a fork-rollback-path issue, whereas the #25618 quantized-target divergence (clm-0094/clm-0104/clm-0106) and #26750 spec-path acceptance collapse (clm-0102) are separate upstream issues on the verify/target side that still stand.
- Sixth arm of the MTP-correctness dossier, superseding the operative claim of clm-0105 (fork rollback exactness). Fork-vendor release-note context only (single-payload verification on gfx1151 Vulkan; speed figures are single runs, NOT the BENCHMARKS.md protocol), so nothing inherits into a HaloBench number. This changes NO comparison boundary: HO-009 plain-vs-n2 results (clm-0097, clm-0101, cfg-0141/cfg-0142) were measured on UPSTREAM ggml-org/llama.cpp commit 7077abb on ROCm0/gfx1151, which never ran the fork's snapshot-plane rollback path, so v0.6.10 does not alter their interpretation; the active HO-009 n2 tau2 lane (t_ca9ef5d9) is untouched. Also flagged as LINK-UP for the ling-30-flash candidate screen: v0.6.10 adds bailingmoe3/Ling 3.0 DSpark spec support (upstream PR #27508, merged 2026-08-22T09:19:49Z). Runtime complete; NOT exercisable end-to-end until Ling-3.0-flash draft GGUFs are published (inclusionAI/Ling-3.0-flash-dspark has model.safetensors up, GGUFs not out as of 2026-08-22T16:19Z). When the DSpark GGUFs land, Ling-3.0-flash (already benched, clm-0086: MTP n_max 1/2/3 engagement real but median decode ratios 1.007x/0.968x/0.850x below the 1.10x promotion threshold, so plain decode stayed production default) can be re-examined against the DSpark draft path once a registered fork build ships it. Keep clm-0105 for the v0.6.9 historical fork record.
- evidence: con-0009
clm-0108
- measured-here high ●●● volatility medium · verified 2026-08-22
- HO-009-AB measured plain-vs-n2 capability A/B on Qwen3.6-35B-A3B-MTP at production agentic depth (c32768), reasoning OFF on both arms (build 7077abb, ROCm0/gfx1151, UD-Q4_K_M, f16/f16 KV). Plain scored 6/26 (mean 0.2308) vs native-MTP n2 (--spec-type draft-mtp --spec-draft-n-max 2) scoring 0/26 (mean 0.0). Fisher exact one-sided n2-worse p=0.0113, two-sided p=0.0226: with the tau2 reasoning_content-drop confound removed, MTP at n_max=2 is NOT capability- safe on this build — n2 is a real, statistically significant capability loss vs plain on the same task set, not noise-around-plain. Both arms clean: 0 infra / 0 empty-assistant / 0 empty-arg tool calls, no single-token cliff (min asst completion_tokens=14, n2). Wilson 95% CIs: plain [0.11,0.42], n2 [0.00,0.13]. n2 terminates via 21 too_many_errors / 2 max_steps / 3 user_stop with 415 tool call messages -> agent-side validation-error spam, a capability loss in practice.
- Plain-vs-n2 capability comparison only (assessed as-run on build 7077abb). This does NOT change the n2 decode-throughput context of clm-0097: n2 still improves server /completion decode throughput through d204800, but that throughput benefit carries a capability cost, so plain remains the capability-safe option. No n_max>=4, KV-quant, alt-backend, or model-change claim is made or inherited. No n2 quality-equivalence, production recommendation, or throughput claim beyond this plain-vs-n2 capability comparison on this card. The reasoning-off requirement itself costs capability on this model (plain 6/26 vs the 122B f16 10/26 / q8_0 15/26 range), consistent with the DIAG confound fix; both arms ran identical reasoning-off protocol. The v0.6.10 MTP rollback re-land is a separate future-facing question tracked in the MTP-correctness dossier (clm-0105/clm-0107); it does not change this as-run verdict on upstream 7077abb.
- evidence: run-0568 run-0569 run-0570 run-0571
clm-0109
- measured-here high ●●● volatility medium · verified 2026-08-23
- Null-update diagnostic: the v0.5.x base-platform Stage-A sentinel empty-content defect is RESOLVED on the reviewed v0.6.10 portable Vulkan/RADV runtime (2586f6edd). The plain/control arm that failed 6/15 Stage-A sentinel on v0.5.x (run-0541/0542/0543) PASSES on v0.6.10 (0 empty-assistant, 0 empty-tool-call, 0 single-token-EOS, code-valid 3/3, tool-correct 3/3) plus the 4/4 house guard, so the plain-vs-DFlash2-vs-stock-MTP comparison becomes legal. Both spec arms (stock-MTP n_max=3, DFlash2 Q8_0 n_max=3) also pass Stage-A sentinel and the 4/4 house guard with log-backed draft activation, but the G2 varied-prompt invariance check is NEGATIVE for both — neither is byte-invariant to the plain control (2/5 prompt classes match; DFlash2 first-diff on code, stock-MTP first-diff on toolish) — so both spec arms are recorded as NOT PROMOTED.
- HO-004-DIAG / clm-0098 successor on v0.6.10. Scope boundary: Stage-A sentinel and house-guard pass only. NOT "DFlash2 works", NOT "DFlash2 beats MTP", NOT a quality-equivalence or throughput-improvement claim, NOT a production recommendation — the exactness gate (G2) is negative for both spec arms, so nothing beyond the empty-content defect resolution is admitted. Deviations: ran Vulkan/RADV per hborchestrator (card ROCm line was boilerplate); DFlash2 n_max=3 because the card forbids n_max>=4; G2 first attempt crashed on a runner slicing bug and re-ran G2-only. Energy: HA wall-meter join is deferred to the energy-join owner (runner lane had no HA MCP) with per-arm windows recorded; join before the ~10-day HA retention closes. No Wh/correct anywhere. This is a sentinel/gate diagnostic, not a capability or throughput benchmark.
- evidence: cfg-0146 cfg-0147 cfg-0148 run-0572 run-0573 run-0574
clm-0110
- measured-here high ●●● volatility medium · verified 2026-08-23
- On the v0.6.10 strix-halo fork under Vulkan/RADV (build 2586f6edd), Qwen3.6-35B-A3B-MTP UD-Q4_K_M with reasoning OFF recovers MTP n2 capability that build 7077abb-upstream-on-ROCm0 destroyed. This is a STACK effect, not a build-only change: AB ran ggml-org 7077abb on ROCm0/gfx1151 (cfg-0144/0145, run-0568/0569/0570/0571), while this retest ran the Nathanw1014 fork on Vulkan/RADV (cfg-0149/0150, run-0575..0578) — repo, backend and build all changed together, so the comparison is strictly "v0.6.10-fork-on-Vulkan vs 7077abb-upstream-on-ROCm0", never a build-only claim. Within v0610, native-MTP n2 (--spec-type draft-mtp --spec-draft-n-max 2) scored 17/26 (mean 0.65385) vs plain 22/26 (mean 0.84615); Fisher exact two-sided n2-worse p=0.1994 — n2 is NOT significantly worse than plain on this stack. Against the ORIGINAL AB n2 failure, n2 17/26 vs AB n2 0/26 (run-0571) p≈2.8e-07 (Fisher two-sided) — MTP n2 capability destroyed on 7077abb is recovered on v0.6.10. The empty-turn/agent-spam mechanism is resolved at the mechanism level: all 26 v0610 n2 terminations were user_stop with 0 too_many_errors, 0 max_steps, 0 empty-assistant and 1 empty-argument tool call (tool_call_messages 192), versus AB n2's 415 tool-call + 21 too_many_errors + 2 max_steps spam on run-0571. Scope is strictly qwen35moe target, UD-Q4_K_M, reasoning-off, n_max=2 only. NO DFlash2, stock-MTP, n_max>=4, reasoning-on, alt-model, alt-backend, or production-throughput claim is made or inherited.
- Capability-recovery claim resolving the HO-009-AB negative verdict (clm-0108) for the v0.6.10 fork/Vulkan stack. Values are RECORD values from the raw tau2 summaries (plain tool_call 215 / empty-arg 1; n2 tool_call 192 / empty-arg 1), NOT the over-reported parent metadata (315/2). MTP draft acceptance range 0.728-1.00 per the n2 guard (run-0577). Fisher p-values: n2-vs-plain within v0610 = 0.1994 (two-sided, not significantly worse), n2 17/26 vs AB n2 0/26 ~2.8e-07, plain v0610 22/26 vs AB plain 6/26 = 1.7e-05. Energy join: eng-0235 (plain) / eng-0236 (n2). The mechanism-level empty-turn recovery supersedes the AB reasoning-off capability loss ONLY on this exact stack+config scope; it does NOT relax the house varied-prompt/raw-token MTP-invariance requirement (clm-0102/clm-0104/ clm-0106 still stand) and does not alter the plain-remains-capability-safe upstream-7077abb recommendation of clm-0108.
- evidence: cfg-0149 cfg-0150 run-0575 run-0576 run-0577 run-0578 run-0569 run-0571
clm-0111
- measured-here high ●●● volatility medium · verified 2026-08-23
- On Qwen3.5-122B-A10B UD-Q4_K_M (build ce7689f, llama.cpp-kvfix tree, ROCm0 Radeon 8060S gfx1151 122880 MiB, q8_0/q8_0 KV baseline) the production-optimisation matrix interior conclusions hold (internal comparison, same model/quant/build): (1) interactive first-token latency (ttft) is flat at ~1420 ms across batch 512-2048 at ub=512 (all N=3, CV<1.6%), and degrades when ubatch drops below 512 — ub256 ttft ~1683-1691 ms, ub128 ttft ~2364 ms — so ub=512 is the interactive optimum and tpot stays flat ~45.8 ms/tok throughout. (2) Depth d262144 LOADS with q8_0 KV, no OOM; pp16384 = 62.25 tok/s, tg256 = 8.448 t/s — the >=200k production-depth probe is satisfied on this config. (3) KV q4_0 is NOT statistically equivalent to f16: ttft +0.14% (t=-0.39, within noise) but decode (tpot) +0.407% is resolvable (t=-21.45, N=3 CV<0.04%) — q4_0 hides a small but measurable decode regression and is UNPROFITABLE at this depth; it is near-free as the protocol expects on a hybrid-attention minority but not free. Boundary: internal matrix only — same model, same quant UD-Q4_K_M, same build ce7689f, same backend ROCm0; only batch/ubatch, KV quant, and context depth vary. NO cross-model, cross-build, or cross-backend-equivalence claim is made or inherited.
- Ingested from HO-012-REVIEW (t_7a165bde) ADMIT-with-corrections verdict. Key corrections applied vs the run card's initial wording: Cell-3 is NOT called "statistically equivalent" — the +0.407% tpot regression is resolvable (t=-21.45) and called out; energy window count is exactly 20, not 21. Per-rep values were independently re-derived by the reviewer from the per-arm JSON. This is a production-optimization-matrix claim for the local-agent stack: it does not establish a new 122B capability result, does not change the model page verdict, and is bounded to the internal matrix boundary above. Energy windows joined via HA counter-diff (eng-0237..0245).
- evidence: cfg-0151 cfg-0152 cfg-0153 cfg-0154 cfg-0155 cfg-0156 cfg-0157 cfg-0158 run-0579 run-0580 run-0581 run-0582 run-0583 run-0584 run-0585 run-0586 run-0587 run-0588 run-0589 run-0590 eng-0237 eng-0238 eng-0239 eng-0240 eng-0241 eng-0242 eng-0243 eng-0244 eng-0245
clm-0112
- measured-here high ●●● volatility medium · verified 2026-08-23
- On Ornith-1.0-35B-UD-Q4_K_XL (v0.6.10 fork build 2586f6ed, Vulkan/RADV, f16/f16 KV, -ngl 999 -fa 1 --parallel 1 --load-mode none -t 16, plain decode) the interactive-decode batch/ubatch tuning result is effectively FLAT at the d131072 production-depth probe: decode runs 34.342-34.430 tps across b512/ub512 (34.357, CV 0.073%), b1024/ub512 (34.430, CV 0.321%), b2048/ub1024 (34.363, CV 0.055%) and b4096/ub2048 (34.342, CV 0.089%) - a 0.26% spread with max CV 0.32%. NO configuration is statistically resolvable as faster than any other; the numeric maximum (b1024-ub512) is explicitly NOT asserted as fastest. Separate depth anchor at 262144 (b2048/ub512, pp1024/tg256, N=1): decode 24.425 t/s, with the paired prefill 185.806 tok/s recorded apart, not merged. Boundary: single-model, single-build, interactive- decode tuning only; this is a performance/no-change result, NOT a capability change, cross-model comparison, tau2 result, production-throughput equivalence, or energy ranking.
- Ingested from HO-005-REVISED-REVIEW (t_b3f466d1) ADMIT-with-corrections verdict. Anchor decode = 24.425 t/s (tg256 row); the 185.806 prefill is a separate figure, not merged. Wording faithfully matches the review: "effectively flat decode across batch/ubatch at d131072 (34.34-34.43 tps, spread 0.26%, max CV 0.32%); no configuration is statistically resolvable as faster." Energy windows (eng-0246..0250) joined this run via HA counter-diff, so energy_unjoined_reason was cleared on the runs.
- evidence: cfg-0159 run-0591 run-0592 run-0593 run-0594 run-0595 run-0596 eng-0246 eng-0247 eng-0248 eng-0249 eng-0250
clm-0113
- measured-here high ●●● volatility medium · verified 2026-08-23
- On the v0.6.10 strix-halo fork under Vulkan/RADV (build 2586f6edd), Qwen3.6-35B-A3B-MTP UD-Q4_K_M: raising native-MTP draft n_max from 2 to 4 (reasoning OFF) is FLAT — n4-off full-26 scored 17/26 (mean 0.653846), identical to n2-off (17/26, run-0578); Fisher exact two-sided n4-vs-n2 p=1.000000. Neither n2 nor n4 is significantly worse than plain (plain 22/26, mean 0.846154, run-0576): n4-vs-plain p=0.199350, n2-vs-plain p=0.199350 — matching the documented HO-009-v0610 baseline. So n_max=4 does NOT narrow the 5-correct deficit to plain, does NOT degrade, and does NOT collapse (EOS-cliff sentinel PASS 0/8 at n_max=4; the clm-0055 accumulation trigger did not reproduce). The reasoning-ON path is BLOCKED, not just degraded: the plain-rea-ON content sentinel FAILED under G6 (reasoning_content_len=1846 but assistant_content_len=0, empty assistant content, rc=2) — the same empty-content/content-drop defect lineage that forced reasoning OFF in HO-009-AB re-emerges under -rea on on this fork/stack. Because the failure is on the PLAIN reasoning-ON control (no MTP contribution required), the entire reasoning-ON lane is do-not-score: n2-rea-on tau2 skipped, gated n4-rea-on cancelled. Within the admitted reasoning-OFF capability matrix, plain remains the capability winner (22/26). Scope is strictly single-model/single-build MTP n2|n4 × reasoning on|off on this fork stack: cfg-0160..0163, build 2586f6edd, UD-Q4_K_M sha 0b21525e, f16/f16 KV, c32768, airline@668d3bcd seed 42 tasks 0-25 trials 1 agent temp0 max_tokens4096 judge claude-haiku-4.5 temp 0. NO cross-model, KV-quant, DFlash2, n>4, degradation- threshold, or production-throughput claim is made or inherited.
- Capability-matrix claim from HO-013-REVIEW (t_ebb2e334) ADMIT verdict: Cell-0 EOS sentinel PASS, Cell-1 n4-off full-26 ADMITTED (17/26, comparable + complete), reasoning-ON REJECTED under G6 (do-not-score) with n2-ON tau2 skipped and gated n4-ON cancelled. Values are RECORD values from the raw cell1-n4-off/tau2-full26-summary.json (tool_call 211 / empty-arg 1, verified on-disk by review) and sentinel reasoning-content files (reasoning_content_len = 1846, NOT 2182 as parent metadata incorrectly stated — provenance corrected in the review verdict). Fisher p-values recomputed by review (n4-vs-plain 0.199, n4-vs-n2 1.0, n2-vs-plain 0.199). Energy join: eng-0251 (57.17 Wh, wh/task 2.199) added this ingest; compare arms eng-0235/eng-0236. This closes the MTP n_max investigation: n2 and n4 are statistically identical, both trail plain. Reasoning-on remains a production-relevant blocker for the config that must be reasoning-ON; a DIAG card mirroring HO-004-DIAG is recommended for the reasoning-ON empty-assistant content-drop lineage on this fork. The plain- remains-capability-safe recommendation of clm-0108/clm-0110 is unchanged within reasoning-OFF scope; this claim does NOT relax the house varied-prompt/raw-token MTP-invariance requirement (clm-0102/clm-0104/clm-0106 still stand).
- evidence: cfg-0160 cfg-0161 cfg-0162 cfg-0163 run-0597 run-0598 run-0599 run-0600 run-0576 run-0578 eng-0251
clm-0114
- community med ●●○ volatility medium · verified 2026-08-23
- The llama.cpp MTP/spec-path multi-GPU fragility corroborates across a second, platform, with a validated workaround. On llama.cpp #27122 a CUDA/RTX-A4000-4x tensor-split build (Threadripper Pro 3945WX, no NVLink/P2P, closed driver 595.84, build 749f688fc) hard-crashes 6/6 on a Qwen3.8-27B Q6_K + draft-mtp n_max=3 deep 131072-token prefill under --split-mode tensor (symptom is a full machine power-off with BMC "Power off/down", no Xid/AER/panic), while the SAME workload on 2 GPUs (internal 2-device AllReduce path, not the meta-backend) is completely stable with MTP enabled. The defect is specific to the multi-GPU meta-backend MTP path. LLAMA_GRAPH_REUSE_DISABLE=1 lets the exact crash config survive the full 131072 prefill (610.82 t/s prefill, 27.53 t/s decode @131K, 38.43 t/s @4K), consistent with PR #24549's mechanism (graph reuse leaves dangling per-device tensor references when MTP/target contexts share memory under SPLIT_MODE_TENSOR). The corollary for the house: a HIGH --spec-draft- n-max interacts with the multi-GPU MTP path where the crash frequency scales with n-max.
- Corroborating arm of the MTP-correctness dossier, independent platform confirmation of the MTP CUDA multi-GPU lockup axis (originally tripletto, now independently reproduced by mazinist on a different GPU/OS/platform, with the earlier zyxyunxin 2-GPU note). Reported surface is CUDA multi-GPU tensor split; HaloBench is single-GPU (aihydra, Strix Halo APU, ROCm/Vulkan), so this does NOT change any HaloBench boundary, config, run or production claim — it confirms the spec-path fragility and the value of a shallow/adaptive draft depth. No HaloBench benchmark number inherits from these figures.
- evidence: con-0010
clm-0115
- community med ●●○ volatility medium · verified 2026-08-23
- A high fixed MTP draft-depth CEILING carries a measurable fixed cost that is INDEPENDENT of the adaptive logic: on llama.cpp PR #27210 (draft-mtp-adaptive), stew675 held the algorithm at fixed depth 3 but kept max depth 10 and found merely having max depth 10 incurred a fixed ~2.6% performance penalty on its own; the subsequent fix (2026-08-23T01:32:47Z, limit the full MTP buffer scan on truncated/short drafts) recovers ~2-3% for --spec-draft-n-max 10 --spec-draft-p-min >0.5 configs. marcusds independently confirms on a single RTX 5090 that depth CHANGING does not degrade CUDA-graph performance (GGML_CUDA_DISABLE_GRAPHS=1 vs default deltas small, baseline C0 -3.0% / -1.9% / -1.8% / -1.7%), and that adaptive 3..10 + p-min is the only config with a POSITIVE recall delta (+2.6%) — so the ~2.6% is the cost of a static high ceiling, not of adaptive switching. The house reading: a deep static or deep blind-sweep n_max is not a safe/default target; a shallow-plus-adaptive draft depth (low floor, bounded ceiling, adaptive climb/drop) is the target.
- Draft-depth-cost arm of the MTP-correctness dossier. Reinforces the standing shallow/adaptive draft-depth guidance (protocol §1a/§0 and the house n_max<=3 caution on qwen38-27b, clm-0055) and the "no deep blind sweep" position: a high fixed n_max ceiling carries a fixed cost even when the model never uses it. Reported surface is CUDA/gfx1201 upstream PR context, not our single-GPU ROCm/Vulkan Strix Halo aihydra box, so no HaloBench benchmark claim or n_max guidance change is made from these figures; the shallow/adaptive + varied-prompt requirement stands intact.
- evidence: con-0011
clm-0116
- measured-here med ●●○ volatility medium · verified 2026-08-23
- HK-001-FULLBENCH-REVIEW record for the affected v1.1 Kairic Edge TheRock configuration (cfg-0164): Qwen3.8-27B-IU4-Kairic-Edge.gguf on the TheRock 7.14 / AMD clang 23.0.0 ROCm runtime, c32768, --kairic-edge --no-mmap. The record-backed guard exercised three probes and passed them, with cross-request isolation N/A under --parallel 1 (run-0601). Record-backed performance evidence retained here is the c32768 FIT (prefill 279.6 t/s, decode 11.73 t/s; run-0602) and the engine-path depth rungs at d0 (prefill 612.3 t/s, decode 13.14 t/s; run-0603) and d32768 (prefill 272.1 t/s, decode 11.66 t/s; run-0604). These are measurements of the affected v1.1 configuration and are not clean/v1.2 capability validation. Served-path, deeper-depth, smoke, and full-26 attribution is on evidence hold in this correction. No replacement run IDs, metrics, or energy allocations are asserted.
- HK-001-FULLBENCH-REVIEW retained only the locally resolved guard and performance records run-0601..run-0604. Other HK-001-FULLBENCH material is held by the correction packet; this record does not make a clean/v1.2 capability claim. [2026-08-24 COMPARABILITY ANNOTATION — HK-V12-ANNOTATE, per t_b05473c9] The retained records use the affected v1.1 Kairic Edge runtime. The server_flags_degraded marker identifies that configuration state; it does not validate v1.2 capability or clear the evidence hold.
- evidence: cfg-0164 run-0601 run-0602 run-0603 run-0604 con-0012
clm-0117
- measured-here high ●●● volatility medium · verified 2026-08-23
- On Qwen3.6-35B-A3B-MTP UD-Q4_K_M (build 2586f6edd strix-halo v0.6.10 fork/Vulkan RADV, capability c32768, plain mode -rea off, no MTP, f16/f16 KV baseline, temp 0 seed 42 max_tokens 4096) the production-optimisation matrix interior conclusions hold (internal, same model/quant/build): (1) interactive first-token latency (ttft) is minimally sensitive to batch once ubatch=512 — b1024/ub512 1353.1 ms (CV 1.46%), b2048/ub512 1357.4 (CV 2.01%), b512/ub512 1364.5 (CV 1.13%), all N=5 d131072 — and degrades when ubatch drops below 512 (ub256 ttft ~1538-1545 ms, ub128 2320 ms); tpot stays flat ~28.4-28.5 ms/tok across all six arms, so b1024/ub512 is the interactive/ttft production choice (35.1 t/s decode). (2) Depth is viable through d262144 (no OOM) at b1024/ub512: pp_tps 356.8/238.8/167.4 and tg_tps 35.0/28.8/24.9 at d131072/d200000/d262144 (N=3). CAVEAT required (house scatter rule, reviewer-required): the d262144 pp_tps 167.4 figure carries pp_cv 8.22% (>3% scatter rule) so it is a COARSE survival figure reported WITH its CV, NOT a clean mean; tg_tps 24.9 (CV 0.32%) is clean; the production config is d131072 so the recommendation is unaffected. The >=200k production-depth target for the agentic stack is satisfied. (3) KV q4_0 is a CANDIDATE production setting at the production config: decode ~24% faster (tpot 21.670 vs 28.503 ms, ttft within noise 1334.2 vs 1357.2 ms, N=5), mean KLD 0.0109±0.0003, PPL ratio 1.0005, same-top agreement 95.33% (control 100%) — small KLD, no cliff, no resolved regression, not rejected. CAVEAT (hybrid-arch, reviewer-required): this is a hybrid/recurrent model — only 11 of 40 layers expose a quantisable full-attention KV (G3 probe blk 3..40 step 4), so the small delta is partly ARCHITECTURAL; the recommendation is bounded to THIS exact fingerprint and is not a claim that q4_0 KV is lossless on a dense-attention or long-KV model. Boundary: internal production matrix only — same model, same quant, same build, same backend; only batch/ubatch, context depth, and KV quant vary. NO cross-model, cross-build, cross-backend-equivalence, MTP n2/n4, or reasoning-ON claim is made or inherited (HO-013 closed n_max flat + reasoning-ON rejected).
- Ingested from the combination of HO-014A-FILL Cell-1/2 (t_315f0104) and HO-014-RUN Cell-3 (t_677f1b47) via the HO-014-REVIEW ADMIT verdict (t_dc235dcd). All reviewer-required corrections applied verbatim: (1) d262144 pp_tps 167.4 labelled a coarse survival figure with its pp_cv 8.22%, never a clean mean; (2) energy join is an explicit HA counter-diff (eng-0256..0266), never a silent null; (3) production recommendation bounded to this fingerprint WITH the hybrid-arch caveat (11/40 layers quantisable). Values are RECORD values re-verified against summary.tsv / cell2-summary.tsv / cell3-summary.tsv and cell3-verdict.md. The Cell-3A KLD earlyoom at ctx=131072 was an instrument logits-buffer wall (~96 GiB RSS, not model OOM) and was repaired at capability c32768 with both arms identical (recorded in cfg-0171/run-0617). q4_0 decode advantage and safe-KV verdict establish a candidate production KV setting that only a full cross-model capability Δ could overturn — none is made here. This is a production-optimisation claim for the local-agent stack; it does not change the capability verdict (clm-0113) and is interior to this fingerprint.
- evidence: cfg-0165 cfg-0166 cfg-0167 cfg-0168 cfg-0169 cfg-0170 cfg-0171 run-0605 run-0606 run-0607 run-0608 run-0609 run-0610 run-0611 run-0612 run-0613 run-0614 run-0615 run-0616 run-0617 eng-0256 eng-0257 eng-0258 eng-0259 eng-0260 eng-0261 eng-0262 eng-0263 eng-0264 eng-0265 eng-0266
clm-0118
- measured-here high ●●● volatility medium · verified 2026-08-23
- HK-RERUN-REASONOFF verdict on the Kairic Edge TheRock build with reasoning DISABLED (-rea off, cfg-0172) — the fix path clm-0116 prescribed for the reasoning-ON empty-assistant defect. Verdict: ADMIT-with-corrections applied verbatim (HK-RERUN-REASONOFF-REVIEW, t_3b03c623). ADMITTED CLAIM 1 — tau2 full-26 capability, reasoning-OFF: reward 0.423 (11/26 passed, total_reward 11.0), 26/26 simulations COMPLETED with ZERO empty-assistant turns and zero empty-argument tool calls — the prior full-26's 60%-abort / "AssistantMessage must have either content or tool_calls" failure mode is eliminated by -rea off. The defect was a server-config interaction, not model capability. Wall 3181 s at ~13 t/s single-stream decode. ADMITTED CLAIM 2 — served-path spec-on/off throughput ratio ~1.0: matched 10/10 rung sweep (rungs 0/32768/131072/204800, ±256 depth-fill probes) of the same server with and without --kairic-edge; tg 13.1-13.2 t/s flat across the entire depth ladder in BOTH arms, per-cell deltas within noise. Includes d32768, stable in both states (the prior served sweep crashed there). --kairic-edge buys NO served-path decode throughput on this workload. Energy counterpoint (eng-0268 vs eng-0269): the spec=on band drew 17.1 Wh vs 22.2 Wh spec=off for identical tokens (~23% less wall energy); one observation, not an established effect. ADMITTED CLAIM 3 — d32768 stability: previously-crashing fill now completes rc=0 in both spec states. Production-config validation: build e1da26bb8 ("version: 80"), binary e121a4d3, gdn_output_fallback=48 (tau2 server) / 0 (all 20 sweep cells), served id qwen3.8-27b-kairic-edge — all re-verified on-box this ingest. BLOCKED CLAIMS (explicitly NOT made): no reasoning-ON reward comparison (prior full-26 incomplete/do-not-score under G6); no reasoning-ON server throughput claim (crashed). No production recommendation flips on energy alone from a single band pair.
- Ingest of t_b7766d91 under the reviewer boundary from t_3b03c623. Every number re-verified against raw aihydra artifacts (/home/aihydra/bench-results/hk001-kairic-rerun-reasonoff-r1/) this session: tau2-reasonoff-full26-summary.json (mean_reward 0.4230769…, 11/26, empty_assistant_turns=0, infrastructure_errors=0), specsweep-summary.txt (20/20 cells, tg flat 13.1-13.2), fingerprint-full26.txt (binary/model/build hashes match cfg-0164 lineage). All 21 energy windows joined from HA history same-day (well within ~10d retention; deadline was 2026-09-02) — no unjoined_pending_ingestion records needed. [2026-08-24 COMPARABILITY ANNOTATION — HK-V12-ANNOTATE, per t_b05473c9] Recorded runs used the v1.1 Kairic Edge runtime, whose native IU4 M65 verifier could flip a greedy target token on low-margin reproduced cases; throughput figures are UNAFFECTED (GGUF/sidecars unchanged), but the correctness of recorded tau2 outputs carries this caveat. Re-runs for capability comparability (full-26 tau2) are justified and gated behind the v1.2 runtime tag. All future Kairic Edge runs MUST use kairic-edge-qwen38-27b-v1.2 @ 205a3e5f (strict-compact M65 verifier default).
- evidence: clm-0116 cfg-0172 run-0618 run-0619 run-0620 eng-0267 eng-0268 eng-0269
clm-0119
- measured-here high ●●● volatility medium · verified 2026-08-23
- HK-001-REASONING-MATRIX verdict on the Kairic Edge TheRock build of Qwen3.8-27B-IU4 (cfg-0173) — 5-cell tau2 smoke matrix isolating the reasoning lever on identical tasks/seeds/judge. Verdict: ADMIT-with-corrections applied verbatim (HK-MATRIX-REVIEW, t_8869ab61). ADMITTED CLAIM 1 — reasoning-ON improves smoke capability: all four reasoning-ON cells (R1-B6k unlimited; R2-B8k unlimited; R3-B8k-L budget=1024; R4-B8k-M budget=2048) scored 0.8 mean reward (4/5) vs reasoning-OFF R0-B4k at 0.6 (3/5). Same tasks 0-4, seed 42, claude-haiku-4.5 simulator — comparison fair. ADMITTED CLAIM 2 — budget capping preserves quality at this depth: capped cells R3 (1024) and R4 (2048) match unlimited R1/R2 at 0.8. ADMITTED CLAIM 3 — operational hygiene held across the matrix: 5/5 cells rc=0 with ZERO infra_errors and ZERO empty_assistant_turns; guard 4/4 PASS on R0-B4k, waived for R1-R4 under a documented binary-invariant rationale (model binary/sidecars/quant unchanged). ADMITTED COST SIGNAL — reasoning-ON roughly doubles wall clock (1808-1942 s vs 901 s) and wall energy (~84-92 Wh vs 43.2 Wh per cell; matrix total 394.4 Wh, eng-0270..eng-0274). RECOMMENDED CONFIG (reviewer): R2-B8k (max_tokens=8192, reasoning=on, budget=-1) for the full-26 confirmation run. REVIEWER CORRECTION applied verbatim: R0 smoke 0.6 differs from the full-26 reasoning-OFF baseline 0.423 (t_2330c08d lineage) — expected 5-task vs 26-task variance, not a protocol issue. BLOCKED CLAIMS (explicitly NOT made): these are 5-task smoke results — no full-26 capability claim, no production recommendation flip beyond the reviewer's config pick for the NEXT run, no energy effect claims beyond the single-matrix observation above.
- Ingest of HK-MATRIX-INGEST under the ADMIT_WITH_CORRECTIONS boundary from t_8869ab61. Every number re-verified against raw aihydra artifacts (/home/aihydra/bench-results/hk001-kairic-reasoning-matrix-r1/) this session: tau2-{cell}-results.json per-task rewards ([1,0,0,1,1] R0; [1,1,1,0,1] R1-R4) match matrix-table.tsv and matrix-summary.md exactly; fingerprint.txt matches cfg-0164/cfg-0172 lineage (binary e121a4d3, model 360caf73, build "version: 80 (e1da26bb8)", gdn_output_fallback=48); guard-waiver.txt rationale recorded. All 5 energy windows joined from HA history same-day (well within retention; cutoff ~2026-08-28 not approached) — no unjoined_pending_ingestion records needed. [2026-08-24 COMPARABILITY ANNOTATION — HK-V12-ANNOTATE, per t_b05473c9] All five matrix cells ran on the v1.1 Kairic Edge runtime, whose native IU4 M65 verifier could flip a greedy target token on low-margin reproduced cases; throughput figures are UNAFFECTED (GGUF/sidecars unchanged), but the correctness of recorded tau2 outputs carries this caveat — the reasoning-ON vs OFF smoke deltas could shift under the v1.2 strict-compact verifier. Capability-comparability re-runs (the matrix and the full-26 confirmation at the reviewer's R2-B8k pick) are justified on v1.2. All future Kairic Edge runs MUST use runtime tag kairic-edge-qwen38-27b-v1.2 @ 205a3e5f (strict-compact M65 verifier default; unsafe native path only via KAIRIC_UNSAFE_NATIVE_M65_VERIFY=1).
- evidence: clm-0116 clm-0118 cfg-0173 run-0621 run-0622 run-0623 run-0624 run-0625 eng-0270 eng-0271 eng-0272 eng-0273 eng-0274
clm-0120
- measured-here med ●●○ volatility medium · verified 2026-08-24
- A Qwen3.8-27B ROCmFPX-Q4 plus DFlash2 server screening sample reported 52.2 t/s on a varied code prompt and 31.5 t/s on a varied prose prompt. A separate five-task airline smoke completed 5/5 with mean reward 1.000. These are bounded source observations, not a production-throughput, quant-equivalence, or full capability verdict.
- The source is an operator-directed autonomous handoff, not an hbrunner run. The served-path samples use distinct prompts, but the handoff supplies neither the mandatory plain/llama-bench floor pair nor a house guard, artifact SHA256s, an energy window, or a blind-judge artifact for prose. The varied-prompt caveat is load-bearing: repetitive prompts inflate speculative acceptance. Do not convert this screening sample into a production figure.
- evidence: con-0013 cfg-0174 run-0626 run-0631 run-0627
clm-0121
- measured-here med ●●○ volatility medium · verified 2026-08-24
- A 45-turn Vulkan cache-reuse trace completed without a logged DeviceLost or lockup signature. Each returned row records only the newly prefetched prompt tokens, which remain approximately two thousand per turn, consistent with prefix reuse. The handoff describes the conversation as about 90k tokens, but its machine-readable trace does not record cumulative context; this is not a precise depth-capability boundary or production-latency result.
- Source: the autonomous handoff's 45-turn stdout, server log, and reproduction script. The script labels its displayed depth as prompt_n for the current turn, not total retained context. It provides neither a standard guard nor independent answer grading, and no energy window. Cold-load failure remains a separate result.
- evidence: cfg-0175 run-0628
clm-0122
- inferred high ●●● volatility medium · verified 2026-08-24
- No Qwen3.8-27B multi-slot serving or concurrency-throughput claim is admitted from this autonomous handoff. It preserves the isolation-test script but not its raw stdout or per-request responses, so the reported leakage and throughput values cannot be independently traced to a run record.
- The missing artifact is the conc_test.py output for each parallel setting, including marker ownership and cross-request bleed. Existing #25992 context is not a substitute for a configuration-specific result. Re-run the isolation guard before any -np > 1 use.
clm-0123
- inferred high ●●● volatility medium · verified 2026-08-24
- No Kairic Edge versus UD-Q4_K_XL head-to-head verdict is admitted from this autonomous handoff. Although it retains a Kairic server log, it does not provide the paired UD raw timing, tau2 output, footprint telemetry, runtime-tag evidence, or required floor/served-path pair needed for a capability or production comparison.
- The exact server-side feature state and any degradation marker are not captured in claimable form. Reproduce both configurations with their actual server flags, the required dual throughput cells, raw tau2 artifacts, and memory telemetry.
clm-0124 superseded
- measured-here high ●●● volatility medium · verified 2026-08-27
- Qwen3.8-Flash-Next UD-Q4_K_XL (Qwen4-arch preview MoE: 125B MoE / ~6B active + 51B N-gram/PLE tables + vision) on the Unsloth qwen4exp Vulkan build (cfg-0177, commit 250b6144, gfx1151/RADV) — first full benchmark. ADMITTED CLAIM 1 (capability leader): on the standard tau2 airline full-26 suite (seed 42, claude-haiku-4.5 simulator, deterministic reward) it scored mean_reward 0.9231, 24/26 tasks at reward 1.0 (run-0632). This LEADS the tau2 airline board — the prior best was Ornith-1.0-35B UD-Q4_K_XL at 22/26 (0.846). Tool-calling was clean: 184 tool calls, 0 empty-argument calls, 1 tool-error message, 24.7 messages/task. Failures = tasks 7 and 20. ADMITTED CLAIM 2 (energy leader per correct answer): wall-metered join (eng-0275, HA counter-difference) = 256.08 Wh total over 6605 s, 237.55 Wh active above the 10.1 W idle floor, = 9.90 Wh per correct answer (0.300 p @ 30.3 p/kWh). This is ~4x cheaper per correct answer than the ho003 stock-f16 tau2 control (42.06 Wh/correct), because it solves more than twice as many tasks in less wall time. ADMITTED CLAIM 3 (footprint): GTT-resident 76.7 GiB with ~43 GiB free for KV. The 26.8 GiB n-gram/PLE table (per_layer_token_embd, iq4_nl) is placed on CPU automatically by the Vulkan backend regardless of -ot, so it never occupies GTT; KV scales ~48 MiB per 1000 tokens (hybrid QSA attention). Performance: prefill 277-313 tok/s at 5-8K, decode ~22 tok/s short falling to 8.5 at 131K; TTFT ~2.9 s. BLOCKED CLAIMS (explicitly NOT made): single trial (N=1), no variance bound. reasoning_effort=low, not the model's default xhigh — a higher-reasoning result is untested and may differ (both tasks 7 and 20 may be reasoning-recoverable). The runtime is a draft/WIP PR (#27742), not a released/immutable build; the GGUF's advertised LICENSE artifact 404s and conversion-base provenance is unpinned, so this is a LAB result with NO weight redistribution or public/production recommendation. MTP speculative decoding is unavailable (the GGUF lacks MTP head layers) and n-gram speculation did not help this quant. 262K context hits a Vulkan workgroup-count assertion (cap -c <= 262140). No UD-Q2 control or cross-host pair run yet.
- First-look benchmark of Qwen3.8-Flash-Next, run under lab custody of AI Hydra. Capability + energy admitted; the WIP- runtime / license / N=1 / reasoning-low caveats are load-bearing and are carried into the candidate page verdict.
- evidence: cfg-0177 run-0632 eng-0275
clm-0125
- measured-here med ●●○ volatility medium · verified 2026-08-28
- The Qwen3.8-Flash-Next records from 2026-08-27 remain visible as a provisional first look: cfg-0177 and run-0632 record the reasoning-low capability arm, cfg-0178 and run-0633 through run-0636 record the non-served engine-floor depth ladder, and eng-0275 records the joined wall-meter window. They are real observations, not erased failures, but they are not admitted as the current capability leader, energy-per-correct leader, production-throughput result, or role-fit recommendation. The first look did not retain a protocol-valid evidence-bearing screen, formal guard, repeated capability series, exact full-run serving transport, or an unambiguous independently admitted energy denominator. A new intended/default-reasoning run is being prepared to supersede the decision boundary while preserving this history.
- Interim current-best-information boundary. The underlying first-look data remain public and citable as historical observations; only their headline admission is held pending protocol-valid superseding evidence.
- evidence: cfg-0177 cfg-0178 run-0632 run-0633 run-0634 run-0635 run-0636 eng-0275 clm-0124
clm-0126 superseded
- measured-here high ●●● volatility low · verified 2026-08-29
- On AI Hydra's admitted HIP/gfx1151/native/Release host boundary, the KingJones R2 Qwen3.8-Flash-Next full-STRIX ROCmFP4 campaign fixed model artifact c6770d7442a06bf1d78edf28cec83e1ec93afdd34664c23ff898807b6b9349fa (121838036032 bytes, publisher revision 069dddb53bab04218d734fa9a771f8a0242ab059), runtime source 36e9acd40e10a87cd3c3ef8ec734668757dc8520, and the independently admitted post-route receipt patch 1f1c3bed910922b415c1be36c9a04c9b7fedc4162aedf504a9da007daebcb4d2. The exact tracked 1,017-byte rotate-bits header was present at blob 75c4881fc322f2e6a6ee9d809e696852531abb8c and the patch clean-applied at +109/-2. The card-authorized targeted llama-server capture build, not an all-target build, then failed because sha256.c could not resolve rotate-bits/rotate-bits.h. Therefore no receipt-backed guard, Tau2 capability, llama-bench floor, served-path, cache, depth or energy result exists for this campaign. This is not a model quality, fit, throughput, capability or source-runtime performance claim. The earlier Unsloth UD-Q4_K_XL/Vulkan records are a separate artifact, runtime and backend series and are not merged with this full-STRIX result.
- R2 terminal negative at the evidence-capture build gate. No model process or energy-bearing arm began, so an energy record would invent a metered model run for this R2 variant; the exact non-energy disposition remains on run-0637.
- evidence: cfg-0179 run-0637 inc-0009
clm-0127
- measured-here high ●●● volatility low · verified 2026-08-30
- The KingJones Qwen3.8-Flash-Next full-STRIX ROCmFP4 series is admitted only as BOUNDED/PARTIAL on AI Hydra (cfg-0180/cfg-0181). Exact identity is artifact c6770d7442a06bf1d78edf28cec83e1ec93afdd34664c23ff898807b6b9349fa, runtime 36e9acd40e10a87cd3c3ef8ec734668757dc8520, Ubuntu 26.04, gfx1151 and HIP 7.1.52801-9999. The corrected one-slot served cohort used c262144, f16 K/V, b2048/ub512, thinking off, 512 context checkpoints and a 2048-token checkpoint cadence under the disclosed runtime-default-mmap exception. Standard Tau2 Airline smoke was 5/5. The full seed-42 N=1 arm attempted all 26 tasks: 24 were evaluated, 19 passed and five scored zero; tasks 23 and 24 were infrastructure errors with no score. Therefore 19/24 is the evaluated result and 19/26 is all-attempt pass coverage, not a zero-imputed suite score. The 24 evaluated trajectories carried 245 tool calls, zero empty-argument calls, zero empty assistant turns without a tool, and two tool outputs marked error inside otherwise scored-pass trajectories. The 8,192-token ceiling was inherited from clm-0119's different five-task Qwen3.8-27B IU4/Kairic Edge matrix, where 6k and 8k both scored 4/5 and 8k was selected as headroom; it was not tail-tested for this full-26 STRIX arm. An original-log audit found exactly two final calls at that ceiling, tasks 23 and 24. A later two-task 16k recovery is held because its retained transport was not lossless, so no mixed-budget aggregate amends this strict 8k result. A separate post-hoc wall-meter join (eng-0276) recovers 419.1478 Wh total and 20.5395 active Wh per 19 correct answers from retained exact UTC edges. It is not campaign-time metering compliance: the dated 10.1 W idle baseline and whole-box attribution remain explicit. Corrected cache_prompt:false prompt processing is admitted at exact server prompt_n 8192, 32768, 65536, 131072 and 260096 (run-0640..run-0644). Generation rates are admitted only for nonzero-output 32768, 65536 and 131072 rows. The zero-byte 8192/260096 synthetic predicted rates are excluded. Full-window d262144 returned HTTP 400 without timing or prompt_n (run-0645). Controlled 32k and 131k prefix-reuse traces are separate records (run-0646/run-0647), not ladder repetitions. Representative whole-file mincore, GTT, MemAvailable, page-cache and process receipts remain bounded observations on the cited runs. Per-PLE mincore is incomplete: 30 of 40 before/after records failed after five of six layout tensors, so no PLE fault-tracking, complete PLE working-set or resident-page claim is made. Raw mmap engine rows remain visible (run-0648..run-0652), but d0/d32768 prompt-processing CVs of 3.731412%/3.131244% exceed the protocol trigger; no unrestricted aggregate floor, best-config or cross-model claim is admitted. Thermal observations reached 98 C but never the confirmed >=100 C stop. The exact owner released cleanly with no lease or model process. Publisher comparison is attribution-bounded by con-0017: that immutable page used ROCm 7.2.4, different prompt depths/context allocation and checkpoint evidence, and STRIX_LEAN for its c262 rows. Similar 32k-131k magnitudes are a caveat, not equivalence or causal attribution to a runtime version.
- Current-best bounded record. It supersedes the earlier receipt-build terminal boundary without deleting it, and it does not create a production recommendation. Prohibited claims remain explicit: no complete PLE series, no zero-byte generation claim, no mixed-budget Tau2 aggregate, no campaign-time metering-compliance claim, no unrestricted engine-floor aggregate, and no direct publisher equivalence.
- evidence: cfg-0180 cfg-0181 run-0638 run-0639 run-0640 run-0641 run-0642 run-0643 run-0644 run-0645 run-0646 run-0647 run-0648 run-0649 run-0650 run-0651 run-0652 eng-0276 inc-0010 inc-0011 inc-0012 con-0017 clm-0119 clm-0126
clm-0128
- measured-here high ●●● volatility medium · verified 2026-09-03
- Qwen3.8-Flash-Next ROCmFP4-FAST-v2-ple16 (agention imatrix quant, 87.06 GiB @ 4.23 bpw, everything GPU-resident incl. the per-head n-gram/PLE table) on the agention/Laurent Vulkan fork (cfg-0182, LaurentZuijdwijk/llama.cpp branch vulkan/qwen4exp-rocmfpx commit 5e085d12, build b10809, gfx1151/RADV) — first benchmark of the ROCmFP4 quant path on this fleet. ADMITTED CLAIM 1 (speed holds at depth, served path): single-slot with adaptive MTP, temp 0, source-code content, tokenizer-exact depths — decode 16.1 / 16.0 / 22.8 / 26.1 tok/s at 32K / 64K / 128K / 200K, prefill 298 -> 127 tok/s (run-0654..run-0657). Decode rises with MTP acceptance (0.49 -> 0.86 across the sweep) rather than falling with depth, i.e. it HOLDS to 200K on code; the ordering is N=1-noisy. Real agentic decode averaged ~34 tok/s over the tau2 run (run-0653); shallow generated code (red-black tree, JSON) reaches 44-47 tok/s at 0.86-0.91 acceptance; prose is lower (~18-22 tok/s, ~0.5 acceptance). Warm agent-turn TTFT ~0.6 s (cached prefix). ADMITTED CLAIM 2 (fits fully on GPU, stable): 87.06 GiB weights stay GPU- resident (GTT ~95-106 GiB at 1 slot, ~16-27 GiB free for KV), loads in 57-64 s, no host-RAM thrash and no OOM. The per-head ple16 layout keeps the ~51B n-gram/PLE table on GPU (no -ot offload); a joined-table variant with --ngram-on-disk exists for off-GPU placement but was not needed. 3-slot + adaptive-MTP loads and serves stably (no crash), but decode is bandwidth- bound so concurrency does not beat single-slot aggregate at depth. ADMITTED CLAIM 3 (capability, with a matched-comparison caveat): tau2 airline full-26 (run-0653, seed 42, claude-haiku-4.5 simulator, deterministic reward) scored 21/25 = 0.84 (26 attempted, task 10 excluded as an infrastructure error with 0 messages — a user-simulator/cloud fault). All 25 scored terminated clean; tool-calling and 196k single-slot context served the full agentic workload with no grammar or overflow failures. This sits BELOW the same-suite UD-Q4_K_XL result (run-0632, 24/26 = 0.9231) and the Ciru-IU4 result (23/26) — but it is NOT reasoning-matched: this run used the model's DEFAULT reasoning effort in the deployed serving config, whereas run-0632 used reasoning_effort=low. ADMITTED CLAIM 4 (energy, same caveat): wall-metered join (eng-0277, HA counter-difference) = 286.3 Wh over 6701 s, 267.5 Wh active above the 10.1 W idle floor, = 12.74 Wh per correct answer (0.386 p @ 30.3 p/kWh). This is HIGHER (worse) than UD-Q4_K_XL's 9.90 Wh/correct (eng-0275) — the raw-decode speed advantage does NOT translate to energy-per-correct, because the default reasoning effort generates more tokens and the run solved fewer tasks (21 vs 24). A reasoning=low re-run is the matched comparison. BLOCKED CLAIMS (explicitly NOT made): single trial (N=1), no variance bound; the decode-depth cells are N=1 served-path measurements, not the n=3 llama-bench ladder used for cfg-0178, and MTP acceptance (hence decode) is noisy cell-to-cell. NOT reasoning-matched to run-0632 (default vs low), so the capability and Wh-per-correct comparisons conflate quant, reasoning depth, and serving-vs-bench config — no "faster AND better/worse" conclusion is drawn. Quant quality "within 2.5% PPL" is a vendor (agention) claim, not independently perplexity-measured here. Lab result under agention's qwen-community license; the fork is a third-party WIP branch (verified branch tip, but not a released/immutable upstream build). This quant IS in production use on the fleet's serving box as of this date, but that is an operational choice for speed, not a capability endorsement over UD-Q4_K_XL.
- First benchmark of the agention ROCmFP4-FAST-v2 quant, under lab custody of AI Hydra, driven by the same operator flow as the UD-Q4_K_XL entry. The headline is speed-at-depth + clean GPU-resident fit; the capability and energy numbers are admitted but carry load-bearing caveats (N=1, default reasoning not matched to the UD-Q4_K_XL reasoning=low baseline). The obvious next step is a reasoning=low, n=3 matched re-run before any "which quant is better" claim.
- evidence: cfg-0182 run-0653 eng-0277
clm-0129
- measured-here high ●●● volatility medium · verified 2026-09-06
- On its proposed dense-WORKER serving config (cfg-0183: UD-Q4_K_XL, the pr27311 leak-fix build, self-speculative draft-mtp n_max 2, --parallel 4, reasoning_effort low, greedy, fronted by the slotpin proxy), Qwen3.8-27B scored 91.7% on tau2-bench airline tasks 0-25 — 22 of 24 scored passed, run end-to-end through the proxy under a real 4-slot concurrent workload, with 0 empty assistant turns over 106 minutes. Two of the 26 attempted tasks were infrastructure errors (empty AssistantMessage from the cloud user-simulator, openrouter/haiku-4.5) and are excluded from the denominator, as run-0271 excluded its too_many_errors task. Harness rollup: Read Actions 24/24, Write Actions 26/27, DB Match 23/24; LLM-judge agent errors 0. This is NOT a clean A/B against the 2026-08 screen's 0.577 (run-0271, same harness 668d3bc, same seed 42, same tasks 0-25): four fingerprint fields changed at once — quant (Q8_0 -> UD-Q4_K_XL), build (3653e6d -> pr27311), reasoning_effort (medium -> low) and agent max_tokens (4096 -> 8192), plus --parallel (1 -> 4). The leading hypothesis for the gap is the output cap — run-0271's 4096-token cap under reasoning=medium plausibly truncated agent turns mid-tool-call, which reasoning=low with an 8192 cap avoids — but that is a hypothesis, not an isolated measurement.
- METHOD — tau2-bench airline, harness 668d3bcd135c02aa3438f987ef45735b7c163ee3, --domain airline --seed 42 --num-trials 1 --task-ids 0..25 --max-concurrency 4 --max-steps 100, user simulator pinned to openrouter/anthropic/claude-haiku-4.5 (temperature 0). Agent: local llama-server on cfg-0183 (pr27311 build c530ea79c, UD-Q4_K_XL, self-spec draft-mtp n_max 2, -c 49152, --parallel 4), agent sampling temperature 0.0 / max_tokens 8192, reasoning_effort low set both at the server (--reasoning-effort low) and in agent-llm-args. Window 2026-09-05 13:41:58-15:27:37Z. reasoning=low VALIDATED FOR CAPABILITY, not just speed — the worker default holds full competence at this task set, so choosing it for throughput does not trade away the agentic score. n_max 2 is safely under this model's n_max>=4 EOS-cliff ceiling (clm-0055). CONFIDENCE high on the number as a measurement of cfg-0183 (n=24 scored, 0 empty turns, 100-step ceiling never truncating a task at max_steps); MEDIUM would apply to any claim that a single lever CAUSED the lift over run-0271, which this deliberately does not make. OPEN QUESTION — a matched single-lever isolation (hold quant/build/parallel, vary only max_tokens; then only reasoning_effort) would attribute the 0.577 -> 0.917 movement. Until then the two numbers are two honest measurements of two different configs, not a before/after. ENERGY: 15.18 Wh per correct answer (eng-0278), see clm-0131.
- evidence: run-0658 run-0660
clm-0130
- measured-here high ●●● volatility medium · verified 2026-09-06
- The Qwen3.8-27B dense worker earns its slot on THROUGHPUT UNDER CONCURRENCY, not single-stream speed, and only on the leak-fixed build. Served single-stream decode falls off gently with depth and never cliffs: 19.2 / 17.0 / 14.7 / 12.5 t/s at 32k / 65k / 131k / 205k, prefill 239 / 141 / 74 / 50 t/s, self-spec MTP acceptance 0.76-0.84 (run-0661..run-0664, cfg-0184, single slot). Under a real 4-slot agentic workload the same config sustains ~12.8 t/s per slot at mean concurrency 3.6 — a derived aggregate of ~46 t/s of useful work (run-0659). BINARY BUILD FINDING: this multi-slot behaviour requires PR #27311. The identical flags on fresh master c5a5535e degraded to EMPTY output within ~4 minutes of 4-slot load (the #25992 cross-request leak; monitor 13:34:25Z OK -> 13:38:29Z EMPTY), whereas the pr27311 build (c530ea79c) held 0 EMPTY across the full 106-minute run (run-0660). The earlier Q8_0 serve config pinned --parallel 1 precisely to avoid #25992; pr27311 is what makes a multi-slot dense worker viable at all.
- DEPTH SWEEP — served path via the slotpin proxy (drive_depth_served, code content, tokenizer-trimmed to exact depths, 256-token generations), cfg-0184, --parallel 1, pr27311 build. Monotonic, no device loss, no cliff to d204800; d262144 (model-max) was not taken on the served path (prompt+gen exceeded -c and returned HTTP 400), so the single-stream served ceiling recorded here is d204800. MULTI-SLOT — read from proxy metrics.jsonl for the τ² window (run-0658): 469 decode turns, per-slot mean 12.8 t/s (min 1.9, peak 23.7), concurrency mean 3.6 with 350 of 471 rows at 4.0. Aggregate ~46 t/s is per-slot-mean x mean-concurrency on a VARIED workload — not the shared-n-gram-pool artifact identical-prompt multi-slot would give. BUILD HAZARD — PR #27311 "Scheduler UMA ring buffer (+ sanitizer and fixes)" is an OPEN, unmerged upstream fix for the gfx1151 UMA async-output race behind #25992; it also restores speculative-decode acceptance under concurrent MTP (#27572). Carried on aihydra as branch pr27311 (con-0018). server_flags_degraded is false ON THIS BUILD; it would be true on stock master under --parallel>1, which is the whole point of the guard (run-0660). Serving this worker multi-slot on a non-pr27311 build is a data-integrity hazard, not a performance preference.
- evidence: run-0659 run-0660 run-0661 run-0662 run-0663 run-0664 con-0018
clm-0131
- measured-here med ●●○ volatility medium · verified 2026-09-06
- The Qwen3.8-27B worker config cost 15.18 Wh per correct answer on the tau2 0-25 run (333.87 Wh whole-session across 22 correct, mean 189.6 W, 0.46 pence at 30.3 p/kWh; eng-0278) — well under the 38.29 Wh per correct of the earlier single-slot Q8_0 run (eng-0081). Two things drive the direction, neither a clean A/B: serving --parallel 4 amortises the box's ~190 W draw across ~4 concurrent tasks, and the run answered more tasks correctly (22 vs 15). Quant, build, reasoning and parallelism all differ between the two windows, so this is an efficiency observation about the whole worker profile, not an attribution to any single lever.
- Wh-per-correct = wh_total / tasks_passed, whole-session (includes user-simulator wait and inter-task gaps), the same method as eng-0081 so the two are comparable. Energy differenced from the aihydra HA cumulative-kWh counter across the recorded run window. Confidence is MEDIUM because the comparison to eng-0081 is confounded (four fingerprint fields differ) and because the delta_w rests on the canonical 10.1 W idle floor while a spot idle reading this session was ~20 W — a fresh post-fan-control idle-baseline would refine delta_w (not wh_per_task). The multi-slot amortisation effect itself is robust: a fixed ~190 W spread over more concurrent useful work is fewer Wh per answer, which is the general reason a worker that actually receives concurrent tasks is cheaper per task than one driven single-stream. See clm-0130 for the throughput side.
- evidence: eng-0278 run-0658
clm-0132
- measured-here high ●●● volatility medium · verified 2026-09-06
- The reproduced pwilkin/ilintar Strix Halo stack (IQ4_XS-imatrix + DFlash2 draft on the retained-PM4 runtime, cfg-0185) scored 83.3% on tau2 airline 0-25 — 20 of 24 scored passed, 4-slot, reasoning low, 2 cloud-user-sim infra errors excluded — versus the worker's 91.7% (run-0658) on the identical task set, harness (668d3bc), seed, reasoning and concurrency. So the worker holds an 8-point capability lead. The DFlash2 correctness gate was clean: 0 empty outputs across 16 probes over the 66-minute run (run-0666), so the prior DFlash2 sentinel-empty concern did not recur at --parallel 4. This is NOT a clean single-lever A/B: quant (UD-Q4_K_XL -> IQ4_XS-imatrix), speculation (self-spec MTP -> DFlash2 draft model) and build/runtime all differ at once. Separately, the build reproduced the author's own headline (31.5k prompt, 256 gen, DFlash2 width 3): decode 26.9 t/s vs the claimed 26.256, prefill 325.7 vs 256.84, acceptance 0.573 vs 0.607.
- METHOD — tau2-bench airline, harness 668d3bc, --seed 42 --num-trials 1 --task-ids 0..25 --max-concurrency 4 --max-steps 100, user-sim openrouter/anthropic/claude-haiku-4.5 (temp 0). Agent: pwilkin llama.cpp d3b5cc4 + custom retained-PM4 runtime, IQ4_XS-imatrix target + DFlash2 draft, temperature 0 / max_tokens 8192, reasoning_effort low. Direct :5803 (capability is proxy-invariant, so comparable to the worker's proxy run-0658). Window 09:27:14-10:33:13Z. Reproduction of the author headline used repro_bench at a 31,497-token prompt with DFlash2 width 3. The 8-point gap is genuine capability, not degradation (0 empties, no max-steps cuts): the IQ4_XS-imatrix quant and/or DFlash2 spec path cost competence that the worker's UD-Q4_K_XL + self-spec MTP retains. Energy: 10.26 Wh per correct answer (eng-0280), clm-0134. Throughput: clm-0133.
- evidence: run-0665 run-0666
clm-0133
- measured-here high ●●● volatility medium · verified 2026-09-06
- The pwilkin candidate wins single-stream throughput decisively but only ties the worker under concurrency. Served single-stream decode: 27.98 / 23.18 / 22.12 / 17.07 t/s at 32k / 64k / 131k / 205k (cfg-0186) — a consistent +37% to +50% over the worker's 19.23 / 16.99 / 14.71 / 12.49 (cfg-0184) across every depth, monotonic with no cliff (and no Vulkan-style device loss, since this is the HIP/retained-PM4 build). But multi-slot aggregate reaches 47.3 t/s at 4 concurrent slots (run-0667) — essentially tied with the worker's ~46 (run-0659) — and it gets there by scaling only 1.3x from a high single-stream base (36.5 -> 47.3 @1->4) with very uneven per-slot rates (DFlash2 draft contention), whereas the worker scales ~2x evenly from ~22. So pwilkin front-loads latency-per-task; the worker scales with concurrency; they converge at 4-slot aggregate.
- Single-stream depth via slotpin (drive_depth_served, code content, 256-token gens), cfg-0186, --parallel 1, DFlash2 width 3, retained-PM4 build. DFlash2 acceptance rose with depth (0.62 -> 0.79). Multi-slot: 4 varied ~4k-prompt requests fired concurrently via slotpin; aggregate decode summed across slots. The single-stream lead is real and depth-robust; the multi-slot tie is the operative fact for a worker that receives concurrent tasks. Why the lead exists is isolated in clm-0135 (it is the quant + spec, not the retained-PM4 runtime).
- evidence: run-0667 run-0668 run-0669 run-0670 run-0671
clm-0134
- measured-here med ●●○ volatility medium · verified 2026-09-06
- The pwilkin candidate cost 10.26 Wh per correct answer on its tau2 0-25 run (205.2 Wh whole-session across 20 correct, mean 187 W, 0.31 pence at 30.3 p/kWh; eng-0280) — below the worker's 15.18 (eng-0278). The direction is not a clean A/B: the pwilkin stack decodes faster so its window was shorter (66 vs 106 min -> less total energy), even though it answered fewer tasks correctly (20 vs 22). So it is cheaper per correct answer but on a lower capability base — the same speed/quality trade seen everywhere in this comparison, expressed in energy.
- Wh-per-correct = wh_total / tasks_passed, whole-session, same method as eng-0081/eng-0278 for comparability. Energy differenced from the aihydra HA cumulative-kWh counter over the recorded window. Confidence medium: the comparison to eng-0278 is confounded (quant, spec, build all differ) and rests on the canonical 10.1 W idle floor. The efficiency direction is robust (a faster decode over a fixed idle floor is fewer Wh per answer), but it buys that efficiency at the 8-point capability cost of clm-0132.
- evidence: eng-0280 run-0665
clm-0135
- measured-here high ●●● volatility low · verified 2026-09-06
- The retained-PM4 runtime is a real but small, lossless dispatch optimisation — NOT the source of the pwilkin speed advantage. Running OUR worker model (UD-Q4_K_XL + self-spec MTP, cfg-0187) on the pwilkin build with retained-PM4 toggled: 20.17 t/s at 32k with it ON (run-0672) vs 19.45 OFF (run-0673) — +3.7%, identical acceptance (0.766), same binary and model so the delta is purely the runtime. That ~20 t/s ceiling for our high-capability model sits far below the pwilkin candidate's 27.98 at the same depth (clm-0133). Therefore the ~45% single-stream lead comes from the IQ4_XS quant (far fewer bytes read per token than UD-Q4_K_XL) plus the DFlash2 draft — i.e. exactly the levers that cost the 8 capability points (clm-0132) — and cannot be recovered by adopting the runtime. The retained-PM4 change is also the one piece of the pwilkin stack with no upstream PR: the author states he does not expect it to be accepted and may retire it for HRX support (con-0019); everything else is upstream PRs (e.g. #27311, con-0018).
- Clean isolation: cfg-0187 is our exact worker model + self-spec draft-mtp n_max 2 on the pwilkin binary + custom retained-PM4 HIP runtime; run-0672 (ENABLE_RETAINED_PM4=1 -> DEBUG_HIP_GRAPH_PM4=1, HIP graphs on) vs run-0673 (=0, graphs disabled) differ only by that env toggle. PM4-off (19.45) matches our worker's own baseline (19.23, run-0661), confirming the runtime is the only variable. Consequence for adoption: there is no free lunch — the worker's speed cannot be lifted to pwilkin's by the runtime; buying that speed means accepting the lower-capability IQ4_XS + DFlash2 stack. The remaining unexhausted lever is a higher-capability fast quant (a calibrated IQ4 that holds competence closer to Q4_K_XL), which is a quant-quality question, not a runtime/spec one.
- evidence: run-0672 run-0673 con-0019
Incidents (9)
inc-0004 — Unified-memory OOM cascade to kernel panic, then a wedged boot
- 2026-07-21 · class hardware · severity blocked · resolved
- detected_by: smart-plug-power-flat-at-45W (lag 46m)
- should_have_caught_it: The box was OOM-crash-looping for 46 minutes before the panic — three llama-server kills between 09:05 and 09:51 — and nothing surfaced any of them. Worse, the morning brief had flagged "does cache-ram 3072 hold?" as a watch item that same day: the risk was identified and then left unmonitored. A process-restart counter or a free-memory alarm would have caught it with most of an hour to spare.
- signature: Smart plug flat at 45 W with no network response; power draw that had been dynamic all morning stopped varying.
- initial: a single OOM kill, expected to self-heal via the restart → actual: Kernel OOM deadlock. An amdgpu svm_range_restore_work worker needed pages; the OOM killer had already swept the process table; llama-server's memory is GTT-pinned and unreclaimable so killing it freed nothing ("oom_reaper: unable to reap"). With no killable processes the kernel panicked with "System is deadlocked on memory". The post-kdump reboot then wedged at a flat 45 W with NPU SMU init errors, a state only an AC pull clears. (false path cost none material)
- blast radius: runs 0 · configs 1
- lesson: On unified memory an OOM is a recovery problem, not a performance problem — it takes the whole box down and needs physical or out-of-band access to clear. Headroom is therefore a stability metric. Mitigations applied: --ctx-checkpoints capped at 4 (was defaulting to 32), desktop stack removed, swap raised to 16G, llama-swap/slotpin restart pairing installed. The ratio at failure was 83.8% of pool, which is the first real calibration point for the OOM warn line.
inc-0005 — Abrupt power loss under sustained load; board never POSTed again — RMA closed (full refund)
- 2026-07-23 · class hardware · severity blocked · resolved
- detected_by: smart-plug-statistics-retrospective (lag Warden itself noticed in ~71 seconds (failover at 09:58:24 after the 09:57:13 attempt). The warm-lane canary did not flag it until 06:33 the NEXT DAY — a ~20 hour gap during which nothing escalated. Warden knew; nothing told anyone. )
- should_have_caught_it: Nothing alerted. The box simply stopped answering and the outage was noticed by its absence. A plug-power threshold alarm — "AI Power fell below 5 W while a model was supposed to be resident" — would have fired within minutes, and the same sensor that diagnosed this retrospectively could have raised it live. The data existed the whole time; nobody was looking at it.
- signature: Hourly statistics show active operation (mean 41.3 W, peak 178.4 W) at 09:00 UTC, then mean 1.0 W / max 3.7 W at 10:00 UTC. Five hours at 0.5-0.6 W, then ~10 W from 15:00 UTC onward, flat ever since and slowly declining to ~5.2 W by 5 Aug.
- initial: same failure as the 2026-07-21 OOM panic → actual: A DIFFERENT failure. The 21 July event wedged at a flat 45 W — powered, hung after POST. This one collapsed to ~0.5 W, which is BELOW the ~10 W standby the board draws today, so even standby power was absent for five hours. Recovery to ~10 W at 15:00 UTC is consistent with a plug cycle clearing a latched state back to standby. The board has not POSTed since. Most consistent with a power-delivery or protection latch under load rather than a software or configuration fault. (false path cost none — diagnosed remotely before travelling)
- blast radius: runs 0 · configs 1
- lesson: Two lessons. First: the mitigations applied after 21 July (--ctx-checkpoints 4, desktop stack removed, swap 16G) target dynamic memory growth and are almost certainly IRRELEVANT to this failure — a patch chosen for the wrong incident buys false confidence. Second and more useful: a smart plug is an availability instrument, not just an energy one. The power curve distinguished "wedged after POST" from "lost power entirely" retrospectively, from another country, with no access to the machine. That signal should be an alarm, not an archaeology tool.
inc-0006 — HG-002 smoke parser rejected a valid task with no required tool call
- 2026-08-19 · class code · severity data-integrity · resolved
- detected_by: summary-review-against-copied-tau2-results (lag immediate at arm close)
- should_have_caught_it: The parser fixture should have included a successful task whose scenario legitimately terminates without a tool call.
- signature: Five rewards and terminations, six valid nonempty tool calls, no empty turns, no infrastructure errors and mean 0.80, but rc 1 solely because task 1 had zero tool calls.
- initial: invalid smoke because one task lacked a nonempty tool call → actual: The non-vacuity rule applies to the smoke as a whole, not every task. Task 1 completed successfully without needing a tool; the harness parser produced a false negative after the scientific result was already complete. (false path cost n1 was not run in the first paired root; a corrected paired retry followed)
- blast radius: runs 0 · configs 1
- lesson: Capability validity gates must test whole-smoke non-vacuity unless the task contract itself requires a call. Preserve a scientifically complete result separately from the harness exit status and do not claim the arm formally passed until the corrected parser closes rc 0.
inc-0007 — HG-003 retrieval runner stopped on an unbound backend variable before model work
- 2026-08-19 · class code · severity data-integrity · resolved
- detected_by: runner-exit-and-supervisor-log (lag immediate)
- should_have_caught_it: A shell syntax/static check or a fixture executing the backend loop with nounset enabled should have caught the uninitialized variable before dispatch.
- signature: Artifact preflight passed, retrieval.tsv remained header-only, and the shell exited rc 1 at line 67 with "backend: unbound variable" in a zero-second run-meta window.
- initial: retrieval depth 2048 failed → actual: The runner failed before server launch or any model-dependent request. safe_depth=0 is only the runner's terminal sentinel and carries no model capability information. (false path cost one zero-second dispatched attempt; no scientific work completed)
- blast radius: runs 1 · configs 1
- lesson: Preflight must execute or statically validate every nounset-sensitive loop variable. Never translate a pre-model harness exit or safe-depth sentinel into a capability fail.
inc-0008 — HG-003 nominal depth prompt exceeded its server context before placement
- 2026-08-19 · class config · severity data-integrity · resolved
- detected_by: server-log-and-prompt-metadata-review (lag immediate at failed-attempt ingestion)
- should_have_caught_it: The runner should have tokenized the complete rendered request, including template and tool overhead, and rejected or resized it before starting the scientific arm.
- signature: The nominal depth-2048 prompt recorded 6691 prompt tokens; the server counted 6779 request tokens against n_ctx=6144 and returned HTTP 400 before placement/generation.
- initial: LFM2 failed retrieval at depth 2048 → actual: Repeated filler text expanded to far more model tokens than the nominal depth and the configured context could hold. The server rejected the request with zero tokens processed, so no retrieval observation or safe depth exists. (false path cost model load plus one rejected request; stage B was correctly skipped)
- blast radius: runs 1 · configs 1
- lesson: Retrieval depth must be defined and gated in model-token space after full request rendering. HTTP admission failures are runner invalidity, never capability failures.
inc-0009 — Q38FN R2 receipt-only capture variant stopped at header include resolution
- 2026-08-29 · class code · severity blocked · UNRESOLVED
- detected_by: capture-build-exit-and-compiler-log (lag immediate at the required build gate)
- should_have_caught_it: The R2 targeted-build preflight did surface the stop before model execution; clean patch application and a tracked-file check alone could not prove the patched include topology was buildable.
- signature: After the exact 1,017-byte tracked header was verified and the receipt patch clean-applied at +109/-2, the targeted Release HIP/gfx1151/native llama-server build exited rc 2 when examples/gguf-hash/deps/sha256/sha256.c could not include rotate-bits/rotate-bits.h.
- initial: Verifying the tracked header and limiting the build to the required llama-server target might clear the earlier mechanics ambiguity. → actual: The tracked header exists at the exact reviewed blob and the patch applied with the independently admitted three-file delta. The targeted compiler path still could not resolve rotate-bits/rotate-bits.h for sha256.c. No include-path or patch repair was authorized, so the evidence admits the exact resolution failure without claiming a broader root cause or testing model execution. (false path cost One targeted receipt-only llama-server build attempt; no all-target build, model server, model request, benchmark cell or energy-bearing arm began.)
- blast radius: runs 1 · configs 1
- lesson: A tracked dependency and clean patch application do not prove the patched include topology is buildable. Build the exact targeted receipt path before model execution; this pre-model stop is an instrumentation result, never a model quality or runtime-performance result.
inc-0010 — Full-STRIX campaign retained corrected QSA evidence with bounded instrumentation gaps
- 2026-08-30 · class data · severity degraded · resolved
- detected_by: independent-package-review-and-raw-receipt-audit (lag Configuration drift was isolated before the corrected cohort; incomplete PLE mincore and engine scatter were contained during package review before publication.)
- should_have_caught_it: The QSA launcher should have asserted thinking-off and checkpoint cadence before its first measurement; the PLE adapter should have validated every tensor range; the engine wrapper already caught the scatter through its CV trigger.
- signature: Earlier QSA attempts used default 32/8192 checkpoint semantics and thinking-on; corrected r1/r2 stopped before valid measurements. Corrected r3 produced admitted cache-false rows, but 30 of 40 before/after receipts had a six-layout/five-mincore mismatch with ValueError("mmap length is greater than file size"). Engine d0 and d32768 prompt-processing CVs were 3.731412% and 3.131244%.
- initial: One campaign could support capability, a comparable engine floor, served QSA, complete PLE fault tracking, cache traces and joined energy. → actual: Tau2 and the corrected thinking-off 512/2048 served cohort are interpretable; prefix reuse is separately interpretable. The earlier default-QSA attempts, incomplete per-PLE adapter and engine trigger rows cannot support their intended aggregate claims. An interrupted d240000 cell was repeated once, and an interrupted 131k cache warm-up was preserved before a distinct unchanged-config resume-r4 pair. Corrected r1/r2 launcher stops and every raw attempt remain in the immutable package. (false path cost Two pre-measurement corrected launch attempts, one interrupted engine cell, one interrupted cache warm-up, and package-only review rework; no valid model attempt was rerun for score.)
- blast radius: runs 0 · configs 2
- lesson: Corrected startup evidence can salvage a distinct served cohort without laundering earlier attempts. Incomplete per-tensor instrumentation and scatter above the house trigger must narrow claims even when whole-file memory receipts and individual raw rows remain useful.
inc-0011 — Authorized energy join handoff was omitted before full-STRIX evidence closure
- 2026-08-30 · class data · severity degraded · resolved
- detected_by: publication-hold-and-post-hoc-energy-recovery (lag The first publication package correctly exposed a structured-unjoined state, but review treated the runner's lack of Home Assistant authorization as a terminal evidence disposition instead of routing the retained edges to hborchestrator.)
- should_have_caught_it: Campaign closure should have required a predeclared meter entity, method and authorized join owner after exact UTC edges landed and before final evidence review.
- signature: The runner was told to join energy but had no authorized Home Assistant reader; package-r2 sealed an unjoined disposition even though exact UTC edges and live Home Assistant history remained recoverable.
- initial: A structured unjoined reason was sufficient because the hardware runner could not access the wall-meter history. → actual: Credential separation was correct; workflow ownership was not. Hbrunner should retain exact UTC edges and hand off the join to authorized hborchestrator before evidence review. Final review may accept joined receipts or an explicit reviewed unjoined reason, but may not silently close a pending handoff. (false path cost One publication package was held and a post-hoc evidence addendum was required; no benchmark rerun or model output change was needed.)
- blast radius: runs 0 · configs 1
- lesson: Keep Home Assistant credentials with hborchestrator. Every campaign must predeclare meter entity/method/join owner, persist exact UTC edges, and close an authorized join handoff before final evidence review. The executable closure check rejects pending handoffs and credential-like artifact material.
inc-0012 — Inherited 8k output ceiling bound two full-STRIX Tau2 calls
- 2026-08-30 · class data · severity degraded · resolved
- detected_by: two-task-budget-recovery-and-independent-transport-review (lag The original package retained tasks 23 and 24 as infrastructure errors. The later targeted review established the shared ceiling signature before publication, so the strict denominator was preserved and no mixed-budget score entered the public record.)
- should_have_caught_it: Full-domain preflight should have demonstrated that the chosen output ceiling was non-binding for the intended reasoning mode, and package review should have audited every final server timing row for an exact-cap token count.
- signature: Programmatic audit of the original server log found exactly two final eval rows at n_decoded=8192, corresponding to tasks 23 and 24; both original task attempts ended with empty assistant envelopes and no reward.
- initial: max_tokens=8192 was treated as ample reasoning headroom for a full Tau2 Airline arm. → actual: The 8k choice came from clm-0119's different five-task Qwen3.8-27B IU4/Kairic Edge matrix: 6k and 8k both scored 4/5, and 8k was selected as headroom for a future full run. That recommendation was generalized to the KingJones full-STRIX artifact without a tail test. A later two-task 16k recovery was retained but not admitted: its strace receipts preserve byte counts and ranges, not lossless OpenAI-compatible payloads, so they cannot independently separate model exhaustion from adapter/parser serialization. (false path cost A targeted two-task control/recovery package and documentary repair were retained, but no recovery score or mixed-budget aggregate is admitted and no original row was changed.)
- blast radius: runs 0 · configs 1
- lesson: Treat max_tokens as a non-binding safety ceiling, not inherited folklore. Predeclare it, inspect every final timing/usage row for structural cap hits, retain reasoning and visible output separately where supported, and keep strict fixed-budget results separate from any intended-config or recovery cohort. A recovery may amend a page only after its transport boundary is independently interpretable.