Home › Evidence › Records
Records
Values marked derived are computed at build time and never stored in YAML — rendering them makes every build a check. Every record id also has a stable per-record page — /records/<id>/, linked from each heading — carrying the full record and its computed backlinks; the #<id> anchors on this page keep resolving.
Nodes (4)
gpt-oss-120b — gpt-oss-120b @ Q4_K_M
- kind model · engine igpu · tier candidate · runs_on aibeast
- config → cfg-0003
- derived current_state: rejected
- derived days_in_production: 0
- lifecycle: candidate 2026-06-27 → rejected 2026-06-28
nemotron3-super — Nemotron 3 Super @ IQ4_XS (not quant-equivalent to peers)
- kind model · engine igpu · tier candidate · runs_on aibeast
- config → cfg-0004
- derived current_state: rejected
- derived days_in_production: 0
- lifecycle: candidate 2026-06-27 → rejected 2026-06-28
qwen35-122b-a10b — Qwen3.5-122B-A10B-MTP @ UD-Q4_K_M, single slot
- kind model · engine igpu · tier daily-driver · runs_on aibeast
- config → cfg-0002
- derived current_state: live
- derived days_in_production: 46
- lifecycle: candidate 2026-06-27 → live 2026-06-28
qwen36-35b-a3b — Qwen3.6-35B-A3B @ UD-Q4_K_XL, reasoning off
- kind model · engine igpu · tier daily-driver · runs_on aibeast
- config → cfg-0001
- derived current_state: retired
- derived days_in_production: 58
- lifecycle: live 2026-05-01 → retired 2026-06-28
Configs (31)
cfg-0001 — unsloth/Qwen3.6-35B-A3B-GGUF @ UD-Q4_K_XL
- host aibeast · engine igpu · runtime ggml-org/llama.cpp@unknown-2026-06 (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; llama.cpp commit, chat-template hash and the memory block were not recorded at test time. Quant, backend, context, KV quant, serving flags and the reasoning setting are from the Phase 3C write-up.
- memory block not recorded at test time
- flags:
-ngl 999 --no-mmap --parallel 1 -fa 1 --jinja -c 131072 (reasoning OFF / enable_thinking:false; q8_0 KV; no MTP — this GGUF ships no heads) - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0002 — unsloth/Qwen3.5-122B-A10B-MTP-GGUF @ UD-Q4_K_M
- host aibeast · engine igpu · runtime ggml-org/llama.cpp@b6142 (rocm)
- derived runtime.tree: fork — carrying con-0001
- derived ctx_total: 200,000(200,000 × 1 slot(s))
- derived total_gb: 80.4 · headroom_gb: 15.6 · pool 96
- derived pool_utilisation (static estimate): 83.8%
- derived risk_utilisation: 83.8% — basis estimate (no observed peak collected; see clm-0001 — the estimate has never predicted an OOM) ⚠ OOM RISK
- acknowledged (ack-0001, review 2026-09-15): Runs at 83.8% of pool and has been the daily driver since 2026-06-28. Accepted deliberately: the 122B does not fit under 80% at any usable context, and the alternative is a materially weaker model. Mitigations after the 2026-07-21 OOM panic: --ctx-checkpoints capped at 4, desktop stack removed, swap raised to 16G. Revisit when aihydra arrives and the load can move.
- flags:
-ngl 999 --no-mmap --parallel 1 -fa 1 --jinja --spec-type draft-mtp --spec-draft-n-max 6 --ctx-checkpoints 4 - chat template: 7b21e9c4a5f60d18
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0003 — openai/gpt-oss-120b-GGUF @ Q4_K_M
- host aibeast · engine igpu · runtime ggml-org/llama.cpp@unknown-2026-06 (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; llama.cpp commit, chat-template hash and the per-model memory block were not recorded at test time. Quant, backend, context, KV quant, serving flags and the reasoning setting are from the Phase 3C write-up.
- memory block not recorded at test time
- flags:
-ngl 999 --no-mmap --parallel 1 -fa 1 --jinja -c 131072 (reasoning OFF / enable_thinking:false; q8_0 KV) - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0004 — nvidia/Nemotron-3-Super-GGUF @ IQ4_XS
- host aibeast · engine igpu · runtime ggml-org/llama.cpp@unknown-2026-06 (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; llama.cpp commit, chat-template hash and the per-model memory block were not recorded at test time. Quant, backend, context, KV quant, serving flags and the reasoning setting are from the Phase 3C write-up.
- memory block not recorded at test time
- flags:
-ngl 999 --no-mmap --parallel 1 -fa 1 --jinja -c 131072 (reasoning OFF / enable_thinking:false; q8_0 KV) - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0005 — unsloth/Qwen3.6-27B-GGUF @ Q4_K_M
- host aibeast · engine igpu · runtime ggml-org/llama.cpp@unknown-2026-06 (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; llama.cpp commit, chat-template hash and the per-model memory block were not recorded at test time. Quant, backend, context, KV quant, serving flags and the reasoning setting are from the Phase 3C write-up.
- memory block not recorded at test time
- flags:
-ngl 999 --no-mmap --parallel 1 -fa 1 --jinja -c 131072 (reasoning OFF / enable_thinking:false; q8_0 KV) - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0006 — unsloth/Qwen3.5-122B-A10B-MTP-GGUF @ UD-Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d6d547ec763317d9ecd0ace334a7e21359 (rocm)
- derived runtime.tree: upstream
- derived ctx_total: 16,384(16,384 × 1 slot(s))
- derived total_gb: 76.98 · headroom_gb: 28.02 · pool 105
- derived pool_utilisation (static estimate): 73.3%
- derived risk_utilisation: 73.3% — basis estimate (no observed peak collected; see clm-0001 — the estimate has never predicted an OOM)
- flags:
-ngl 999 -fa on -c 16384 --parallel 1 --load-mode none --jinja -ctk f16 -ctv f16 - chat template: not recorded
- capability: measured
cfg-0007 — unsloth/Qwen3.5-122B-A10B-MTP-GGUF @ UD-Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d6d547ec763317d9ecd0ace334a7e21359 (rocm)
- derived runtime.tree: upstream
- derived ctx_total: 16,384(16,384 × 1 slot(s))
- derived total_gb: 76.98 · headroom_gb: 28.02 · pool 105
- derived pool_utilisation (static estimate): 73.3%
- derived risk_utilisation: 73.3% — basis estimate (no observed peak collected; see clm-0001 — the estimate has never predicted an OOM)
- flags:
-ngl 999 -fa on -c 16384 --parallel 1 --load-mode none --jinja -ctk f16 -ctv f16 --spec-type draft-mtp --spec-draft-n-max 3 - chat template: not recorded
- capability: inherited from cfg-0006 across neutral lever(s): speculation
cfg-0008 — Qwen3.5-122B-A10B-UD-Q4_K_M-00001-of-00003.gguf @ UD-Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this tool. Everything present here — build commit, backend, KV quant, batch sizes, flash attention — is machine-reported.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 0 -b 2048 -ub 512 -ctk f16 -ctv f16 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0009 — Qwen3.5-122B-A10B-UD-Q4_K_M-00001-of-00003.gguf @ UD-Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this tool. Everything present here — build commit, backend, KV quant, batch sizes, flash attention — is machine-reported.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk f16 -ctv f16 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0010 — Qwen3.5-122B-A10B-UD-Q4_K_M-00001-of-00003.gguf @ UD-Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this tool. Everything present here — build commit, backend, KV quant, batch sizes, flash attention — is machine-reported.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk q8_0 -ctv q8_0 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0011 — Qwen3.5-122B-A10B-UD-Q4_K_M-00001-of-00003.gguf @ UD-Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@ce7689f (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this tool. Everything present here — build commit, backend, KV quant, batch sizes, flash attention — is machine-reported.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk q8_0 -ctv q8_0 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0012 — Qwen3.6-35B-A3B-UD-Q4_K_XL.gguf @ UD-Q4_K_XL
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this tool. Everything present here — build commit, backend, KV quant, batch sizes, flash attention — is machine-reported.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk f16 -ctv f16 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0013 — Qwen3.5-122B-A10B-UD-Q4_K_M-00001-of-00003.gguf @ UD-Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@ce7689f (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this tool. Everything present here — build commit, backend, KV quant, batch sizes, flash attention — is machine-reported.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk q8_0 -ctv q8_0 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0014 — Qwen3.5-122B-A10B-UD-Q4_K_M-00001-of-00003.gguf @ UD-Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@ce7689f (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this tool. Everything present here — build commit, backend, KV quant, batch sizes, flash attention — is machine-reported.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk q8_0 -ctv q8_0 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0015 — NVIDIA-Nemotron-3-Super-120B-A12B-UD-IQ4_XS-00001-of-00003.gguf @ UD-IQ4_XS
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this tool. Everything present here — build commit, backend, KV quant, batch sizes, flash attention — is machine-reported.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk f16 -ctv f16 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0016 — gpt-oss-120b-UD-Q4_K_XL-00001-of-00002.gguf @ UD-Q4_K_XL
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this tool. Everything present here — build commit, backend, KV quant, batch sizes, flash attention — is machine-reported.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk f16 -ctv f16 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0017 — gpt-oss-120b-UD-Q4_K_XL-00001-of-00002.gguf @ UD-Q4_K_XL
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this tool. Everything present here — build commit, backend, KV quant, batch sizes, flash attention — is machine-reported.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk f16 -ctv f16 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0018 — gpt-oss-120b-UD-Q4_K_XL-00001-of-00002.gguf @ UD-Q4_K_XL
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this tool. Everything present here — build commit, backend, KV quant, batch sizes, flash attention — is machine-reported.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk q8_0 -ctv q8_0 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0019 — NVIDIA-Nemotron-3-Super-120B-A12B-UD-IQ4_XS-00001-of-00003.gguf @ UD-IQ4_XS
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this tool. Everything present here — build commit, backend, KV quant, batch sizes, flash attention — is machine-reported.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 0 -b 2048 -ub 512 -ctk f16 -ctv f16 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0020 — NVIDIA-Nemotron-3-Super-120B-A12B-UD-IQ4_XS-00001-of-00003.gguf @ UD-IQ4_XS
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this tool. Everything present here — build commit, backend, KV quant, batch sizes, flash attention — is machine-reported.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk f16 -ctv f16 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0021 — NVIDIA-Nemotron-3-Super-120B-A12B-UD-IQ4_XS-00001-of-00003.gguf @ UD-IQ4_XS
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this tool. Everything present here — build commit, backend, KV quant, batch sizes, flash attention — is machine-reported.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk q8_0 -ctv q8_0 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0022 — Qwen3.6-35B-A3B-UD-Q4_K_XL.gguf @ UD-Q4_K_XL
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this tool. Everything present here — build commit, backend, KV quant, batch sizes, flash attention — is machine-reported.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 0 -b 2048 -ub 512 -ctk f16 -ctv f16 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0023 — Qwen3.6-35B-A3B-UD-Q4_K_XL.gguf @ UD-Q4_K_XL
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this tool. Everything present here — build commit, backend, KV quant, batch sizes, flash attention — is machine-reported.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk f16 -ctv f16 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0024 — Qwen3.6-35B-A3B-UD-Q4_K_XL.gguf @ UD-Q4_K_XL
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; Ingested from llama-bench, which measures throughput and not footprint: the memory block, chat-template hash and model revision are absent by construction. Peak memory comes from a Tier 0 capability probe, not from this tool. Everything present here — build commit, backend, KV quant, batch sizes, flash attention — is machine-reported.
- memory block not recorded at test time
- flags:
-ngl 999 -fa 1 -b 2048 -ub 512 -ctk q8_0 -ctv q8_0 -t 16 - chat template: not recorded
- capability: unmeasured — performance numbers for this config stand on a guard alone, not an established capability
cfg-0025 — Qwen3.5-122B-A10B-UD-Q4_K_M-00001-of-00003.gguf @ UD-Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d6d547ec763317d9ecd0ace334a7e21359 (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; tau2 serving config for the Aug 10-11 protocol arms (0.545 headline, anchor pair, sim-sensitivity, task9-regression). Flags come from the runner scripts — powered.sh on-box for the Aug 9-10 arms, bench/queues/queue-2026-08-11.sh / queue-f / queue-g in this repo — with the build asserted fatal at launch, not trusted. Memory block, HF revision and chat-template hash were not recorded at test time; GGUF sha256 is ledgered in audits/environment.lock-aihydra-2026-08-12.json.
- memory block not recorded at test time
- flags:
-ngl 999 -fa on -c 32768 --parallel 1 --load-mode none --jinja -rea off -ctk f16 -ctv f16 - chat template: not recorded
- capability: measured
cfg-0026 — Qwen3.5-122B-A10B-UD-Q4_K_M-00001-of-00003.gguf @ UD-Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d6d547ec763317d9ecd0ace334a7e21359 (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; cfg-0025 with quantised KV — the arms differ ONLY in -ctk/-ctv, which is the whole design of the clm-0038 paired comparison. Same provenance and gaps as cfg-0025.
- memory block not recorded at test time
- flags:
-ngl 999 -fa on -c 32768 --parallel 1 --load-mode none --jinja -rea off -ctk q8_0 -ctv q8_0 - chat template: not recorded
- capability: measured
cfg-0027 — NVIDIA-Nemotron-3-Super-120B-A12B-UD-IQ4_XS-00001-of-00003.gguf @ UD-IQ4_XS
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d6d547ec763317d9ecd0ace334a7e21359 (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; tau2 serving config for the Nemotron IQ4_XS arms (Aug 10 comparison arm, Stage E2). Flags from powered.sh (on-box) and bench/queues/queue-e-nemotron-quant.sh; build asserted at launch. Memory block, HF revision and chat-template hash not recorded at test time; GGUF sha256 ledgered in audits/environment.lock-aihydra-2026-08-12.json.
- memory block not recorded at test time
- flags:
-ngl 999 -fa on -c 32768 --parallel 1 --load-mode none --jinja -rea off -ctk f16 -ctv f16 - chat template: not recorded
- capability: measured
cfg-0028 — NVIDIA-Nemotron-3-Super-120B-A12B-UD-Q4_K_M-00001-of-00003.gguf @ UD-Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@3653e6d6d547ec763317d9ecd0ace334a7e21359 (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; cfg-0027 with the model at UD-Q4_K_M — Stage E1 of the quant-pair deconfound (clm-0047): the two arms differ only in the model file's quantisation. Flags from bench/queues/queue-e-nemotron-quant.sh; build asserted at launch. Same gaps as cfg-0027.
- memory block not recorded at test time
- flags:
-ngl 999 -fa on -c 32768 --parallel 1 --load-mode none --jinja -rea off -ctk f16 -ctv f16 - chat template: not recorded
- capability: measured
cfg-0029 — Qwen3.5-122B-A10B-UD-Q4_K_M-00001-of-00003.gguf @ UD-Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@ce7689f (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; cfg-0025 on the KV-dequant-patched build (clm-0022's lever, clm-0045's arms). ce7689f is a LOCAL commit object — the cherry-pick of Nathanw1014's 2a24abc onto 3653e6d, tree ~/src/llama.cpp-kvfix on aihydra; recipe recorded in the audit, exact SHA not re-derivable bit-for-bit off-box (same caveat as cfg-0011/13/14). Build self-report asserted fatal at launch by the runner. Same recording gaps as cfg-0025.
- memory block not recorded at test time
- flags:
-ngl 999 -fa on -c 32768 --parallel 1 --load-mode none --jinja -rea off -ctk f16 -ctv f16 - chat template: not recorded
- capability: measured
cfg-0030 — Qwen3.5-122B-A10B-UD-Q4_K_M-00001-of-00003.gguf @ UD-Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@ce7689f (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; cfg-0029 with quantised KV — the patched half of the clm-0045 paired design (arms differ only in -ctk/-ctv on the SAME patched build). Same provenance and gaps as cfg-0029.
- memory block not recorded at test time
- flags:
-ngl 999 -fa on -c 32768 --parallel 1 --load-mode none --jinja -rea off -ctk q8_0 -ctv q8_0 - chat template: not recorded
- capability: measured
cfg-0031 — Qwen3.5-122B-A10B-UD-Q4_K_M-00001-of-00003.gguf @ UD-Q4_K_M
- host aihydra · engine igpu · runtime ggml-org/llama.cpp@62bf73d (rocm)
- derived runtime.tree: upstream
- backfilled reconstructed from the archive; cfg-0025 on the fleet-candidate minimal build min-62bf73d — the NEW side of the Stage A anchor pair (clm-0046's build-delta calibration: +2.8% pp / +0.9% tg with identical tau2 capability). Build tree ~/src/llama.cpp-min62bf73d on aihydra, asserted at launch by bench/queues/queue-2026-08-11.sh. Same recording gaps as cfg-0025.
- memory block not recorded at test time
- flags:
-ngl 999 -fa on -c 32768 --parallel 1 --load-mode none --jinja -rea off -ctk f16 -ctv f16 - chat template: not recorded
- capability: measured
Runs (111)
Split by what is being established. Capability is what the model can do; performance is how fast it delivers that; a guard is the cheap cliff-detector that makes a performance number worth recording. Seeprotocol §1.
capability (15)
run-0003 — conceptual-c1-c8-blind@v1
- 2026-06-28 · config → cfg-0001
- depth: 0 tokens (conversational — says nothing about behaviour at 100k+) · scoring: blind-human
- time_to_correct_answer_s=1.7
run-0098 — tau2-bench-airline@v1.0.1
run-0099 — tau2-bench-airline-nothink@v1.0.1
run-0100 — tau2-bench-airline@668d3bc
- 2026-08-09 · config → cfg-0025
- depth: 32,768 tokens · scoring: deterministic
- tasks_passed=12 · tasks_total=22 · mean_score=0.5455
run-0101 — tau2-bench-airline@668d3bc
- 2026-08-10 · config → cfg-0026
- depth: 32,768 tokens · scoring: deterministic
- tasks_passed=8 · tasks_total=9 · mean_score=0.8889
run-0102 — tau2-bench-airline@668d3bc
- 2026-08-10 · config → cfg-0027
- depth: 32,768 tokens · scoring: deterministic
- tasks_passed=10 · tasks_total=16 · mean_score=0.625
run-0103 — tau2-bench-airline@668d3bc
run-0104 — tau2-bench-airline@668d3bc
run-0105 — tau2-bench-airline@668d3bc
run-0106 — tau2-bench-airline@668d3bc
run-0107 — tau2-bench-airline@668d3bc
run-0108 — tau2-bench-airline@668d3bc
run-0109 — tau2-bench-airline@668d3bc
run-0110 — tau2-bench-airline@668d3bc
run-0111 — tau2-bench-airline@668d3bc
- 2026-08-11 · config → cfg-0025
- depth: 32,768 tokens · scoring: deterministic
- tasks_passed=1 · tasks_total=8 · mean_score=0.125
guard (2)
run-0004 — tool-loop@v2.1
- 2026-06-27 · config → cfg-0001
- tool_call_success=0.94 · turns_before_degradation=10+ · empty_arg_calls=0
run-0005 — capability-guard@v2
- 2026-08-08 · config → cfg-0006
- tool_call_success=1 · empty_arg_calls=0 · tasks_passed=4 · tasks_total=4
performance (94)
run-0001 — tool-loop@v2.1
run-0002 — mtp-decode-sweep@v1
- 2026-07-20 · config → cfg-0002
- decode_tps=36.5 · prefill_tps=144 · ttft_ms=0 · empty_arg_calls=0
- ⚠ no guard — No guard was run. MTP is a neutral lever so capability loss is not expected by construction, but that is an argument, not a measurement.
run-0007 — spec-ab@v1
run-0008 — llama-bench@1
- 2026-08-08 · config → cfg-0008
- prefill_tps=312.037487
- ⚠ no guard — Guarded at run-0005 for this model+build+backend at --parallel 1. llama-bench cells vary KV-quant (lossy) and flash-attn (binary), across which capability cannot be inherited, and llama-bench serves no endpoint to guard directly. Recorded as a known gap rather than an implied pass.
- energy window: eng-0002
run-0009 — llama-bench@1
- 2026-08-08 · config → cfg-0008
- decode_tps=21.676081
- ⚠ no guard — Guarded at run-0005 for this model+build+backend at --parallel 1. llama-bench cells vary KV-quant (lossy) and flash-attn (binary), across which capability cannot be inherited, and llama-bench serves no endpoint to guard directly. Recorded as a known gap rather than an implied pass.
- energy window: eng-0002
run-0010 — llama-bench@1
- 2026-08-08 · config → cfg-0008
- prefill_tps=283.473901
- ⚠ no guard — Guarded at run-0005 for this model+build+backend at --parallel 1. llama-bench cells vary KV-quant (lossy) and flash-attn (binary), across which capability cannot be inherited, and llama-bench serves no endpoint to guard directly. Recorded as a known gap rather than an implied pass.
- energy window: eng-0002
run-0011 — llama-bench@1
- 2026-08-08 · config → cfg-0008
- decode_tps=19.258807
- ⚠ no guard — Guarded at run-0005 for this model+build+backend at --parallel 1. llama-bench cells vary KV-quant (lossy) and flash-attn (binary), across which capability cannot be inherited, and llama-bench serves no endpoint to guard directly. Recorded as a known gap rather than an implied pass.
- energy window: eng-0002
run-0012 — llama-bench@1
- 2026-08-08 · config → cfg-0008
- prefill_tps=170.811067
- ⚠ no guard — Guarded at run-0005 for this model+build+backend at --parallel 1. llama-bench cells vary KV-quant (lossy) and flash-attn (binary), across which capability cannot be inherited, and llama-bench serves no endpoint to guard directly. Recorded as a known gap rather than an implied pass.
- energy window: eng-0002
run-0013 — llama-bench@1
- 2026-08-08 · config → cfg-0008
- decode_tps=9.867545
- ⚠ no guard — Guarded at run-0005 for this model+build+backend at --parallel 1. llama-bench cells vary KV-quant (lossy) and flash-attn (binary), across which capability cannot be inherited, and llama-bench serves no endpoint to guard directly. Recorded as a known gap rather than an implied pass.
- energy window: eng-0002
run-0014 — llama-bench@1
- 2026-08-08 · config → cfg-0009
- prefill_tps=321.722166
- ⚠ no guard — Guarded at run-0005 for this model+build+backend at --parallel 1. llama-bench cells vary KV-quant (lossy) and flash-attn (binary), across which capability cannot be inherited, and llama-bench serves no endpoint to guard directly. Recorded as a known gap rather than an implied pass.
- energy window: eng-0003
run-0015 — llama-bench@1
- 2026-08-08 · config → cfg-0009
- decode_tps=21.910456
- ⚠ no guard — Guarded at run-0005 for this model+build+backend at --parallel 1. llama-bench cells vary KV-quant (lossy) and flash-attn (binary), across which capability cannot be inherited, and llama-bench serves no endpoint to guard directly. Recorded as a known gap rather than an implied pass.
- energy window: eng-0003
run-0016 — llama-bench@1
- 2026-08-08 · config → cfg-0009
- prefill_tps=300.673419
- ⚠ no guard — Guarded at run-0005 for this model+build+backend at --parallel 1. llama-bench cells vary KV-quant (lossy) and flash-attn (binary), across which capability cannot be inherited, and llama-bench serves no endpoint to guard directly. Recorded as a known gap rather than an implied pass.
- energy window: eng-0003
run-0017 — llama-bench@1
- 2026-08-08 · config → cfg-0009
- decode_tps=21.393789
- ⚠ no guard — Guarded at run-0005 for this model+build+backend at --parallel 1. llama-bench cells vary KV-quant (lossy) and flash-attn (binary), across which capability cannot be inherited, and llama-bench serves no endpoint to guard directly. Recorded as a known gap rather than an implied pass.
- energy window: eng-0003
run-0018 — llama-bench@1
- 2026-08-08 · config → cfg-0009
- prefill_tps=208.771771
- ⚠ no guard — Guarded at run-0005 for this model+build+backend at --parallel 1. llama-bench cells vary KV-quant (lossy) and flash-attn (binary), across which capability cannot be inherited, and llama-bench serves no endpoint to guard directly. Recorded as a known gap rather than an implied pass.
- energy window: eng-0003
run-0019 — llama-bench@1
- 2026-08-08 · config → cfg-0009
- decode_tps=18.174796
- ⚠ no guard — Guarded at run-0005 for this model+build+backend at --parallel 1. llama-bench cells vary KV-quant (lossy) and flash-attn (binary), across which capability cannot be inherited, and llama-bench serves no endpoint to guard directly. Recorded as a known gap rather than an implied pass.
- energy window: eng-0003
run-0020 — llama-bench@1
- 2026-08-08 · config → cfg-0010
- prefill_tps=321.771492
- ⚠ no guard — Guarded at run-0005 for this model+build+backend at --parallel 1. llama-bench cells vary KV-quant (lossy) and flash-attn (binary), across which capability cannot be inherited, and llama-bench serves no endpoint to guard directly. Recorded as a known gap rather than an implied pass.
- energy window: eng-0004
run-0021 — llama-bench@1
- 2026-08-08 · config → cfg-0010
- decode_tps=21.63496
- ⚠ no guard — Guarded at run-0005 for this model+build+backend at --parallel 1. llama-bench cells vary KV-quant (lossy) and flash-attn (binary), across which capability cannot be inherited, and llama-bench serves no endpoint to guard directly. Recorded as a known gap rather than an implied pass.
- energy window: eng-0004
run-0022 — llama-bench@1
- 2026-08-08 · config → cfg-0010
- prefill_tps=302.605133
- ⚠ no guard — Guarded at run-0005 for this model+build+backend at --parallel 1. llama-bench cells vary KV-quant (lossy) and flash-attn (binary), across which capability cannot be inherited, and llama-bench serves no endpoint to guard directly. Recorded as a known gap rather than an implied pass.
- energy window: eng-0004
run-0023 — llama-bench@1
- 2026-08-08 · config → cfg-0010
- decode_tps=20.609235
- ⚠ no guard — Guarded at run-0005 for this model+build+backend at --parallel 1. llama-bench cells vary KV-quant (lossy) and flash-attn (binary), across which capability cannot be inherited, and llama-bench serves no endpoint to guard directly. Recorded as a known gap rather than an implied pass.
- energy window: eng-0004
run-0024 — llama-bench@1
- 2026-08-08 · config → cfg-0010
- prefill_tps=204.485278
- ⚠ no guard — Guarded at run-0005 for this model+build+backend at --parallel 1. llama-bench cells vary KV-quant (lossy) and flash-attn (binary), across which capability cannot be inherited, and llama-bench serves no endpoint to guard directly. Recorded as a known gap rather than an implied pass.
- energy window: eng-0004
run-0025 — llama-bench@1
- 2026-08-08 · config → cfg-0010
- decode_tps=15.256436
- ⚠ no guard — Guarded at run-0005 for this model+build+backend at --parallel 1. llama-bench cells vary KV-quant (lossy) and flash-attn (binary), across which capability cannot be inherited, and llama-bench serves no endpoint to guard directly. Recorded as a known gap rather than an implied pass.
- energy window: eng-0004
run-0026 — llama-bench@2
- 2026-08-08 · config → cfg-0011
- prefill_tps=322.369017
- ⚠ no guard — Patched build (3653e6d + ce7689f KV-dequant). Speculation/KV levers only; guard run-0005 covers the same model+backend at --parallel 1, but KV quant is a lossy lever so capability does not inherit. Recorded as a gap.
- energy window: eng-0010
run-0027 — llama-bench@2
- 2026-08-08 · config → cfg-0011
- decode_tps=21.83139
- ⚠ no guard — Patched build (3653e6d + ce7689f KV-dequant). Speculation/KV levers only; guard run-0005 covers the same model+backend at --parallel 1, but KV quant is a lossy lever so capability does not inherit. Recorded as a gap.
- energy window: eng-0010
run-0028 — llama-bench@2
- 2026-08-08 · config → cfg-0011
- prefill_tps=211.351215
- ⚠ no guard — Patched build (3653e6d + ce7689f KV-dequant). Speculation/KV levers only; guard run-0005 covers the same model+backend at --parallel 1, but KV quant is a lossy lever so capability does not inherit. Recorded as a gap.
- energy window: eng-0010
run-0029 — llama-bench@2
- 2026-08-08 · config → cfg-0011
- decode_tps=18.087085
- ⚠ no guard — Patched build (3653e6d + ce7689f KV-dequant). Speculation/KV levers only; guard run-0005 covers the same model+backend at --parallel 1, but KV quant is a lossy lever so capability does not inherit. Recorded as a gap.
- energy window: eng-0010
run-0030 — llama-bench@1
run-0031 — llama-bench@1
run-0032 — llama-bench@1
run-0033 — llama-bench@1
run-0034 — llama-bench@1
run-0035 — llama-bench@1
run-0036 — llama-bench@2
- 2026-08-08 · config → cfg-0013
- decode_tps=12.152979
- ⚠ no guard — Patched build at 131072 depth, decode-only (-p 0 -r 1). Guard run-0005 covers this model+backend at --parallel 1; KV quant is a lossy lever so capability does not inherit. Depth exceeds the guard's 8000-token probe — recorded as a gap.
- energy window: eng-0015
run-0037 — llama-bench@2
- 2026-08-08 · config → cfg-0014
- decode_tps=9.6866
- ⚠ no guard — Patched build at 204800 depth (production context), decode-only. Guard run-0005 covers this model+backend at --parallel 1; KV quant is a lossy lever so capability does not inherit, and depth far exceeds the guard's 8000-token probe. Recorded as a gap.
- energy window: eng-0018
run-0038 — llama-bench@1
run-0039 — llama-bench@1
run-0040 — llama-bench@1
run-0041 — llama-bench@1
run-0042 — llama-bench@1
run-0043 — llama-bench@1
run-0044 — llama-bench@1
- 2026-08-08 · config → cfg-0016
- prefill_tps=450.832446
- ⚠ no guard — GUARD FAILED on this model: retrieval at 8000 tokens returned a stale answer from a previous request. Coherence and tool-calling passed. Numbers recorded because throughput is unaffected by the anomaly, but they are NOT cleared by a guard — see clm-0025.
- energy window: eng-0022
run-0045 — llama-bench@1
- 2026-08-08 · config → cfg-0016
- decode_tps=55.454588
- ⚠ no guard — GUARD FAILED on this model: retrieval at 8000 tokens returned a stale answer from a previous request. Coherence and tool-calling passed. Numbers recorded because throughput is unaffected by the anomaly, but they are NOT cleared by a guard — see clm-0025.
- energy window: eng-0022
run-0046 — llama-bench@1
- 2026-08-08 · config → cfg-0016
- prefill_tps=421.546647
- ⚠ no guard — GUARD FAILED on this model: retrieval at 8000 tokens returned a stale answer from a previous request. Coherence and tool-calling passed. Numbers recorded because throughput is unaffected by the anomaly, but they are NOT cleared by a guard — see clm-0025.
- energy window: eng-0022
run-0047 — llama-bench@1
- 2026-08-08 · config → cfg-0016
- decode_tps=52.793591
- ⚠ no guard — GUARD FAILED on this model: retrieval at 8000 tokens returned a stale answer from a previous request. Coherence and tool-calling passed. Numbers recorded because throughput is unaffected by the anomaly, but they are NOT cleared by a guard — see clm-0025.
- energy window: eng-0022
run-0048 — llama-bench@1
- 2026-08-08 · config → cfg-0016
- prefill_tps=308.546401
- ⚠ no guard — GUARD FAILED on this model: retrieval at 8000 tokens returned a stale answer from a previous request. Coherence and tool-calling passed. Numbers recorded because throughput is unaffected by the anomaly, but they are NOT cleared by a guard — see clm-0025.
- energy window: eng-0022
run-0049 — llama-bench@1
- 2026-08-08 · config → cfg-0016
- decode_tps=41.698106
- ⚠ no guard — GUARD FAILED on this model: retrieval at 8000 tokens returned a stale answer from a previous request. Coherence and tool-calling passed. Numbers recorded because throughput is unaffected by the anomaly, but they are NOT cleared by a guard — see clm-0025.
- energy window: eng-0022
run-0050 — llama-bench@1
run-0051 — llama-bench@1
run-0052 — llama-bench@1
run-0053 — llama-bench@1
run-0054 — llama-bench@1
run-0055 — llama-bench@1
run-0056 — llama-bench@1
run-0057 — llama-bench@1
run-0058 — llama-bench@1
run-0059 — llama-bench@1
run-0060 — llama-bench@1
run-0061 — llama-bench@1
run-0062 — llama-bench@1
run-0063 — llama-bench@1
run-0064 — llama-bench@1
run-0065 — llama-bench@1
run-0066 — llama-bench@1
run-0067 — llama-bench@1
run-0068 — llama-bench@1
run-0069 — llama-bench@1
run-0070 — llama-bench@1
run-0071 — llama-bench@1
run-0072 — llama-bench@1
run-0073 — llama-bench@1
run-0074 — llama-bench@1
run-0075 — llama-bench@1
run-0076 — llama-bench@1
run-0077 — llama-bench@1
run-0078 — llama-bench@1
run-0079 — llama-bench@1
run-0080 — llama-bench@1
run-0081 — llama-bench@1
run-0082 — llama-bench@1
run-0083 — llama-bench@1
run-0084 — llama-bench@1
run-0085 — llama-bench@1
run-0086 — llama-bench@1
run-0087 — llama-bench@1
run-0088 — llama-bench@1
run-0089 — llama-bench@1
run-0090 — llama-bench@1
run-0091 — llama-bench@1
run-0092 — llama-bench@1
run-0093 — llama-bench@1
run-0094 — llama-bench@1
run-0095 — llama-bench@1
run-0096 — llama-bench@1
run-0097 — llama-bench@1
Energy (64)
Wh-denominated, wall-metered, read as a cumulative-counter difference over the window. delta_w over the idle baseline is the only figure that means anything.
idle-aihydra-2026-08 — idle baseline, aihydra
- period 2026-08-07 → 2026-08-10 · method wall-meter
- box-idle-all-empty: 10.1 W · 14 h
- igpu-active: 168 W · 53 h
- igpu-resident-quiet: 13.6 W · 0.17 h
- npu-resident-quiet: unmeasured — an honest gap beats an invented number
- standing: 7.4 kWh/mo · £2.24/mo · duty cycle 0.79
- First idle_baseline record for aihydra, covering its first days of service. The box-idle floor of 10.1 W is remarkably low for a 128 GB machine and is the figure every delta_w in the energy records subtracts from. igpu-resident-quiet MEASURED 2026-08-11 (eng-0035: 10 minutes, 73 GB model loaded, zero requests): **13.6 W mean, peak 35.1 W** — just +3.5 W over the empty-box floor. The keep_alive question is answered: holding the 122B resident costs ~2.5 kWh/month, ~76p at 30.3 p/kWh. This also anchors energy attribution for cloud-simulator arms: with resident-quiet this low, a high mean during such arms means genuine computation, not wait-state draw. `npu-resident-quiet` remains unmeasured — still the question an always-on NPU triage loop turns on. Duty cycle is 0.79 — this box has been working, not idling, across the period. That is unrepresentative of steady state and will fall once the benchmark programme ends; the standing-cost figures should be re-derived then rather than quoted from this window. Standing cost assumes 30.3 p/kWh grid import. With solar and battery the realised cost is lower and sometimes zero, which is what the energy records' tariff_window field exists to capture.
eng-0001
- 2026-08-08 · wall-meter · TP-Link smart plug via Home Assistant · window reconstructed
- wh_total 4.977 · mean 59.7 W · baseline 10.1 W (box-idle-all-empty) · delta 49.6 W
- lead-in assumed 5min (no prior invocation within 1h). invocation JSON is 0 bytes (empty output) — the run failed to record results, so no run can be linked.
eng-0002
- 2026-08-08 · wall-meter · TP-Link smart plug via Home Assistant · window reconstructed
- wh_total 26.164 · mean 88.1 W · baseline 10.1 W (box-idle-all-empty) · delta 78 W
- covers: run-0008 run-0009 run-0010 run-0011 run-0012 run-0013
- window = previous invocation end. matched to cfg-0008 (commit 3653e6d, fa=0, ctk/ctv=f16) by exact avg_ts value match against all 6 rows.
eng-0003
- 2026-08-08 · wall-meter · TP-Link smart plug via Home Assistant · window reconstructed
- wh_total 12.32 · mean 168.5 W · baseline 10.1 W (box-idle-all-empty) · delta 158.4 W
- covers: run-0014 run-0015 run-0016 run-0017 run-0018 run-0019
- window = previous invocation end. matched to cfg-0009 (commit 3653e6d, fa=1, ctk/ctv=f16) by exact avg_ts value match against all 6 rows.
eng-0004
- 2026-08-08 · wall-meter · TP-Link smart plug via Home Assistant · window reconstructed
- wh_total 13.087 · mean 129.3 W · baseline 10.1 W (box-idle-all-empty) · delta 119.2 W
- covers: run-0020 run-0021 run-0022 run-0023 run-0024 run-0025
- window = previous invocation end. matched to cfg-0010 (commit 3653e6d, fa=1, ctk/ctv=q8_0) by exact avg_ts value match against all 6 rows.
eng-0005
- 2026-08-08 · wall-meter · TP-Link smart plug via Home Assistant · window reconstructed
- wh_total 29.199 · mean 175.5 W · baseline 10.1 W (box-idle-all-empty) · delta 165.4 W
- window = previous invocation end. model_filename is Qwen3.6-27B-Q4_K_M.gguf — no content/configs record exists for this model at all; appears to be an exploratory flash-attn cliff-detector sweep (depths 0/32768/65536) that was never promoted to a run.
eng-0006
- 2026-08-08 · wall-meter · TP-Link smart plug via Home Assistant · window reconstructed
- wh_total 38.37 · mean 173.4 W · baseline 10.1 W (box-idle-all-empty) · delta 163.3 W
- window = previous invocation end. model_filename is Qwen3.6-27B-Q4_K_M.gguf — no content/configs record exists for this model at all; appears to be an exploratory flash-attn cliff-detector sweep (depths 0/32768/65536) that was never promoted to a run.
eng-0007
- 2026-08-08 · wall-meter · TP-Link smart plug via Home Assistant · window reconstructed
- wh_total 15.225 · mean 139.6 W · baseline 10.1 W (box-idle-all-empty) · delta 129.5 W
- window = previous invocation end. signature (commit 3653e6d, fa=1, ctk/ctv=f16, model qwen35-122b) matches cfg-0009's config, but avg_ts values (0-depth prefill 326.36 vs run-0014's 321.72) don't exactly match any promoted run — looks like a preliminary A/B rerun superseded by the sweep2 sweep, not itself promoted.
eng-0008
- 2026-08-08 · wall-meter · TP-Link smart plug via Home Assistant · window reconstructed
- wh_total 10.678 · mean 160.6 W · baseline 10.1 W (box-idle-all-empty) · delta 150.5 W
- window = previous invocation end. signature (commit 3653e6d, fa=1, ctk/ctv=q8_0, model qwen35-122b) matches cfg-0010's config, but avg_ts values don't exactly match any promoted run — preliminary A/B rerun, not itself promoted.
eng-0009
- 2026-08-08 · wall-meter · TP-Link smart plug via Home Assistant · window reconstructed
- wh_total 10.22 · mean 158.8 W · baseline 10.1 W (box-idle-all-empty) · delta 148.7 W
- window = previous invocation end. signature is commit ce7689f, fa=1, ctk/ctv=f16 — no content/configs record combines that commit with f16 KV (cfg-0011/0013/0014 on ce7689f are all q8_0), so no run to link.
eng-0010
- 2026-08-08 · wall-meter · TP-Link smart plug via Home Assistant · window reconstructed
- wh_total 10.167 · mean 151.2 W · baseline 10.1 W (box-idle-all-empty) · delta 141.1 W
- covers: run-0026 run-0027 run-0028 run-0029
- window = previous invocation end. matched to cfg-0011 (commit ce7689f, fa=1, ctk/ctv=q8_0) by exact avg_ts value match against all 4 rows.
eng-0011
- 2026-08-08 · wall-meter · TP-Link smart plug via Home Assistant · window reconstructed
- wh_total 9.025 · mean 127 W · baseline 10.1 W (box-idle-all-empty) · delta 116.9 W
- covers: run-0030 run-0031 run-0032 run-0033 run-0034 run-0035
- window = previous invocation end. matched to cfg-0012 by exact avg_ts value match against all 6 rows.
eng-0012
- 2026-08-08 · wall-meter · TP-Link smart plug via Home Assistant · window reconstructed
- wh_total 45.193 · mean 169.4 W · baseline 10.1 W (box-idle-all-empty) · delta 159.3 W
- window = previous invocation end. commit 3653e6d + ctk/ctv f16 signature, single row (depth 131072) — no cfg record combines 3653e6d with a 131072-depth f16 test, and its avg_ts (12.031) matches no run.
eng-0013
- 2026-08-08 · wall-meter · TP-Link smart plug via Home Assistant · window reconstructed
- wh_total 45.466 · mean 168.8 W · baseline 10.1 W (box-idle-all-empty) · delta 158.7 W
- window = previous invocation end. commit 3653e6d + ctk/ctv q8_0, single row (depth 131072) — avg_ts (7.812) matches no run; appears superseded by the deepkv200-082221 rerun at the larger 204800 depth.
eng-0014
- 2026-08-08 · wall-meter · TP-Link smart plug via Home Assistant · window reconstructed
- wh_total 44.91 · mean 175.1 W · baseline 10.1 W (box-idle-all-empty) · delta 165 W
- window = previous invocation end. commit ce7689f + ctk/ctv f16, single row (depth 131072) — no cfg record combines that commit with f16 KV, and avg_ts (12.043) matches no run.
eng-0015
- 2026-08-08 · wall-meter · TP-Link smart plug via Home Assistant · window reconstructed
- wh_total 43.512 · mean 174.4 W · baseline 10.1 W (box-idle-all-empty) · delta 164.3 W
- covers: run-0036
- window = previous invocation end. matched to cfg-0013 (commit ce7689f, ctk/ctv q8_0, depth 131072) by exact avg_ts value match.
eng-0016
- 2026-08-08 · wall-meter · TP-Link smart plug via Home Assistant · window reconstructed
- wh_total 0.842 · mean 10.1 W · baseline 10.1 W (box-idle-all-empty) · delta 0 W
- lead-in assumed 5min (no prior invocation within 1h). invocation JSON is 0 bytes (empty output) — the run failed to record results, so no run can be linked.
eng-0017
- 2026-08-08 · wall-meter · TP-Link smart plug via Home Assistant · window reconstructed
- wh_total 146.082 · mean 171.7 W · baseline 10.1 W (box-idle-all-empty) · delta 161.6 W
- window = previous invocation end. commit 3653e6d + ctk/ctv q8_0, single row (depth 204800) — avg_ts (5.686) matches no run; no cfg record combines 3653e6d with a 204800-depth q8_0 test.
eng-0018
- 2026-08-08 · wall-meter · TP-Link smart plug via Home Assistant · window reconstructed
- wh_total 85.248 · mean 178.5 W · baseline 10.1 W (box-idle-all-empty) · delta 168.4 W
- covers: run-0037
- window = previous invocation end. matched to cfg-0014 (commit ce7689f, ctk/ctv q8_0, depth 204800) by exact avg_ts value match.
eng-0019
- 2026-08-08 · wall-meter · TP-Link smart plug via Home Assistant · window reconstructed
- wh_total 89.553 · mean 177.7 W · baseline 10.1 W (box-idle-all-empty) · delta 167.6 W
- window = previous invocation end. commit 3653e6d + ctk/ctv f16, single row (depth 204800) — avg_ts (9.599) matches no run.
eng-0020
- 2026-08-08 · wall-meter · TP-Link smart plug via Home Assistant · window reconstructed
- wh_total 90.427 · mean 177.7 W · baseline 10.1 W (box-idle-all-empty) · delta 167.6 W
- window = previous invocation end. commit ce7689f + ctk/ctv f16, single row (depth 204800) — no cfg record combines that commit with f16 KV, and avg_ts (9.539) matches no run.
eng-0021
- 2026-08-08 · wall-meter · TP-Link smart plug via Home Assistant · window reconstructed
- wh_total 19.658 · mean 46.9 W · baseline 10.1 W (box-idle-all-empty) · delta 36.8 W
- window = previous invocation end. cold.json/warm.json/diverged.json are empty and slots.json is a raw /slots endpoint dump ({n_ctx, speculative, is_processing}), not a llama-bench result — this is a disk-warmth/slot-restore diagnostic, not a benchmark run, so there is no run collection entry to link.
eng-0022
- 2026-08-08 · wall-meter · TP-Link smart plug via Home Assistant · window reconstructed
- wh_total 7.714 · mean 121 W · baseline 10.1 W (box-idle-all-empty) · delta 110.9 W
- covers: run-0044 run-0045 run-0046 run-0047 run-0048 run-0049
- window = previous invocation end. matched to cfg-0016 by exact avg_ts value match against all 6 rows.
eng-0023
- 2026-08-08 · wall-meter · TP-Link smart plug via Home Assistant · window reconstructed
- wh_total 14.858 · mean 151.1 W · baseline 10.1 W (box-idle-all-empty) · delta 141 W
- covers: run-0038 run-0039 run-0040 run-0041 run-0042 run-0043
- window = previous invocation end. matched to cfg-0015 by exact avg_ts value match against 5 rows; the depth-0 decode row (17.506 t/s) is shared verbatim with cfg-0020's depth-0 decode row, disambiguated by the other 5 rows in this file all keying to cfg-0015.
eng-0024
- 2026-08-08 · wall-meter · TP-Link smart plug via Home Assistant · window reconstructed
- wh_total 16.943 · mean 37.7 W · baseline 10.1 W (box-idle-all-empty) · delta 27.6 W
- window = previous invocation end. cold.json/warm.json/diverged.json are empty and slots.json is a raw /slots endpoint dump, not a llama-bench result — disk-warmth/slot-restore diagnostic, no run collection entry to link.
eng-0025
- 2026-08-08 · wall-meter · TP-Link smart plug via Home Assistant · window reconstructed
- wh_total 44.231 · mean 103.8 W · baseline 10.1 W (box-idle-all-empty) · delta 93.7 W
- covers: run-0050 run-0051 run-0052 run-0053 run-0054 run-0055
- window = previous invocation end. matched to cfg-0017 by exact avg_ts value match against all 6 rows.
eng-0026
- 2026-08-08 · wall-meter · TP-Link smart plug via Home Assistant · window reconstructed
- wh_total 142.65 · mean 173.3 W · baseline 10.1 W (box-idle-all-empty) · delta 163.2 W
- covers: run-0056 run-0057 run-0058 run-0059 run-0060 run-0061
- window = previous invocation end. matched to cfg-0018 by exact avg_ts value match against all 6 rows.
eng-0027
- 2026-08-08 · wall-meter · TP-Link smart plug via Home Assistant · window reconstructed
- wh_total 22.565 · mean 159.7 W · baseline 10.1 W (box-idle-all-empty) · delta 149.6 W
- covers: run-0086 run-0087 run-0088 run-0089 run-0090 run-0091
- window = previous invocation end. matched to cfg-0023 by exact avg_ts value match against all 6 rows.
eng-0028
- 2026-08-08 · wall-meter · TP-Link smart plug via Home Assistant · window reconstructed
- wh_total 28.854 · mean 179.2 W · baseline 10.1 W (box-idle-all-empty) · delta 169.1 W
- covers: run-0080 run-0081 run-0082 run-0083 run-0084 run-0085
- window = previous invocation end. matched to cfg-0022 by exact avg_ts value match against all 6 rows.
eng-0029
- 2026-08-08 · wall-meter · TP-Link smart plug via Home Assistant · window reconstructed
- wh_total 20.58 · mean 176.5 W · baseline 10.1 W (box-idle-all-empty) · delta 166.4 W
- covers: run-0092 run-0093 run-0094 run-0095 run-0096 run-0097
- window = previous invocation end. matched to cfg-0024 by exact avg_ts value match against all 6 rows.
eng-0030
- 2026-08-08 · wall-meter · TP-Link smart plug via Home Assistant · window reconstructed
- wh_total 48.097 · mean 172.2 W · baseline 10.1 W (box-idle-all-empty) · delta 162.1 W
- covers: run-0068 run-0069 run-0070 run-0071 run-0072 run-0073
- window = previous invocation end. matched to cfg-0020 by exact avg_ts value match against 5 rows; the depth-0 decode row (17.506 t/s) is shared verbatim with cfg-0015's depth-0 decode row, disambiguated by the other 5 rows in this file all keying to cfg-0020.
eng-0031
- 2026-08-08 · wall-meter · TP-Link smart plug via Home Assistant · window reconstructed
- wh_total 63.404 · mean 178 W · baseline 10.1 W (box-idle-all-empty) · delta 167.9 W
- covers: run-0062 run-0063 run-0064 run-0065 run-0066 run-0067
- window = previous invocation end. matched to cfg-0019 by exact avg_ts value match against all 6 rows.
eng-0032
- 2026-08-08 · wall-meter · TP-Link smart plug via Home Assistant · window reconstructed
- wh_total 48.764 · mean 178.3 W · baseline 10.1 W (box-idle-all-empty) · delta 168.2 W
- covers: run-0074 run-0075 run-0076 run-0077 run-0078 run-0079
- window = previous invocation end. matched to cfg-0021 by exact avg_ts value match against all 6 rows.
eng-0033
- 2026-08-09 · wall-meter · TP-Link smart plug via Home Assistant · window reconstructed
- wh_total 38.246 · mean 156.5 W · baseline 10.1 W (box-idle-all-empty) · delta 146.4 W
- per task: 12.749 Wh · 0.3863 p (tariff 30.3 p/kWh)
- covers: run-0098
- tau2 sim dir 20260809_144711 (airline, n=3, reasoning_content present -> think); picked as best- effort match for run-0098 (tasks_total=3, max_steps=40 matching the nothink group's config) out of 3 candidate reruns with identical reward pattern.
eng-0034
- 2026-08-09 · wall-meter · TP-Link smart plug via Home Assistant · window reconstructed
- wh_total 28.564 · mean 169.2 W · baseline 10.1 W (box-idle-all-empty) · delta 159.1 W
- per task: 5.713 Wh · 0.1731 p (tariff 30.3 p/kWh)
- covers: run-0099
- tau2 sim dir 20260809_161821 (airline, n=5, no reasoning_content -> nothink); picked as best- effort match for run-0099 (tasks_total=5, all rewards 1.0) as the earliest fully-passing nothink rerun following the chosen think run, out of 5 candidate fully-passing reruns.
eng-0035
- 2026-08-11 · wall-meter · AI Hydra energy/power sensors via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy; peak/mean drawn from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 2.275 · mean 13.6 W · baseline 10.1 W (box-idle-all-empty) · delta 3.5 W
- aihydra run-meta.jsonl entry: suite=idle-baseline, model=qwen35-122b, variant="resident- quiet". no run-NNNN file exists yet for this window; run left unlinked ([]).
eng-0036
- 2026-08-11 · wall-meter · AI Hydra energy/power sensors via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy; peak/mean drawn from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 0.494 · mean 29.6 W · baseline 10.1 W (box-idle-all-empty) · delta 19.5 W
- aihydra run-meta.jsonl entry: suite=tau2-kvpatch, model=qwen35-122b, variant="kvfix-f16". rc=2 (failed fast, 53s) — early-exit failure, window kept as it is a real recorded draw, not a zero-duration preflight guard. no run-NNNN file exists yet for this window; run left unlinked ([]).
eng-0037
- 2026-08-11 · wall-meter · AI Hydra energy/power sensors via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy; peak/mean drawn from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 0.434 · mean 31.3 W · baseline 10.1 W (box-idle-all-empty) · delta 21.2 W
- aihydra run-meta.jsonl entry: suite=tau2-kvpatch, model=qwen35-122b, variant="kvfix-q8". rc=2 (failed fast, 55s) — early-exit failure, window kept as it is a real recorded draw, not a zero-duration preflight guard. no run-NNNN file exists yet for this window; run left unlinked ([]).
eng-0038
- 2026-08-11 · wall-meter · AI Hydra energy/power sensors via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy; peak/mean drawn from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 116.723 · mean 65.3 W · baseline 10.1 W (box-idle-all-empty) · delta 55.2 W
- aihydra run-meta.jsonl entry: suite=llama-bench, model=nemotron3-super, variant="UD-Q4_K_M-first-look". no run-NNNN file exists yet for this window; run left unlinked ([]).
eng-0039
- 2026-08-11 · wall-meter · AI Hydra energy/power sensors via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy; peak/mean drawn from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 620.052 · mean 154.5 W · baseline 10.1 W (box-idle-all-empty) · delta 144.4 W
- covers: run-0103
- aihydra run-meta.jsonl entry: suite=tau2-kvpatch, model=qwen35-122b, variant="kvfix-f16". rc=124 (timed out) — window is real, arm did not complete but drew power the whole time. linked to run-0103 (ingested 2026-08-12 via scripts/ingest-tau2.mjs).
eng-0040
- 2026-08-11 · wall-meter · AI Hydra energy/power sensors via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy; peak/mean drawn from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 170.496 · mean 164 W · baseline 10.1 W (box-idle-all-empty) · delta 153.9 W
- covers: run-0104
- aihydra run-meta.jsonl entry: suite=tau2-kvpatch, model=qwen35-122b, variant="kvfix-q8". linked to run-0104 (ingested 2026-08-12 via scripts/ingest-tau2.mjs).
eng-0041
- 2026-08-11 · wall-meter · AI Hydra energy/power sensors via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy; peak/mean drawn from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 97.594 · mean 63.1 W · baseline 10.1 W (box-idle-all-empty) · delta 53 W
- aihydra run-meta.jsonl entry: suite=anchor-bench, model=qwen35-122b, variant="OLD-3653e6d". no run-NNNN file exists yet for this window; run left unlinked ([]).
eng-0042
- 2026-08-11 · wall-meter · AI Hydra energy/power sensors via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy; peak/mean drawn from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 55.131 · mean 162.7 W · baseline 10.1 W (box-idle-all-empty) · delta 152.6 W
- covers: run-0105
- aihydra run-meta.jsonl entry: suite=anchor-tau2, model=qwen35-122b, variant="OLD-3653e6d". linked to run-0105 (ingested 2026-08-12 via scripts/ingest-tau2.mjs).
eng-0043
- 2026-08-11 · wall-meter · AI Hydra energy/power sensors via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy; peak/mean drawn from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 118.131 · mean 61.1 W · baseline 10.1 W (box-idle-all-empty) · delta 51 W
- aihydra run-meta.jsonl entry: suite=anchor-bench, model=qwen35-122b, variant="NEW-62bf73d". no run-NNNN file exists yet for this window; run left unlinked ([]).
eng-0044
- 2026-08-11 · wall-meter · AI Hydra energy/power sensors via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy; peak/mean drawn from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 55.132 · mean 158.8 W · baseline 10.1 W (box-idle-all-empty) · delta 148.7 W
- covers: run-0106
- aihydra run-meta.jsonl entry: suite=anchor-tau2, model=qwen35-122b, variant="NEW-62bf73d". linked to run-0106 (ingested 2026-08-12 via scripts/ingest-tau2.mjs).
eng-0045
- 2026-08-11 · wall-meter · AI Hydra energy/power sensors via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy; peak/mean drawn from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 0.153 · mean 27.5 W · baseline 10.1 W (box-idle-all-empty) · delta 17.4 W
- aihydra run-meta.jsonl entry: suite=muse-fit, model=muse-glimmer-30b, variant="min-62bf73d". no run-NNNN file exists yet for this window; run left unlinked ([]).
eng-0046
- 2026-08-11 · wall-meter · AI Hydra energy/power sensors via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy; peak/mean drawn from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 1.861 · mean 167.5 W · baseline 10.1 W (box-idle-all-empty) · delta 157.4 W
- aihydra run-meta.jsonl entry: suite=muse-guard, model=muse-glimmer-30b, variant="min-62bf73d". rc=1 (failed) — window kept, real recorded draw. no run-NNNN file exists yet for this window; run left unlinked ([]).
eng-0047
- 2026-08-11 · wall-meter · AI Hydra energy/power sensors via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy; peak/mean drawn from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 86.118 · mean 172.2 W · baseline 10.1 W (box-idle-all-empty) · delta 162.1 W
- aihydra run-meta.jsonl entry: suite=muse-smoke, model=muse-glimmer-30b, variant="min-62bf73d". rc=124 (timed out) — window is real, arm did not complete but drew power the whole time. no run-NNNN file exists yet for this window; run left unlinked ([]).
eng-0048
- 2026-08-11 · wall-meter · AI Hydra energy/power sensors via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy; peak/mean drawn from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 1.061 · mean 54.6 W · baseline 10.1 W (box-idle-all-empty) · delta 44.5 W
- aihydra run-meta.jsonl entry: suite=laguna-fit, model=laguna-s21, variant="min-62bf73d". no run-NNNN file exists yet for this window; run left unlinked ([]).
eng-0049
- 2026-08-11 · wall-meter · AI Hydra energy/power sensors via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy; peak/mean drawn from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 0.995 · mean 89.6 W · baseline 10.1 W (box-idle-all-empty) · delta 79.5 W
- aihydra run-meta.jsonl entry: suite=bf16pr-matrix, model=qwen36-35b, variant="ROCm0-f16f16-p1024n256". no run-NNNN file exists yet for this window; run left unlinked ([]).
eng-0050
- 2026-08-11 · wall-meter · AI Hydra energy/power sensors via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy; peak/mean drawn from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 15.113 · mean 183.5 W · baseline 10.1 W (box-idle-all-empty) · delta 173.4 W
- aihydra run-meta.jsonl entry: suite=bf16pr-matrix, model=qwen36-35b, variant="ROCm0-f16f16-p32768proxy". no run-NNNN file exists yet for this window; run left unlinked ([]).
eng-0051
- 2026-08-11 · wall-meter · AI Hydra energy/power sensors via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy; peak/mean drawn from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 1.492 · mean 107.4 W · baseline 10.1 W (box-idle-all-empty) · delta 97.3 W
- aihydra run-meta.jsonl entry: suite=bf16pr-matrix, model=qwen36-35b, variant="ROCm0-bf16bf16-p1024n256". no run-NNNN file exists yet for this window; run left unlinked ([]).
eng-0052
- 2026-08-11 · wall-meter · AI Hydra energy/power sensors via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy; peak/mean drawn from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 14.041 · mean 179.1 W · baseline 10.1 W (box-idle-all-empty) · delta 169 W
- aihydra run-meta.jsonl entry: suite=bf16pr-matrix, model=qwen36-35b, variant="ROCm0-bf16bf16-p32768proxy". no run-NNNN file exists yet for this window; run left unlinked ([]).
eng-0053
- 2026-08-11 · wall-meter · AI Hydra energy/power sensors via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy; peak/mean drawn from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 1.449 · mean 104.3 W · baseline 10.1 W (box-idle-all-empty) · delta 94.2 W
- aihydra run-meta.jsonl entry: suite=bf16pr-matrix, model=qwen36-35b, variant="Vulkan0-f16f16-p1024n256". no run-NNNN file exists yet for this window; run left unlinked ([]).
eng-0054
- 2026-08-11 · wall-meter · AI Hydra energy/power sensors via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy; peak/mean drawn from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 13.352 · mean 176.8 W · baseline 10.1 W (box-idle-all-empty) · delta 166.7 W
- aihydra run-meta.jsonl entry: suite=bf16pr-matrix, model=qwen36-35b, variant="Vulkan0-f16f16-p32768proxy". no run-NNNN file exists yet for this window; run left unlinked ([]).
eng-0055
- 2026-08-11 · wall-meter · AI Hydra energy/power sensors via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy; peak/mean drawn from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 1.097 · mean 98.7 W · baseline 10.1 W (box-idle-all-empty) · delta 88.6 W
- aihydra run-meta.jsonl entry: suite=bf16pr-matrix, model=qwen36-35b, variant="Vulkan0-bf16bf16-p1024n256". no run-NNNN file exists yet for this window; run left unlinked ([]).
eng-0056
- 2026-08-11 · wall-meter · AI Hydra energy/power sensors via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy; peak/mean drawn from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 17.168 · mean 176.5 W · baseline 10.1 W (box-idle-all-empty) · delta 166.4 W
- aihydra run-meta.jsonl entry: suite=bf16pr-matrix, model=qwen36-35b, variant="Vulkan0-bf16bf16-p32768proxy". no run-NNNN file exists yet for this window; run left unlinked ([]).
eng-0057
- 2026-08-11 · wall-meter · AI Hydra energy/power sensors via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy; peak/mean drawn from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 4.157 · mean 149.7 W · baseline 10.1 W (box-idle-all-empty) · delta 139.6 W
- aihydra run-meta.jsonl entry: suite=bf16pr-matrix-depthfix, model=qwen36-35b, variant="ROCm0-f16f16-p1024n256d32768-realdepth". no run-NNNN file exists yet for this window; run left unlinked ([]).
eng-0058
- 2026-08-11 · wall-meter · AI Hydra energy/power sensors via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy; peak/mean drawn from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 3.916 · mean 156.6 W · baseline 10.1 W (box-idle-all-empty) · delta 146.5 W
- aihydra run-meta.jsonl entry: suite=bf16pr-matrix-depthfix, model=qwen36-35b, variant="ROCm0-bf16bf16-p1024n256d32768-realdepth". no run-NNNN file exists yet for this window; run left unlinked ([]).
eng-0059
- 2026-08-11 · wall-meter · AI Hydra energy/power sensors via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy; peak/mean drawn from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 3.781 · mean 151.2 W · baseline 10.1 W (box-idle-all-empty) · delta 141.1 W
- aihydra run-meta.jsonl entry: suite=bf16pr-matrix-depthfix, model=qwen36-35b, variant="Vulkan0-f16f16-p1024n256d32768-realdepth". no run-NNNN file exists yet for this window; run left unlinked ([]).
eng-0060
- 2026-08-11 · wall-meter · AI Hydra energy/power sensors via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy; peak/mean drawn from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 4.834 · mean 158.2 W · baseline 10.1 W (box-idle-all-empty) · delta 148.1 W
- aihydra run-meta.jsonl entry: suite=bf16pr-matrix-depthfix, model=qwen36-35b, variant="Vulkan0-bf16bf16-p1024n256d32768-realdepth". no run-NNNN file exists yet for this window; run left unlinked ([]).
eng-0061
- 2026-08-11 · wall-meter · AI Hydra energy/power sensors via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy; peak/mean drawn from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 427.416 · mean 165.2 W · baseline 10.1 W (box-idle-all-empty) · delta 155.1 W
- covers: run-0107
- aihydra run-meta.jsonl entry: suite=nemotron-quant, model=nemotron3-super, variant="E1". linked to run-0107 (ingested 2026-08-12 via scripts/ingest-tau2.mjs).
eng-0062
- 2026-08-11 · wall-meter · AI Hydra energy/power sensors via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy; peak/mean drawn from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 234.142 · mean 168 W · baseline 10.1 W (box-idle-all-empty) · delta 157.9 W
- covers: run-0108
- aihydra run-meta.jsonl entry: suite=nemotron-quant, model=nemotron3-super, variant="E2". linked to run-0108 (ingested 2026-08-12 via scripts/ingest-tau2.mjs).
eng-0063
- 2026-08-11 · wall-meter · AI Hydra energy/power sensors via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy; peak/mean drawn from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 240.779 · mean 152.7 W · baseline 10.1 W (box-idle-all-empty) · delta 142.6 W
- covers: run-0109
- aihydra run-meta.jsonl entry: suite=sim-sensitivity, model=qwen35-122b, variant="F1". linked to run-0109 (ingested 2026-08-12 via scripts/ingest-tau2.mjs).
eng-0064
- 2026-08-11 · wall-meter · AI Hydra energy/power sensors via Home Assistant (counter-difference on sensor.hardware_ai_hydra_energy; peak/mean drawn from sensor.hardware_ai_hydra_power) · window recorded
- wh_total 65.027 · mean 137.6 W · baseline 10.1 W (box-idle-all-empty) · delta 127.5 W
- covers: run-0110
- aihydra run-meta.jsonl entry: suite=sim-sensitivity, model=qwen35-122b, variant="F2". linked to run-0110 (ingested 2026-08-12 via scripts/ingest-tau2.mjs).
Contributions (1)
con-0001 — ggml-org/llama.cpp (carrying)
- fork · opened 2026-07-21 · link
- carrying reproducibility hazard — in use but not upstream
- Disk slot save/restore silently loses all prompt reuse on hybrid/recurrent models because context checkpoints are never persisted. Fixed with a .ckpt sidecar and pushed to headbouyJB/llama.cpp@fix-25913. Independently confirmed working by two community testers. A third-party PR (#26004) fixes the same bug by appending checkpoints inside the save file instead; production carries our fork until one of the two lands upstream.
Claims (51)
clm-0001
- measured-here low ●○○ volatility medium · verified 2026-08-03
- Static pool utilisation did not predict the 2026-07-21 OOM. The box crashed at 83.8% estimated utilisation, and cfg-0002 sits at that same 83.8% today.
- total_gb is weights + KV + a fixed overhead — a static estimate. What actually exhausted memory was runtime growth: context checkpoints defaulting to 32, cache-ram, and fragmentation. None of it appears in the record. The mitigations worked precisely because they targeted dynamic growth (checkpoints capped at 4, desktop stack removed, swap raised to 16G) — which is why they moved the ratio not at all. Until observed_peak_gb is collected from telemetry, the 0.8 warn line is a threshold on a number that has never predicted an OOM. Falsifiable: collect peak RSS/GTT under load and compare against the estimate.
- evidence: inc-0004
clm-0002
- community med ●●○ volatility medium · verified 2026-08-03
- Multi-token prediction gains MORE at heavier quantisation, not less. On this silicon Q8_0 saw 2.44x (7.7 -> 18.1 t/s) against Q4_K_M's 1.81x (12.1 -> 21.2).
- Counterintuitive, and it matters for quant selection: baseline decode here is bandwidth-saturated, so MTP's only lever — fewer memory passes per token — pays more the heavier the weights. Partially offsets the cost of a higher quant. Not yet reproduced on our own hardware; the quant x MTP matrix is queued for the lab box. Our own measured figure is the 122B at UD-Q4_K_M: 20.96 t/s single-slot baseline to 31.17 with MTP n=3 (+49%), and 36.5 t/s after tuning to n=6.
clm-0003
- measured-here med ●●○ volatility low · verified 2026-07-20
- MTP's speedup tracks how predictable the text is: +81% on code, +21% on freeform, with the real agent workload mix landing around +29-40%.
- The headline "+50%" is the favourable end of a range, not a constant. Draft acceptance measured at 75%, averaging ~5.0 tokens per speculative call. Prefill paid a ~10% tax, discounted in practice by a ~90% cache-hit rate — so the tax lands only on cold turns. Single-slot cost zero decode, which is what made the architecture viable.
- evidence: run-0002
clm-0004
- measured-here med ●●○ volatility high · verified 2026-06-28
- gpt-oss-120b placed last on our blind conceptual set — 3/8, mean 1.38 — against Qwen3.6-35B-A3B's 7/8, mean 2.12.
- FINGERPRINT, without which this is a slight rather than a finding. Config cfg-0003: Q4_K_M, ROCm, -c 131072, q8_0 KV, reasoning OFF, `-ngl 999 --no-mmap --parallel 1 -fa 1 --jinja`, on a ~82 GiB budget. Tasks C1-C8, one response per model per task. Judging: a separate, fresh Claude Opus sub-agent per response, given only the task prompt, the rubric and one anonymised answer — no model name, no comparison set, no project context. Every rationale preserved and re-checkable. SCOPE: eight tasks, N=1 per cell, one quant, one date, one llama.cpp build, and a rubric aimed at proactive/generative assistant work. It is not a general capability verdict. The same model was the FASTEST reactive performer in the field (~5 s). Volatility high: this was true of that build on that date and has not been re-tested.
- evidence: run-0003
clm-0005
- measured-here high ●●● volatility low · verified 2026-06-28
- Nemotron 3 Super's benchmark result is NOT quant-equivalent to its peers and must not be read as a like-for-like verdict on the model.
- Every model was tested at the best quant that fits 128K in ~82 GiB. For the field that was a Q4_K-class quant; for Nemotron it was IQ4_XS, because its Q4_K needs 77 GiB at 4K context and OOMs by 16K. It was additionally the only entry with no speculative decoding available — llama.cpp does not support the Mamba-hybrid MTP path — while peers could use it. So it was handicapped twice, and came last on speed (~28 s). The comparison is honest about what THIS BOX can run; it is not honest about the model. This caveat travels with the result rather than sitting in a footnote, which is the whole reason it is a record and not a sentence.
- evidence: run-0003
clm-0006
- community med ●●○ volatility high · verified 2026-08-03
- Quantized KV cache is mis-implemented for Strix Halo in stock llama.cpp: the code dequantizes to full precision repeatedly during inference, which a discrete GPU hides in cache and this box cannot. Community fixes report ROCm text generation +75% to +203% depending on depth, and q8_0 KV generating 23-53% FASTER than f16.
- Source: independent 128GB verification (Nathanw1014) of a fork by another user, four controlled builds — patched and stock, Vulkan and ROCm — flash attention pinned, build hash and flags carried per row. Vulkan prefill +45%/+71%/+87% at 32k/64k/128k on Qwen3-Coder-30B. Full 262,144-token native context runs on both backends; at 262k the compressed cache generates 65% faster for a 2.6% prefill cost. One reported regression: MoE loses 0.3-2.4% ROCm prompt speed with the fix. DIRECTLY LOAD-BEARING FOR US: cfg-0002 runs q8_0/q8_0 KV on ROCm at 200K context, which is the deep end where the reported gain is largest — potentially a bigger win than MTP's +50%. NOT REPRODUCED HERE, and not upstream: no PR exists in ggml-org/llama.cpp, so adopting it means carrying a SECOND fork alongside con-0001. The author states plainly that none of it measures output quality.
clm-0007
- inferred low ●○○ volatility medium · verified 2026-08-03
- A GPU cache wall around 32-40 MB, past which read speed drops roughly 4x, is a candidate explanation for our own prefill degradation from ~354 tok/s at small context to ~144 tok/s at 128K.
- The wall was measured on another 128GB Strix Halo box with a purpose-built benchmark; the originating theory about where the collapse should begin did NOT match the measured curves, so the mechanism is real but not yet modelled. Our ~2.5x prefill drop is the right order for a ~4x bandwidth cliff partially amortised. TWO REASONS TO BE CAUTIOUS: this is a GPU-cache effect, entirely distinct from the pool-utilisation metric in clm-0001 — conflating them would be a mistake; and the reporter's dense model collapsed while his MoE shrugged the same conditions off, which is unexplained. We run MoE, so the q4_0 mitigation that helped his dense model may not transfer. Falsifiable: run the same cache benchmark on aibeast and compare the knee against our measured prefill curve.
- evidence: run-0003
clm-0008
- community low ●○○ volatility high · verified 2026-08-05
- Identity-grounded persistent sessions with a shared memory and peer-to-peer messaging can self-organise into useful working groups without an orchestrator. Evaluated 2026-08-05 and DEFERRED: the prerequisite is many warm concurrent lanes, which our single-KV-slot architecture cannot provide.
- Source: a white paper describing 8 autonomous OpenClaw sessions ("Council of Minions") self-organising into 3 teams over ~3 hours, coordinated via a shared MongoDB brain and sessions_send, at ~$0.10 of tokens. THE INSIGHT WORTH KEEPING: "an orchestrator coordinates tasks; a council coordinates attention." Their strongest evidence is that nobody had been ASSIGNED to notice auth was added to a new engine and never backfilled to the old one — a task decomposition can only cover work you can name in advance. That limitation is real and matches our own experience: the R1 scorer artefact was caught by reading transcripts, not by any metric we had designed. SCEPTICISM: the headline ROI ($0.10 vs EUR 1,500-4,000 of consultancy) prices output volume as professional deliverable. "79 of 98 routes unauthenticated" is a grep. The cross-verification is agents agreeing with agents, with no independent validation. And the emergence is partly pre-seeded — an identity file that says "Auth, Secrets and Identity Security" makes a credential audit declared rather than emergent. Their gateway was also at ~50% RAM across 555 sessions, which is not a healthy system. WHY WE CANNOT DO IT TODAY, AND IT IS NOT ABOUT POWER: their architecture trades model capability for concurrency; ours does the opposite. MTP forces -np 1, so we have ONE KV slot behind one very large model. Eight concurrent sessions would each cold-prefill and destroy the warm lane. The prerequisite is many CHEAP WARM LANES, not more compute. THE OPERATOR'S EXTENSION (2026-08-05), which is the stronger form: use several DIFFERENT small specialist models rather than many instances of one. The paper ran 8 copies of a single model, so all 8 shared its blind spots — diversity of attention, not of capability. Different models have genuinely different failure modes, which is the same principle underpinning our blind-judging methodology. Our own field showed real profile differences: gpt-oss-120b terse to the point of under-delivery, 35B-A3B best on conceptual quality, 27B strongest on code. FEASIBLE SHAPE once aihydra lands: 122B deep reasoner (aibeast) + 35B-A3B fast generalist and a coder model (aihydra) + a small classifier on the M4 mini. Four genuinely different skills across three machines — heterogeneous by hardware as well as by prompt. ⚑ CORRECTED 2026-08-05 (operator catch). I filed this as a design conflict between self-cycling and Warden's restraint. That was the wrong layer. Restraint lives at the DELIVERY GATE (loop3: daily_cap, quiet hours, validity check, teaser-first), not in deliberation. Deliberation is already separated from delivery by design — lots may happen in the background; the operator is involved only when something of value needs delivering AND it is the right moment. So a self-cycling council in loop2 costs COMPUTE, not attention, and the conflict dissolves. THE RESULTING SHAPE — self-cycling as the first escalation tier: sentinel (cheap, continuous, high-recall, deliberately low-precision) flags a change -> council self-cycles on that one thing, bounded -> exits, and only then does loop3 decide if and when the operator hears about it. THREE EXITS, not two. The operator named "nothing is wrong" and "value delivered"; the third is BUDGET EXHAUSTED, INCONCLUSIVE — and it is the one that bites, because both clean exits are conclusions and real investigations often reach neither. Without it a council does not fail loudly, it quietly eats the box. BOUNDARY PRECISION: the council does not deliver. It produces a CANDIDATE WORTH QUEUEING. Loop3 remains the only path to the operator. A council that believes it can deliver has escaped the gate. FAILURE MODES PER TIER: sentinel too sensitive -> constant convening, cost; too conservative -> the 6h blind spot returns; council unbounded -> runs forever; council concluding "nothing wrong" too eagerly -> misses the thing. OUR OWN INCIDENT IS THE SPECIFICATION: inc-0004 saw the box crash-looping for 46 minutes with three llama-server kills before the kernel panicked, and nothing surfaced any of it — with the risk flagged in that morning's brief. Three process kills in 46 minutes is exactly what a change-detector flags, and "why does this keep restarting" is exactly a bounded investigation with a clear conclusive exit. TRANSFERABLE ENGINEERING (the most useful part): their inter-session sends timed out at a ~10s gateway default against 30-60s cold starts, and the fix was retry-with-backoff plus spill-to-disk staging rather than a larger timeout. That is directly applicable to our own scheduler-preemption work.
clm-0009
- community med ●●○ volatility high · verified 2026-08-05
- A ROCmFP4 iMatrix quant of our exact model (Qwen3.5-122B-A10B) reports 60.70 GiB and 28.505 tok/s decode with MTP OFF — against our measured 20.96 tok/s MTP-off at UD-Q4_K_M. That is roughly +36% decode for ~5-6 GB less memory, if it reproduces.
- Source: vmlinux/Qwen3.5-122B-A10B-ROCmFP4-iMatrix-GGUF on HuggingFace, updated 2026-07-26. Built with ROCmFPX's Q4_0_ROCMFP4_STRIX_LEAN recipe and tested ONLY on gfx1151 — our exact silicon. Published figures: 60.70 GiB, greedy decode 28.505 tok/s, 4,277-token prefill 356.9 tok/s, and BF16 mean KLD 0.041366 +/- 0.002531. NOTABLE ON METHOD: they publish KL divergence WITH an uncertainty interval — the same instrument we selected for KV-quant quality, applied to weight quantisation. That gives us a comparable reference point for what a small quality delta looks like on this model, which we did not previously have. WHY IT MATTERS FOR PLACEMENT: 60.70 GiB against our ~66 GB frees ~5-6 GB. Config 2 in model-phases.md has the 122B and 35B co-resident at ~108 of ~120 GiB, which we flagged as uncomfortably tight given our OOM history. This quant would meaningfully relieve that — possibly the difference between viable and reckless. ⚠ THE COST: it uses custom ROCmFP4 tensor types and **stock llama.cpp cannot load it**. It requires the ROCmFPX runtime. That is a THIRD carried fork alongside con-0001 (our checkpoint sidecar) and the community quantized-KV fix — and our reproducibility debt is already a stated concern. Worse for measurement: a different runtime is a different config fingerprint, so ROCmFP4 numbers CANNOT sit in the same series as our existing ones (protocol section 8). It is a new experiment, not a faster version of the old one. Our Tier 1 pipeline also assumes stock llama-bench, which would need the ROCmFPX build. ALSO: the MTP heads ship as a SEPARATE 2.14 GiB companion file, where our current GGUF embeds them. Flags: "experimental", 1,929 downloads, 6 likes — low adoption, so we would be early rather than following. TEST IT AS A LAB-BOX CANDIDATE, not a production swap: same suite, same context, against our own UD-Q4_K_M baseline, with the runtime difference recorded as the config change it is. If +36% MTP-off holds and MTP scales on top, it plausibly lands near 42 tok/s against our current 36.5.
clm-0010
- community med ●●○ volatility high · verified 2026-08-05
- Draft-free ngram speculation (ngram-mod) reportedly beats MTP by a wide margin on repetitive/agentic work on Strix Halo — 71 t/s unspeculated to 216 solo on a code-edit probe, with a shared hash pool letting concurrent streams feed each other's drafts (247 pooled across 4 streams, later 302 end-to-end on HIP).
- Source: a detailed Strix Halo write-up (HP ZBook, Ryzen AI MAX+ PRO 395, 128 GB, Windows, Vulkan then HIP) on a Qwen3.6-35B-A3B finetune in ROCmFP4 via the ciru-ai/ROCmFPX fork. MECHANISM: the speculator matches n-grams against text already in the context window, so the draft costs nothing to produce and only verify rows are paid. Reported 86% acceptance on 48-64 token copied drafts, and on fresh content the gate simply stays closed with no penalty. That asymmetry is what makes it different from a draft model. WHY IT MATTERS TO US MORE THAN MOST: our heaviest repetitive workload is exactly their best case — Claude Code doing code edits, re-emitting large chunks of a file. And the shared-pool effect (acceptance RISING with stream count, 93% to 97%) is a genuine argument for the council pattern in clm-0008, where several agents work adjacent problems. ⚠ MTP IS THROUGHPUT-NEGATIVE AT BATCH per the same source: roughly breakeven at 2 streams, -13% at 4, because every accepted draft still costs a verify row through the MoE and the expert-union already saturates the bus. We run -np 1, so we are in MTP's favourable regime — but this is a further argument against multi-slot, on top of llama.cpp #25992 and the memory cost. CAVEATS: single reporter; the 4-stream figures used IDENTICAL prompts across streams, which is maximum pool sharing and flattering (they say so). Requires the ROCmFPX fork, which would be a third carried runtime (see clm-0009). Their own updates report ngram fighting flash attention at larger contexts, and crashing when combined with MTP. TEST: ngram-mod vs MTP vs neither, on OUR code-edit workload, single stream. It is the cheapest large win available if it reproduces.
clm-0011
- community med ●●○ volatility medium · verified 2026-08-05
- Greedy decode is NOT run-to-run deterministic on the Vulkan backend — same config, same prompt, temperature 0, batch 1, no speculation, three different outputs.
- Reported on Strix Halo/gfx1151 with outputs diverging mid-generation. Deep MTP (n>=4) amplifies it to first-token divergence via draft-length to batch-shape variance. The reporter is explicit that outputs stay coherent — this is variance, not corruption. DIRECT CORRECTION TO OUR PROTOCOL: benchmark-protocol.md adopted "temperature 0, fixed seed" as a determinism control, borrowed from homebench. On this backend that control does not hold, which has two consequences. First, N=1 is not defensible for anything output-dependent even where we assumed determinism. Second, and worse: **speculative losslessness cannot be verified by diffing outputs on Vulkan**, which is exactly how one would naturally check that MTP or ngram speculation is not changing results. UNVERIFIED HERE, and worth checking on ROCm rather than assuming it transfers — the report is Vulkan-specific and our production backend is ROCm. If it holds on ROCm too, quality comparisons need a distributional instrument (KL divergence) rather than output equality, which is the approach we had already chosen for KV quant.
clm-0012
- community low ●○○ volatility high · verified 2026-08-05
- 128 GB Strix Halo systems appear to be repricing upward materially — from a roughly $2,500-4,000 enthusiast tier toward $4,500-6,000 — with 128 GB SKUs scarce while 64 GB configurations remain available.
- Source: a community PSA aggregating six vendors (Bosgame, GMKtec, ACEMAGIC, Framework, Corsair, MSI). Direction is well-evidenced across independent listings; MAGNITUDE is not — the post itself flags ACEMAGIC's $4,999 as possibly a placeholder, and Corsair's jump came with reopened preorders, which could be a reset rather than a trend. ROOT CAUSE MATCHES WHAT WE FOUND ELSEWHERE: this is the same memory squeeze that delayed NVIDIA's RTX 50 Super refresh to CES 2027, where 3 GB GDDR7 modules reportedly cost $60-70 against $20 for 2 GB parts. Here it lands on high-density LPDDR5X for the 256-bit bus. The memory being SOLDERED is what gives manufacturers pricing power specifically on the high-capacity SKUs — a buyer cannot take a cheaper configuration and upgrade later — and it explains 64 GB staying in stock while 128 GB does not. WHAT IT CHANGES FOR US, which is why it is recorded at all: · A THIRD box becomes unlikely, so the two-box architecture is a durable design rather than a transitional one. · It WEAKENS the earlier "wait for Medusa Halo (LPDDR6, ~460 GB/s, 2027)" argument. A part needing even denser memory will be priced in the new tier, not the old one. · It STRENGTHENS the cross-host overlap case in the integration plan. That was justified on "most people replace rather than overlap, so this comparison is rare". If replacement is now expensive, keeping both boxes long-term is likely — and aibeast's 96 GB appreciates in usefulness rather than depreciating.
clm-0013
- community high ●●● volatility medium · verified 2026-08-05
- Qwen3.6-35B-A3B — our chosen reflex model — is published in FastFlowLM's .q4nx NPU format (23.2 GB, plus a 1.0 GB vision encoder), making a genuinely capable MoE, not a 1-2B classifier, runnable on the XDNA2 NPU.
- FastFlowLM/Qwen3.6-35B-A3B-NPU2, created 2026-07-08, ~2,000 downloads. WHY A 35B IS SUDDENLY PLAUSIBLE ON AN NPU: decode bandwidth is proportional to ACTIVE parameters, not total. At A3B the model reads roughly 2 GB per token at 4-bit, where a dense 35B would read ~20 GB. Low-active-parameter MoE is the ideal shape for a bandwidth-poor accelerator — which is exactly why the NPU catalogue was previously 1-8B dense models. WHAT IT CHANGES: the second-lane argument in npu-fastflowlm-2026-08-02 assumed the NPU could host a small classifier. If it can host our actual reflex model, the sentinel tier in clm-0008 could be a capable council member rather than a change-detector. It also ships a vision encoder, which could relieve the Mac mini. ⚑ CORRECTED 2026-08-05 with measured Strix Halo data (Framework Desktop, gfx1151, LMDE 7 / xanmod 7.1.3). I had framed the NPU as "slower at decode but wins prefill, so the value is concurrency". **On a 35B, the NPU loses BOTH phases decisively.** | ctx | NPU decode | NPU prefill | NPU TTFT | |-----|-----------|-------------|----------| | 1k | 12.11 t/s | 90.7 t/s | 10.8 s | | 8k | 10.94 | 209.9 | 36.9 s | | 32k | 8.11 | 245.2 | **126.3 s** | Same box, same model at a HIGHER quant (Q8_0, 35.21 GiB) on the iGPU: **pp4096 709-840 t/s on ROCm, 964-1039 on Vulkan; tg128 43 t/s ROCm, 53 t/s Vulkan.** So the iGPU is roughly **4x faster at decode AND 4-6x faster at prefill**, while running a heavier quantisation. My earlier "NPU wins prefill 2.3x" came from an ~8-9B dense comparison; that advantage does not survive at 35B on this silicon. And TTFT of **126 seconds at 32k** rules the NPU out of anything interactive or escalation-shaped outright. WHAT SURVIVES: the NPU lane is only defensible as genuinely CONCURRENT background work where latency is irrelevant — Phase A/D in model-phases.md — and even then it costs the iGPU about 14% of its own decode. The "capable second council member" framing is much weaker than it looked an hour ago: a member that answers in two minutes is not a peer. ALSO STILL TRUE: mutually exclusive with amd_iommu=off (+5-12%), and aibeast's NPU logged SMU init errors on the 21 July wedged boot, so NPU health is unverified.
clm-0014
- community med ●●○ volatility high · verified 2026-08-05
- On gfx1151, Vulkan measured ~22-24% faster than ROCm on the same 35B-A3B model and build — pp4096 1039 vs 840 t/s, tg128 53.1 vs 43.4 — but the Vulkan build produced GARBAGE OUTPUT for that model, so the numbers describe a broken configuration.
- Measured on a Framework Desktop (Strix Halo, Q8_0, 35.21 GiB) with llama-bench. The reporter states plainly that the current llama.cpp Vulkan build returns a single garbage Chinese character when actually asked anything with Qwen3.6. ⚠ THE METHODOLOGICAL POINT MATTERS MORE THAN THE NUMBERS. **llama-bench measures throughput and never checks that the output is valid.** A backend can be 24% "faster" while emitting nothing usable, and the benchmark cannot tell. Our Tier 1 is built on llama-bench, so we inherit that blind spot exactly. Fix adopted: an OUTPUT SANITY GATE before any throughput number is recorded — one real generation per (model, backend, build), checked for coherent text, and the throughput discarded if it fails. Cheap, and it is the difference between "Vulkan is faster" and "Vulkan is broken but quick about it". SECONDARY, AND DIRECTLY USEFUL: the same runs sweep ubatch. Prefill peaks around **ub 1024-2048** on both backends (ROCm 715 -> 833 t/s from 512 to 2048; Vulkan 964 -> 1039 at 1024) and DEGRADES at 4096. Decode is flat across ubatch, as expected. That narrows our own planned sweep to a sensible range rather than guessing.
clm-0015
- community high ●●● volatility high · verified 2026-08-06
- Ling-3.0-flash (Ant Group, 26 Jul 2026) is a 124B/5.1B-active hybrid MoE that is a near-ideal controlled comparison for our Qwen3.5-122B-A10B — same footprint, half the active parameters — but it cannot be benchmarked on a stock llama.cpp, and its MTP head ships INACTIVE, which would silently rig any decode comparison in Qwen's favour.
- STRUCTURAL FACTS (from the GGUF model card and the upstream PR; high confidence). 124B total, **5.1B active** per token — note 5.1, not 5.2. 42 layers: **35 KDA (Kimi Delta Attention) + 7 gated MLA** at layers 5/11/17/23/29/35/41, plus an MTP/nextn head at layer 42. Native 256K context; the GGUFs expose 262,144 tokens, double the 131,072 in config.json. Q4_K_M = **69.70 GB**, Q5_K_M = 84.18 GB, Q8_0 = 126.32 GB. NO PERFORMANCE CLAIM HERE IS VERIFIED. The vendor line — matching their 1T flagship on most benchmarks at 1/8 the total and 1/12 the active parameters — is marketing, and every number in it is exactly what we would be measuring. Recorded as motivation for a benchmark, not as a result. WHY IT IS WORTH THE SLOT: against Qwen3.5-122B-A10B this is almost a controlled experiment. 124B vs 122B total (near-identical memory footprint), 5.1B vs 10B active (half), both MoE, both hybrid, both carrying MTP heads. It isolates ACTIVE PARAMETER COUNT more cleanly than any other pairing available to us. ⚠ THE TRAP — THE MTP HEAD IS PRESENT BUT "UNUSED BY THE GRAPH". It loads without consuming VRAM and can be enabled later without requantisation, but today it is inert. Our production Qwen runs MTP and we measured +50% decode from it (mtp-single-slot decision, 2026-07-20). So a naive Ling-vs-Qwen decode comparison is **MTP-off versus MTP-on**, and would flatter Qwen by roughly the size of the effect we are trying to measure. This promotes the already-open "MTP-off Qwen build for the ratio" question from nice-to-have to a PRECONDITION for this comparison being meaningful at all. THE HYBRID CONFOUND, NOW QUANTIFIED. bench/kv-quality.sh warns that -ctk/-ctv only touch full-attention layers, so a near-zero KL divergence would be an ARCHITECTURE result rather than a quantisation result. Here the number is **7 of 42 layers**. KV-quant should barely register, and we can state that up front instead of discovering it. The corollary is the attractive part: with 35 of 42 layers holding constant-size recurrent state, KV growth at 256K should be far below what our current model demands. BUILD STATUS — NOT UPSTREAM. Files declare `general.architecture = bailing-hybrid`, a provisional name; upstream PR #26608 (opened **2026-08-05**, unmerged) proposes `bailingmoe3`, continuing bailingmoe → bailingmoe2. Stock builds refuse the model with "unknown model architecture". A patch against llama.cpp commit 6ea215d17 ships with the GGUFs; a fork exists at aetherbird/llama.cpp:bailingmoe3-support. Because the arch name is provisional, **today's GGUFs may need re-downloading** once upstream lands — 70 GB of reason not to rush. PROTOCOL CONSEQUENCE 1 — fingerprint. `build_commit` is a CONFIG_KEY. A patched-build Ling result and our patched-build Qwen result come from two DIFFERENT forks, so they are not the same experiment and must not be presented as one row set without saying so. PROTOCOL CONSEQUENCE 2 — the sanity gate is not strong enough here. The GGUF notes `rope_interleave: true` resolving to NORM rope rather than the NEOX that DeepSeek-style MLA usually uses. A subtly wrong rope on a fork build produces output that passes sweep.sh's 24-token coherence check and then degrades at depth. **For any unsupported-architecture build, the precondition should be a long-context retrieval check (RULER-lite), not a short coherence check.** This is a gap in our own protocol that this model exposed. FIT, AND AN ARGUMENT FOR THE 128 GB BOX. Q4_K_M at 69.7 GB fits aibeast's 96 GB with room for KV and compute — helped by only 7 layers holding a real cache. Q5_K_M at 84.18 GB does not fit aibeast's ~84 GiB GTT but is comfortable on aihydra's 128 GB. That is the first concrete case where the larger box buys a quantisation level rather than just headroom. UNVERIFIED RISK: the published quants use "SM_120-safe" types chosen around NVIDIA Blackwell (no iq1_s/iq2_s/iq3_s). Irrelevant to gfx1151 in intent, but it means these files were built and tested against CUDA, and their behaviour on ROCm is untested. A GAP WE COULD ACTUALLY FILL: the model card states reasoning quality beyond 128K is **unmeasured**. We have the KL-divergence and RULER-lite tooling and a box that fits 256K. That is a contribution-shaped hole, not just a benchmark. RECOMMENDATION: add to the bench queue, do NOT chase the fork. aibeast is dead — the 2026-08-07 hands-on attempt found a board that will not POST and a warranty claim is open (inc-0005), so the benchmark host is now aihydra. A second fork alongside our #25913 patch is real maintenance cost. Wait for #26608 to merge and fold it into the llama.cpp rebuild already queued for return (Step 1b) — that gets the stable arch name, the final GGUFs and the support in one move.
clm-0016
- community med ●●○ volatility medium · verified 2026-08-06
- Measured on Strix Halo: Qwen3.6-35B-A3B under ROCmFP4 + HIP + ngram-mod at parallel 4 sustains 121 tok/s across 500 varied IFEval prompts against a 64.8 tok/s no-speculation floor — a 1.87x production speedup with IFEval-strict at 78.6% — while prefill collapses from 1,211 tok/s cold to 136 tok/s at ~243k depth.
- Single community source (the ACE-SABER follow-up post), self-reported, on someone else's Strix Halo. Confidence is medium DESPITE the detail, because it is one operator on one box and none of it is independently reproduced. The internal arithmetic does check out: 1.32M tokens / 3.0 hours = 122 tok/s, consistent with the stated 121. CONFIGURATION (all five are fingerprint fields, so this is ONE config point, not a decomposition): ROCmFP4 quant, HIP backend, ngram-mod speculation with shared hash pool, --parallel 4, f16 KV. THE NUMBER THAT MATTERS: **121 tok/s sustained vs a 64.8 tok/s no-speculation floor = 1.87x**. Both figures come from the same operator under the same conditions, which makes this a genuine isolation of the speculation effect rather than a headline. The author explicitly labels 121 as "the production number". WHAT THE AUTHOR DISCARDS, AND WHY IT MATTERS MORE THAN WHAT THEY KEEP: 430 tok/s (same-prompt repeat) is flagged as an artifact — only one session doing real work while three ride the ngram pool. 302-305 tok/s (4 streams, identical content) is flagged as "inflated by max pool sharing". Only 380 tok/s (single stream, real agentic session) and 121 tok/s (sustained, varied) are claimed as trustworthy. This is exactly the discipline our protocol demands and it is why this post is usable at all. ⚠ IT ALSO CONFIRMS A HAZARD WE ALREADY FLAGGED. model-phases §3a notes that ngram-mod's shared pool only works within one llama-server, so isolation and speed conflict. This data shows the pool ALSO corrupts measurement: identical content across slots inflates throughput because the slots feed each other's n-gram cache. **Any multi-slot benchmark of ngram-mod must use VARIED prompts or it measures the pool, not the model.** Our sweep.sh does not currently guard against this. PREFILL DECAY AT DEPTH — THE FINDING WITH THE MOST CONSEQUENCE FOR US: | prompt | prefill | |---|---| | 8.1k cold | 1,211 tok/s | | 2.2k blob at ~3k depth | 1,186 tok/s | | 2.2k blob at ~13k depth | 899 tok/s | | 2.2k blob at ~243k depth | **136 tok/s** | Read carefully: that last row is the MARGINAL cost of appending 2.2k tokens when 243k are already resident — roughly **16 seconds of prefill before a single token is generated**, every turn, at depth. It is not the cost of prefilling 243k from scratch. This is the strongest external validation yet of the warm-lane and disk-restore work: our measured restore was 105 ms for 44k tokens, so restore-versus-reprefill is the difference between milliseconds and tens of seconds per turn. TTFT corroborates it — 6.7 s for a cold 8.1k prompt versus **71 ms** for turn 2 on a cached prefix. QUALITY — AND WHAT THE IFEval NUMBER ACTUALLY TESTS. IFEval strict 78.6% / loose 79.8%, reported as no measurable drop. Worth being precise about what that establishes: correctly implemented speculative decoding is mathematically exact — the verification step reproduces the base model's distribution — so an unchanged score is the EXPECTED result and mainly evidences that the ngram implementation is not lossy. The more interesting half is that it also passes under **ROCmFP4**, which is a genuine quantisation change and could have cost accuracy. Note too that IFEval measures instruction-following only: it says nothing about long-context retrieval at 243k, nor about tool calling, which are the two axes we care most about. The author says more benchmarks are coming. SCOPE LIMIT — 27B GETS NO SIMILAR BOOST. Reported plainly, and the mechanism is plausible: speculation pays in proportion to how memory-bound decode is, because batch verification of drafted tokens is nearly free when you are bandwidth-starved rather than compute-starved. A3B activates 3B; a dense 27B activates all 27B and is far more compute-bound, so there is less idle bandwidth for speculation to exploit. **Prediction worth testing rather than assuming:** the benefit should scale inversely with active parameters — largest at Ling-3.0-flash (5.1B active, clm-0015), smaller but real on our Qwen3.5-122B-A10B (10B active), negligible on dense models. VULKAN, SECOND INDEPENDENT REPORT: "Vulkan? Nope, has a bug that's kicking in, HIP is the fix." This is now the second unrelated account of Vulkan misbehaving on gfx1151 with a Qwen3.6 model, alongside our own note about a backend being faster while emitting garbage. Two sources is not proof, but it is enough to make HIP the default for our sweeps and to treat any Vulkan throughput advantage on this silicon as suspect until output is verified. WHAT WE SHOULD TAKE FROM IT: a quality anchor (IFEval strict 78.6%) and a throughput target (121 tok/s sustained, 64.8 floor) that our own runs can be compared against, on the same silicon we own — provided we state the five-way config difference rather than presenting it as a like-for-like row.
clm-0017
- community high ●●● volatility high · verified 2026-08-06
- On Strix Halo Vulkan, a single unmerged patch — contiguizing strided f16 KV data before the flash-attention prefill — removes a dense-model prefill collapse worth 2.5x at 32k and 6.6x at 65k, changes decode not at all, and collapses run-to-run scatter from 5-10% to 0.3%. The variance change is the more useful finding: high scatter at depth is a SYMPTOM of a broken code path, not noise to be averaged away.
- Single community source, but the confidence is high because of the METHOD rather than the reputation: five binaries built from the same upstream base, changing only which patches were present, same cells on each. That is proper attribution — the variable is isolated rather than argued about. It is the standard our own §9 asks for and rarely sees in the wild. | dense 27B, f16 KV, FA on, Vulkan | pp512 @ 32k | pp512 @ 65k | |---|---|---| | stock upstream | 94.5 t/s | 29.8 t/s | | nine patches, contiguize REMOVED | 134.5 | 45.1 | | stock + contiguize and prerequisites | 237.4 | 180.4 | | all nine patches | 252.0 | 198.2 | Removing one patch from the set reinstates the collapse; adding it to an otherwise plain build removes it. Everything else in the set is worth a few percent. Baseline for scale: ~350 t/s at empty context, so stock loses >90% of prefill purely for holding a long conversation. ⚑ THE FINDING THAT CHANGES HOW WE MEASURE, not just what we build. On the broken path identical repetitions of the same deep cell varied by **5-10%**; on the fixed path they reproduce within **0.3%**. The author reports this retroactively explained a 20% standard deviation they had flagged months earlier and never understood. We currently treat `--reps 3` as a way to average noise out. That is exactly wrong when the variance is the signal. **A cell whose repetitions disagree by more than a few percent at depth should be treated as evidence of a broken code path and investigated, not reported as a clean mean with an error bar.** Actioned: sweep.sh now computes and reports per-cell coefficient of variation and flags high-scatter cells; protocol §0 carries the rule. DECODE IS UNAFFECTED — within 0.09% across all five builds. So this is purely a prefill path issue, which also means it would be invisible to any benchmark that only reports tokens/sec of generation. HOW MUCH OF THIS TRANSFERS TO US — genuinely uncertain, and the differences are large: · **Backend.** Measured on VULKAN. The author dropped ROCm entirely because Fedora 44 version locks blocked the upgrade, so they CANNOT test the HIP path. Whether an analogous collapse exists on ROCm is open. · **Architecture.** Measured on a DENSE 27B. Our models are MoE and mostly hybrid. For Ling-3.0-flash only 7 of 42 layers hold a real KV cache (clm-0015), so a flash-attention prefill defect would touch a minority of layers and the effect size should be much smaller — but "smaller" is not "zero", and we run `-fa 1`. · **Our own unexplained decay.** clm-0016 records prefill falling from 1,211 t/s to 136 t/s at ~243k depth on HIP with a 35B-A3B MoE. Different backend, different architecture, but the same SHAPE. Worth testing as a hypothesis rather than assumed related: is there an analogous contiguity problem on the HIP prefill path? IT ALSO COMPLICATES "VULKAN IS BROKEN". We hold two reports of Vulkan misbehaving on gfx1151 (clm-0014 garbage output, clm-0016's author reporting a bug that HIP avoids). This author runs Vulkan exclusively and gets good results once patched. The honest reading is that these are DIFFERENT defects — a model-specific correctness bug and a general FA prefill inefficiency — and we should stop collapsing them into "prefer HIP". FLASH ATTENTION IS NOW A LEVER, NOT A SETTING. The author's practical advice: on a patched build FA wins every cell and can be left on; on stock the old caution still applies. Our sweep.sh hardcoded `FA=1` with no way to vary it, which on a stock build means silently measuring the collapse. Actioned: `--fa` is now a sweepable argument. FORK COST. The patch is unmerged and under discussion upstream. Adopting it would make a THIRD fork alongside our #25913 .ckpt sidecar and a possible bailingmoe3 build — and these would have to be combined, not merely chosen between. Each fork is a permanent reproducibility tax: `build_commit` is a fingerprint field, so a combined tree is a configuration nobody else can reproduce. A CONTRIBUTION WE ARE UNUSUALLY WELL PLACED TO MAKE. The author explicitly asks for replication on stock builds — depth 32k+, FA explicitly on, three runs, because a single run tells you little on the broken path. We will shortly have two Strix Halo boxes and, critically, **ROCm — which the author no longer has**. The unanswered question is not "does this reproduce on Vulkan" but "does the collapse exist on the HIP path at all", and we are positioned to answer exactly that. Qwen3.6-27B is dense and available, so we can replicate the architecture rather than only approximating it with an MoE. Verified 2026-08-06 that the branch `strix-halo-fa-fixes` exists on the Nathanw1014/llama.cpp fork. Individual commit contents were NOT verified — the patch is described here as the author describes it, not as we have read it.
clm-0018
- community med ●●○ volatility high · verified 2026-08-06
- DeepGrove Maple-Preview (20.2B-A1.49B ternary, 5.31 GB, MIT) is a credible second-lane candidate for the Mac mini, but not a Warden candidate — its own model card concedes underperformance on agentic benchmarks. The "on-device weight adaptation (dreaming)" feature is undocumented at source, and even if it works it is architecturally opposed to how Warden's state is built.
- PROVENANCE MATTERS HERE MORE THAN USUAL — the claims arrive at three different levels of evidence and the most interesting one is the weakest: · **Model card (primary).** 20B-A1B, 24 layers, 256 experts / 8 active, ternary weights 2-bit packed as {-a, 0, +a} with one a per row, 5.31 GB checkpoint, 131,072 context, 3:1 SWA-512:global attention, MIT licence. Quantisations exist for llama.cpp, Ollama, LM Studio, Jan. Apple Silicon runs via the authors' own MLX fork (deepgrove-ai/mlx-lm-deepgrove). · **Model card, self-reported, NO NUMBERS.** Evaluated on LCBv6, AIME 2026, HMMT 2026, GPQA-D — but published as a *visual comparison chart* with no scores. Claims a new point on the Pareto frontier for memory-to-performance. · **Launch demo on X + aggregator blog only.** The "dreaming" weight adaptation, the ~5.9 GB adaptation peak, and the 200+ tok/s Mac mini figure. The model card says **nothing** about weight adaptation. The company describes adaptive behaviour partly in FUTURE TENSE ("plans to enhance"), so this is a demo and a roadmap, not a released, reproducible capability. 419 downloads in the last month — very low adoption. For scale, the NPU build in clm-0013 had ~2,000. Nothing here has been independently reproduced. ⛔ DISQUALIFYING FOR WARDEN'S PRIMARY ROLE, on the authors' own evidence. The model card acknowledges **underperformance on agentic benchmarks** and notes minimal post-training for agentic tasks. Warden is an agentic system whose stated failure mode is tool selection. A model that is weak at exactly that is not a candidate for the main lane, however good its reasoning scores turn out to be. This is the rare case where the vendor tells you the disqualifying fact themselves. ✅ WHERE IT IS GENUINELY INTERESTING — THE MAC MINI. wardenmac is an M4 Mac mini with 16 GB currently carrying only vision (MLX) and STT; no LLM lane. A 5.31 GB checkpoint at a claimed 200+ tok/s would fit with enormous headroom, and MLX is a path already proven on that box. That makes it a candidate SECOND LANE that does not depend on aibeast at all — which is the point, because the single-slot bottleneck is an aibeast problem. It is a more plausible route than the XDNA2 NPU option (clm-0013), where measured data showed the iGPU winning both phases by ~4x. ON "DREAMING" — THE NAME COLLIDES WITH OURS AND THE MECHANISM IS OPPOSITE. Warden's dreaming consolidates recalls into MEMORY.md: text, dated, diffable, revertible. Maple's dreaming runs a local fine-tune that embeds a preference **into the weights**. The architectural appeal is real — persistent preference that costs no context is a direct attack on our dominant cost, since the entire warm-lane programme exists because prefill of a large persistent prompt is expensive. But it is opposed to the principle the whole system is built on: **legible, auditable, revertible state.** `anticipatory-flags.md` declares the operator as its owner. Every recovery this month depended on being able to read state and diff it — the 2026-08-06 audit found a SESSION-STATE that had asserted a wrong host address for six weeks, and it was fixable precisely because it was text. A preference baked into weights cannot be inspected, diffed, dated, or switched off with a flag; you would discover a wrong one only by noticing bad behaviour, and then have no way to attribute it. You cannot `git diff` a weight delta into an explanation. If we ever pursue this, the defensible split is: weights may hold **stable, low-stakes, stylistic** adaptation (tone, formatting); anything with consequences stays in text. And even then the audit story is poor enough that it should be a deliberate experiment on a lab box, never on the production agent. TERNARY CAVEAT: replacing multiplies with additions is real, but the speedup needs kernels that exploit it. Community llama.cpp quants exist (e.g. stamsam/maple-preview-gguf), but ternary support in llama.cpp is narrower than standard K-quants and the Apple Silicon numbers come from the authors' own MLX fork rather than a neutral runtime. Treat the 200+ tok/s figure as vendor-path until measured on a stock runtime. IF WE TEST IT, the honest first question is not speed but capability: does a 1.49B-active model actually hold up? That is precisely what the capability tier exists for — and per protocol §1a it must be measured, not inherited, since this differs from everything we run by architecture, quant class and runtime simultaneously.
clm-0019
- measured-here high ●●● volatility medium · verified 2026-08-08
- On aihydra (gfx1151, ROCm 7.1, llama.cpp 3653e6d), running llama-server with `--parallel 4` destroys long-context needle retrieval — 0/8 across controlled trials — while `--parallel 1` on the same build, model and prompt succeeds 8/8. Throughput is unaffected and reports nothing wrong, so no performance benchmark would ever see it.
- MEASURED HERE, controlled, same session, same model file, same server binary, only the slot count varying: | server flags | needle before concurrency | after | |---|---|---| | `-c 65536 --parallel 4` (cache-idle-slots default) | 0/3 | 0/5 | | `-c 65536 --parallel 4 --no-cache-idle-slots` | 0/3 | 0/5 | | `-c 16384 --parallel 1` | **3/3** | **5/5** | Probe: a needle at position 0 ("The maintenance codeword is chartreuse-viper-88."), ~5,368 prompt tokens of whole-sentence filler, question at the end, `temperature 0`, thinking disabled. Failure mode is not garbage — the model answers coherently *from the filler* ("The maintenance codeword is: **routine**") or echoes the instruction ("codeword"). It behaves exactly as if the first part of its context is absent. WHY IT MATTERS MORE THAN IT LOOKS: · **Throughput is unaffected and silent.** decode/prefill numbers at `--parallel 4` look entirely normal. A sweep-only benchmark would have published them. · **It is a `hazard`-class lever behaving exactly as protocol §1a predicts** — a performance setting silently corrupting correctness. This is the first time our own guard framework has caught a real one, and it is the whole argument for the guard tier existing. · **It threatens the multi-slot plans directly** — the second-lane idea, the council architecture, and any use of parallel slots for concurrent agents. ⚠ DISTINCT FROM llama.cpp #25992. That issue is cross-request *leakage* between concurrent requests. We tested for it explicitly — four concurrent requests with disjoint markers — and found **no contamination**, repeatedly. This is a different failure: single sequential requests on a multi-slot server lose their own early context. HONEST ABOUT THE MESSY PATH: earlier ad-hoc observations in the same session were inconsistent (the same prompt passed 10/10 on one server instance and failed 6/6 on another with the same flags), and I twice attributed failures to my own harness before running a controlled comparison. The three-way matrix above is the only evidence that should be relied on; the earlier anecdotes are recorded here only so nobody re-derives them and thinks they contradict this. NOT YET ESTABLISHED — the open questions that would make this reportable upstream: · Does it depend on slot COUNT (2? 8?) or merely on >1? · Does it depend on `n_ctx_slot`, or on depth relative to it? Our failures were at ~5.4k tokens against a 16,384-token slot, so it is not a simple overflow. · Does it reproduce on CPU or Vulkan, or is it HIP/gfx1151-specific? · Does it reproduce on a stock upstream build on other hardware? If yes, this is an upstream bug worth filing; if no, it is a gfx1151 backend issue. IMMEDIATE OPERATIONAL CONSEQUENCE: **run `--parallel 1` for anything that depends on long-context recall** until the above is answered. That is what production already did (`cfg-0002` used `--parallel 1`), so nothing shipped is affected — but the plan to use multiple slots for a second lane is on hold pending this.
clm-0020
- measured-here high ●●● volatility medium · verified 2026-08-08
- On ROCm/gfx1151 with a stock llama.cpp, flash attention is unambiguously BETTER at depth — at 32k it is worth 1.22x prefill and 1.84x decode on the 122B MoE — which is the opposite of the Vulkan cliff reported in clm-0017. Quantised KV on a STOCK build costs 16% decode at 32k versus f16 and cannot create a context at all without flash attention — which is what clm-0006 predicts for stock, and makes testing its fix the highest-value remaining experiment.
- Qwen3.5-122B-A10B UD-Q4_K_M, aihydra, ROCm 7.1.0, llama.cpp 3653e6d (STOCK — no fork, no patches), `--load-mode none`, 3 reps. Full rows ingested as run-0008..run-0023. | depth | fa | KV | pp512 | tg128 | |---|---|---|---|---| | 0 | 0 | f16 | 312.04 | 21.68 | | 0 | 1 | f16 | 321.72 | 21.91 | | 4096 | 0 | f16 | 283.47 | 19.26 | | 4096 | 1 | f16 | 300.67 | 21.39 | | 32768 | 0 | f16 | 170.81 | **9.87** | | 32768 | 1 | f16 | 208.77 | **18.17** | | 32768 | 1 | q8_0 | 204.49 | 15.26 | **FLASH ATTENTION AT 32k: 1.22x prefill, 1.84x decode.** Without it, decode collapses from 21.68 at empty context to 9.87 at 32k — it loses 54% of its speed just for holding a long conversation. With it, the same span costs only 17% (21.91 → 18.17). ⚑ THIS IS THE OPPOSITE OF clm-0017, AND THAT IS THE POINT. That report — dense 27B, **Vulkan**, stock build — found flash attention CAUSING a prefill collapse at depth, fixed only by an unmerged contiguize patch. Its author explicitly could not test ROCm, having lost it to a Fedora 44 upgrade, and asked for exactly this replication. On the HIP path, on a stock build, **there is no cliff — FA is the thing preventing one.** Scope honestly: this is a hybrid MoE, not the dense 27B they used, so architecture is uncontrolled. A dense-model run is in flight and will settle whether the difference is the backend or the architecture. Either answer is worth reporting back. PRACTICAL: leave `-fa on`. It wins every cell measured here and wins hugely at depth. QUANTISED KV — TWO FINDINGS, BOTH NEGATIVE: · **q8_0 costs 16% decode at 32k** (15.26 vs 18.17). At empty context the two are identical (21.63 vs 21.91), so the penalty is depth-dependent and would be invisible to a shallow benchmark. · **q8_0 with flash attention OFF cannot create a context at all** — llama-bench exits with "failed to create context". Quantised KV *requires* FA on this build. Recorded as a failed cell against its fingerprint rather than retried into a pass (protocol §9). ⚑ CORRECTED 2026-08-08 (operator catch). I first wrote that these results "do not support clm-0006". That is backwards. clm-0006's claim is that **stock llama.cpp dequantizes the KV cache to full precision repeatedly during inference on this silicon**, and that a community FIX then makes q8_0 run 23-53% FASTER than f16. **We are on a stock build.** So measuring q8_0 as *slower* is exactly what clm-0006 predicts for stock — it CONFIRMS the diagnosis and says nothing at all about the fixed build, which we have not tested. That inverts what this result means. Rather than deflating clm-0006, it raises the value of testing the patch: stock costs 16% decode at 32k, the fix claims +23-53%, so the available swing at depth is roughly 40-70%. **That is potentially larger than MTP's 1.45x**, and it compounds with it rather than competing. AND IT IMPLICATES PRODUCTION. `cfg-0002` ran **q8_0/q8_0 KV at 200,000 context** on a fork carrying only the `.ckpt` sidecar fix (`con-0001`) — nothing touching KV. So production sat on the slow dequantisation path, at the deep end where clm-0006 says the penalty is largest. Together with the `--spec-draft-n-max 6` finding above, that is two independent production settings measurably below optimum. q8_0 remains a memory-saving lever on a stock build, not a speed one — but on a patched build that may reverse entirely, and it is now the highest-value untested experiment. SCATTER: every surviving cell reproduced within **1.3% CV**, most under 1%. Per clm-0017 that signature indicates a healthy code path; the broken Vulkan path scattered 5-10%. Our sweep now reports CV per cell precisely so this is checked rather than assumed.
clm-0021
- measured-here high ●●● volatility medium · verified 2026-08-08
- The dense-model flash-attention prefill cliff reported in clm-0017 does NOT exist on ROCm. A stock ROCm build of a dense 27B reaches 214.66 t/s prefill at 32k and 153.63 at 65k — close to their PATCHED Vulkan numbers (237.4 / 180.4) and 5.2x their broken stock Vulkan at 65k (29.8). Flash attention is not the problem on HIP; it is what prevents one, worth 3.4x decode at 65k.
- THIS IS THE REPLICATION clm-0017's AUTHOR ASKED FOR AND COULD NOT RUN. They dropped ROCm because Fedora 44 version locks blocked their release upgrade, leaving them Vulkan-only and unable to answer whether the collapse exists on the HIP path. It does not. Qwen3.6-27B Q4_K_M (dense), aihydra, ROCm 7.1.0, **stock** llama.cpp 3653e6d, f16 KV, `--load-mode none`, 3 reps: | depth | fa | pp512 | tg128 | |---|---|---|---| | 0 | 1 | 355.52 | 12.04 | | 0 | 0 | 353.93 | 11.95 | | 32768 | 1 | **214.66** | 10.86 | | 32768 | 0 | 175.87 | **4.76** | | 65536 | 1 | **153.63** | 9.90 | | 65536 | 0 | 119.08 | **2.92** | SIDE BY SIDE WITH THE COMMUNITY REPORT (their dense 27B, f16 KV, FA on, pp512): | | @32k | @65k | |---|---|---| | Vulkan STOCK (theirs, broken) | 94.5 | 29.8 | | Vulkan PATCHED (theirs) | 237.4 | 180.4 | | **ROCm STOCK (ours)** | **214.66** | **153.63** | Our unpatched ROCm lands near their patched Vulkan and **5.2x their unpatched Vulkan at 65k**. Whatever the contiguize patch repairs, the HIP path does not suffer from it. FLASH ATTENTION IS LOAD-BEARING FOR DECODE, NOT PREFILL. At empty context FA is irrelevant (355.52 vs 353.93 prefill, 12.04 vs 11.95 decode — both within noise). The divergence is entirely depth-driven, and decode is hit far harder than prefill: · prefill at 65k: 1.29x with FA · **decode at 65k: 3.4x with FA** (9.90 vs 2.92) Without FA, decode falls 76% from empty context to 65k (11.95 → 2.92). With it, 18%. A benchmark that only reported prefill would have understated this badly. SCATTER: every cell reproduced within **0.6% CV**. Per clm-0017 the broken path scatters 5-10% while a healthy one holds ~0.3% — our numbers carry the healthy signature, which is independent corroboration that we are not on the defective code path. UNCONTROLLED: clm-0017 says "dense 27B" without naming the model; we used Qwen3.6-27B. If theirs was a different 27B, architecture is not perfectly matched — but the effect size (5.2x) is far larger than any plausible model-to-model difference. CONSEQUENCE, COMBINED WITH clm-0020: leave `-fa on` everywhere on this hardware. It wins at every depth measured, on both a dense 27B and a 122B hybrid MoE, and its absence is catastrophic for decode at depth. The clm-0017 caution applies to Vulkan only. WORTH REPORTING BACK. The author explicitly requested this test. Answering it costs us nothing and it resolves whether their patch is a general fix or a Vulkan-backend repair — which changes whether it belongs upstream as a backend fix or a cross-cutting one.
clm-0022
- measured-here high ●●● volatility medium · verified 2026-08-08
- The community KV-dequantisation fix is real, large, and scales monotonically with depth: one cherry-picked commit recovers +18.3% at 32k, +55.6% at 131k and **+70.3% at 204,800 — production's own context** — while leaving f16 unchanged at every depth. Warden's local model was therefore running at **59% of its achievable decode speed** at the context it actually used.
- METHOD — single-variable attribution. Both binaries built from the SAME upstream commit (3653e6d); the only difference is one cherry-picked patch, `ce7689f` = Nathanw1014's `2a24abc` "CUDA: dequantize KV on load in the tile FA kernel, use it for quantized decode" (6 files, +201/-45). Building the fork wholesale was DELIBERATELY avoided: it sits on upstream #25xxx while our tree is 2026-08-07, which would have confounded the patch with the base version — the attribution error clm-0017's author avoided with five same-base builds. Qwen3.5-122B-A10B UD-Q4_K_M, ROCm 7.1.0, `-fa 1`, `--load-mode none`, `--parallel 1`, 3 reps, page cache dropped between arms. | build | KV | depth | pp512 | tg128 | |---|---|---|---|---| | stock | f16 | 0 | 326.36 | 21.90 | | kvfix | f16 | 0 | 319.24 | 21.88 | | stock | f16 | 32768 | 210.21 | **18.16** | | kvfix | f16 | 32768 | 204.48 | **18.16** | | stock | q8_0 | 32768 | 200.26 | **15.29** | | kvfix | q8_0 | 32768 | 211.35 | **18.09** | **THE CONTROL IS THE POINT.** f16 decode at 32k is 18.16 on BOTH builds — identical to four significant figures. The patch touches the tile flash-attention kernel, so without an f16 arm a general FA speedup could have been misread as a quantised-KV win. It moved q8_0 by 18.3% and f16 by nothing, which is exactly what a correct KV-specific fix looks like. q8_0 prefill also gains (+5.5%); at depth 0 nothing moves, so the effect is depth-dependent as the mechanism predicts. WHAT IT CHANGES OPERATIONALLY: on stock, q8_0 KV costs 16% decode at 32k and **35% at 131k** — you pay steeply increasing speed for memory, exactly where long context is the point. **With the patch that cost disappears entirely** and slightly reverses. So quantised KV becomes free: roughly half the KV footprint at no speed penalty at any depth measured. On `cfg-0002`'s 200,000-token configuration that is ~12 GB of headroom recovered for nothing, which on a fixed-memory box buys context or a second resident model — and it is the concrete enabler for the 35B + 122B co-residency in clm-0023. ⚑ DEPTH TEST RUN — AND IT CHANGES THE CONCLUSION. My first write-up said the fix reached "parity, not superiority" and that clm-0006's "23-53% faster than f16" did not reproduce. That was an artefact of testing at only 32k. Decode at **131,072**: | build | KV | tg128 | |---|---|---| | stock | f16 | 12.03 | | stock | q8_0 | **7.81** | | kvfix | f16 | 12.04 (control — unchanged) | | kvfix | q8_0 | **12.15** | **+55.6% from the patch at 131k**, and q8_0 now sits marginally ABOVE f16 — which is clm-0006's claim, reproduced. The gain scales monotonically with depth exactly as the mechanism implies (dequantize once on load versus repeatedly during inference: the more cache you touch per token, the more you save): | depth | patch gain on q8_0 | |---|---| | 0 | none | | 32,768 | +18.3% | | 131,072 | +55.6% | | **204,800** | **+70.3%** | AT PRODUCTION DEPTH (204,800 — `cfg-0002` ran n_ctx 200,000): | build | KV | tg128 | |---|---|---| | stock | q8_0 | **5.69** | | kvfix | q8_0 | **9.69** | | stock | f16 | 9.60 | | kvfix | f16 | 9.54 (control — unchanged, -0.6% is noise) | The control holds at every depth tested: the patch moves q8_0 and leaves f16 alone. At 204,800 the patched q8_0 (9.69) now **exceeds** f16 (9.54-9.60), which is clm-0006's claim in full. Their +75% to +203% was measured at 262k, deeper than anything we ran, so our +70.3% sits just below their range on a shallower test — consistent, not conflicting. ALSO ESTABLISHED, AND WE DID NOT HAVE IT BEFORE: **f16 KV FITS at 204,800.** 9.60 t/s with memory stable at 79 GiB and no swap growth. The f16 arm was expected to be the one that might OOM against the 104 GiB ceiling; it did not. That is a real fit data point of the kind clm-0001 says we lack. 🔴 **MEASURED AT PRODUCTION'S OWN CONTEXT, AND IT IS WORSE THAN THE EXTRAPOLATION.** `cfg-0002` ran **q8_0/q8_0 KV at n_ctx 200,000** on a fork carrying only the `.ckpt` sidecar — nothing touching KV. At 204,800 that configuration produces **5.69 tok/s where 9.69 was available**: Warden's local model was running at **59% of its achievable decode speed** at the context it actually used. On a 500-token reply that is 88 seconds instead of 52. Two independent, compounding mistunings are now measured in the shipped configuration: this, and `--spec-draft-n-max 6` at ~18% below the optimum of 2-3. Neither was visible without an A/B, and neither would have shown up in any throughput number taken alone — the config simply looked like the speed of the hardware. FORK COST, NOW CONCRETE: adopting this means a SECOND divergence alongside `con-0001` (the `.ckpt` sidecar). They touch different files and both would have to be carried together. `build_commit` is a fingerprint field, so a combined tree is a configuration nobody else can reproduce — the tax is real. But a **55.6% decode recovery at 131k**, rising with depth, plus ~12 GB of headroom, on a box whose entire purpose is long context, makes this the strongest case for carrying a fork that we have. SCATTER: every cell ≤0.8% CV. Healthy code paths on both builds.
clm-0023
- measured-here high ●●● volatility medium · verified 2026-08-08
- Qwen3.6-35B-A3B is 2.3x the 122B's decode on identical hardware (51.01 vs 21.90 tok/s at empty context) and holds 42.45 at 32k, with prefill above 1000 tok/s. It passes the capability guard 4/4.
- QWEN3.6-35B-A3B (UD-Q4_K_XL, 21 GB), aihydra, ROCm 7.1.0, stock llama.cpp 3653e6d, `-fa 1`, f16 KV, `--parallel 1`, 3 reps. Ingested as cfg-0012 / run-0030..0032. | depth | pp512 | tg128 | |---|---|---| | 0 | **1079.38** | **51.01** | | 4096 | 965.37 | 49.80 | | 32768 | 598.68 | 42.45 | All cells ≤0.6% CV. Guard 4/4 at depth 8000. THE ACTIVE-PARAMETER PREDICTION, TESTED. Decode is bandwidth-bound, so throughput should scale roughly inversely with ACTIVE parameters: A3B (3B) against A10B (10B) predicts ~3.3x. We measure **2.3x** (51.01 vs 21.90). The direction and rough magnitude hold, and the shortfall is expected — per-token overhead does not scale down with active parameters. Good enough to keep using active-parameter count as a first-order predictor, not good enough to quote as a rule. It also retains speed at depth better than the 122B: 51.01 → 42.45 is a 17% loss to 32k, against the 122B's 21.90 → 18.16, also 17%. Proportionally identical, which suggests the depth penalty is an attention-cost property rather than a model-size one. REFLEX-TIER VERDICT: at 42-51 tok/s with sub-second prefill on 21 GB, this is a credible fast lane. Combined with clm-0022's finding that quantised KV is free on a patched build, a 35B and the 122B could plausibly be co-resident within the 105 GiB ceiling — which is the council architecture's first concrete opportunity on this hardware. ⚑ GRAMMAR CEILING: see clm-0024, which re-measures this properly and finds the ceiling does not reproduce — the constraint is context size, not the grammar builder. Kept below as the process record of how the wrong number nearly got published. NOT MEASURED HERE, AND THE FIRST ANSWER WAS WRONG. The probe reported "highest OK: 30 tools (127,415 B)" against the real 353 KB Home Assistant schemas. That is **not** the grammar ceiling. The 400 carried `{"type":"exceed_context_size_error", "n_prompt_tokens":16651, "n_ctx":16384}` — the server was started at `-c 16384` and 31 real tools need ~16.6k tokens. We measured the CONTEXT limit and nearly published it as a grammar limit, which would have understated the real ceiling badly and sent us hunting a regression that does not exist. The script's own `grammar-related error text: no` line caught it — and then printed a ceiling anyway. **A caveat nobody acts on is not a safeguard.** Fixed 2026-08-08: the probe now exits 2 and refuses to report any number when the failure is a context overflow, naming the token count it needed. Re-running at `-c 131072`. The real ceiling is somewhere above 30 tools and currently unknown. For scale: at ~4 KB per real MCP tool, probing 120 tools needs roughly 60k tokens of context — which is itself a useful fact, because production runs nowhere near that much context devoted to tool schemas. ALSO PASSED TONIGHT: tau2-bench smoke test. LiteLLM reaches our endpoint and gets a correct native tool call (`get_booking {"reference":"ABC123"}`), so the adopted agentic-benchmark path works end to end and is no longer an untested assumption.
clm-0024
- measured-here med ●●○ volatility high · verified 2026-08-08
- The tool-grammar ceiling does not reproduce on llama.cpp 3653e6d. Given adequate context, 200 real MCP tools totalling 752 KB were accepted with HTTP 200. The binding constraint is CONTEXT SIZE — tool schemas consume prompt tokens — not the GBNF grammar builder. This supersedes the ~55-59 tool figure that has shaped our MCP tool budget since 2026-07-05.
- Qwen3.5-122B-A10B UD-Q4_K_M, aihydra, ROCm 7.1.0, stock llama.cpp 3653e6d, `--jinja`, `--parallel 1`, real Home Assistant MCP schemas (78 tools, 353 KB, cycled to reach higher counts). Fresh server, `-c 131072`: | tools | request bytes | result | |---|---|---| | 60 | 226,121 | HTTP 200 | | 80 | 304,532 | HTTP 200 | | 100 | 367,660 | HTTP 200 | | 120 | 464,680 | HTTP 200 | | 160 | 593,137 | HTTP 200 | | **200** | **752,383** | **HTTP 200** | No failure was found at any count tested. WHAT THE OLD CEILING ACTUALLY WAS. Tonight's first probe reported "30 tools" — that was the server at `-c 16384` returning `exceed_context_size_error`, pure context exhaustion. A second probe at `-c 131072` reported 59/60, and a fresh server then accepted 60 AND 61. Only the third run — fresh server, ladder rather than binary search — established that nothing fails up to 200. The consistent story across all three is **context**: 200 tools is ~94k tokens of schema, which fits in 131k and cannot fit in 16k. So the mental model changes. It was "llama.cpp cannot build a grammar above ~55-59 tools by schema SIZE". It should be "tool schemas are prompt tokens, and you need context to hold them". That is a far more tractable constraint — it trades against context budget rather than being a hard wall. ⚠ CONFIDENCE IS medium, NOT high, AND VOLATILITY IS high. Three reasons to hold this loosely: · **One unexplained failure.** n=60 genuinely returned 400 during the binary search at `-c 131072`, then passed twice on fresh servers. Something is state-dependent and we have not characterised it — the same flavour of instability as clm-0019. · **We only proved acceptance, not correctness.** HTTP 200 means the grammar compiled and inference started. It does NOT mean the model selects correctly among 200 tools. Tool-selection accuracy at high tool counts is a CAPABILITY question and completely untested here. · **Build-specific by construction.** The ceiling is a property of the build; ours is one day old. It may differ on the version production runs. WHAT THIS WOULD CHANGE IF IT HOLDS (do not act before re-verifying on the production build): the ~55-59 figure has directly shaped Warden's configuration — the Home Assistant `toolFilter` cut 42 tools to 9, apple-mail was curated, and the live count sits around 51 deliberately close to the limit. If the wall is really context, that curation buys tokens rather than avoiding a cliff, and the trade can be reasoned about instead of feared. The **silent cloud fallback** risk (HTTP 400 → OpenRouter, looking healthy) is the thing that made the ceiling frightening; a context error is at least loud. NEXT: re-run on the production llama.cpp build, and — more importantly — measure whether tool-SELECTION accuracy degrades as tool count rises. Acceptance was never the interesting question; it was just the one that used to fail first.
clm-0025
- measured-here high ●●● volatility medium · verified 2026-08-08
- Four models measured on identical hardware give decode from 17.51 to 55.45 tok/s, and the bandwidth model predicts the ORDER but not the magnitude — realised efficiency ranges from 34% to 62% of the theoretical ceiling. Separately, gpt-oss-120b returns STALE ANSWERS FROM PREVIOUS REQUESTS at --parallel 1, which no other model tested does.
- All on aihydra, ROCm 7.1.0, stock llama.cpp 3653e6d, `-fa 1`, f16 KV, `--parallel 1`, `--load-mode none`, 3 reps, page cache dropped between models. | model | active | quant | pp512 @0 | tg128 @0 | tg128 @32k | |---|---|---|---|---|---| | gpt-oss-120b | ~5.1B | UD-Q4_K_XL | 450.83 | **55.45** | 41.70 | | Qwen3.6-35B-A3B | 3B | UD-Q4_K_XL | 1079.38 | 51.01 | 42.45 | | Qwen3.5-122B-A10B | 10B | UD-Q4_K_M | 326.36 | 21.90 | 18.16 | | Nemotron-3-Super-120B-A12B | 12B | UD-IQ4_XS | 225.36 | **17.51** | 17.16 | THE BANDWIDTH MODEL, TESTED PROPERLY FOR THE FIRST TIME. Decode should scale inversely with ACTIVE parameters at ~256 GB/s. Taking ~0.56 bytes/param at 4-bit: | model | active | predicted ceiling | measured | realised | |---|---|---|---|---| | Qwen3.6-35B-A3B | 3B | ~152 t/s | 51.01 | **34%** | | gpt-oss-120b | 5.1B | ~90 t/s | 55.45 | **62%** | | Qwen3.5-122B-A10B | 10B | ~46 t/s | 21.90 | **48%** | | Nemotron-120B-A12B | 12B | ~38 t/s | 17.51 | **46%** | **The ORDER is right — more active parameters, slower decode, monotonically. The MAGNITUDE is not.** Realised efficiency varies nearly two-fold, and the smallest model is the least efficient: the 35B-A3B converts only 34% of its theoretical bandwidth into tokens against gpt-oss's 62%. Per-token overhead does not shrink with active parameters, so a very sparse model spends proportionally more time on everything that is not weight streaming. This refines clm-0023, which took 2.3x on a single pair as broad support for the model. With four points the honest statement is: **active-parameter count predicts ranking reliably and throughput poorly.** Useful for choosing which model to try; useless for predicting what it will do. ⚠ gpt-oss-120b RETURNS STALE ANSWERS — AND ONLY gpt-oss DOES. Its guard failed retrieval in a way none of the others did: · On a **virgin server**, needle at 8000 tokens returns `'chartreuse'` — the first word of `chartreuse-viper-88`. Retrieval WORKS; the model simply answers partially. That is a probe-strictness issue on our side, not a model failure. · **After any prior request**, the same probe returns `'the quick brown fox'` — verbatim the answer to the COHERENCE check that ran earlier. Reproduced 6/6 across two sequences. The server hands back a previous response. · Control questions are unaffected: "2+2" → `4`, "capital of France" → `Paris`. This is at `--parallel 1`, so it is NOT clm-0019 (multi-slot) and NOT llama.cpp #25992 (concurrent leakage). ⚑ **PROMPT CACHING RULED OUT, 2026-08-08.** I predicted `cache_prompt` prefix reuse. Tested both ways with a priming request in between: | cache_prompt | priming request | needle | needle | |---|---|---|---| | true | 'the quick brown fox' | 'the quick brown fox' | 'the quick brown fox' | | **false** | 'the quick brown fox' | **'the quick brown fox'** | **'the quick brown fox'** | Disabling the prompt cache changes nothing. The model returns the previous answer regardless. Combined with the virgin-server result (first request after start answers correctly with 'chartreuse'), the pattern is: **the FIRST request on a fresh server is correct and every subsequent request returns the first one's answer.** That is slot state not being reset between requests, not a caching optimisation misfiring. Mechanism still unidentified. gpt-oss uses the harmony chat format, which llama.cpp handles through a separate code path, and that remains the most likely locus — but it is a hypothesis, not a finding. Reproducible in three lines against a fresh server, so it is cheap for anyone to confirm and worth an upstream report once characterised. ADDITIONAL, from the 2026-08-08 matrix run: · **gpt-oss REQUIRES flash attention.** Both `fa=0` arms failed outright (f16 and q8_0), where the 122B only failed the q8_0/fa=0 combination. · **q8_0 costs 56% decode at 131k** (10.70 vs f16's 24.53) — substantially worse than the 122B's 35% at the same depth, on the same stock build. If the KV-dequant patch (clm-0022) helps proportionally, gpt-oss has more to gain from it than anything else measured. · f16 at 131k flagged **3.5% CV**, above our 3% scatter threshold. Per clm-0017 that is a defective-path signature and warrants a look rather than a shrug. CONSEQUENCE: **gpt-oss-120b's throughput rows are recorded but NOT guard-cleared.** They are honest numbers for how fast it generates and say nothing about whether it generates the right thing. On this evidence it is not a candidate for any Warden role until the staleness is understood — a model that occasionally serves a previous answer is far worse than a slow one. Nemotron-3-Super passed its guard 4/4 and is the slowest of the four at 17.51 tok/s, which its 12B active parameters predict. It is also the flattest with depth (17.51 → 17.16, a 2% loss to 32k, against 17-20% for the others) — worth a second look if long-context stability ever matters more than raw speed.
clm-0026
- measured-here high ●●● volatility medium · verified 2026-08-08
- Speculation is close to worthless on Qwen3.6-35B-A3B with varied prompts — ngram-mod gives 1.11x with a 29% coefficient of variation, ngram-cache gives nothing, and MTP is unavailable because the model carries no NextN layers. The community's 121 tok/s is not a single-stream figure we can chase with configuration: their no-speculation FLOOR alone is 27% above ours, which points at the quant format, not at tuning.
- Qwen3.6-35B-A3B UD-Q4_K_XL, aihydra, ROCm 7.1.0, stock llama.cpp 3653e6d, `-fa on`, f16 KV, `--parallel 1`, ctx 16384, 3 reps x 5 varied prompts (n=15 per arm). | spec type | tok/s | stddev | ratio | |---|---|---|---| | none | **50.91** | 0.10 | 1.00x | | ngram-mod | 56.42 | **16.53** | 1.11x | | ngram-cache | 50.22 | 1.36 | 0.99x | | draft-mtp | FAILED | | | The floor agrees with llama-bench's independent 51.01 to 0.2%, so both harnesses are measuring the same thing. **draft-mtp failed cleanly and for a good reason:** `context type MTP requested but model doesn't contain MTP layers`. Unlike Qwen3.5-122B-A10B-MTP, this build of the 35B carries no NextN heads, so the 1.45x that MTP delivers on the 122B is simply unavailable here. A recorded failure against its fingerprint, not a retry candidate. **ngram-mod's 29% CV is the finding, not its 1.11x mean.** On the 122B it was 1.17x with 38% CV; here 1.11x with 29%. Draft-free speculation only pays when the output repeats text the n-gram pool has already seen, so on five deliberately dissimilar prompts some runs gain substantially and others gain nothing. The mean is not a number to plan with. This QUALIFIES clm-0016 rather than contradicting it: their 1.87x came from 500 IFEval prompts, which are short, structured and highly repetitive — the best case for the technique. On open-ended generation it largely evaporates. WHY WE WILL NOT REACH THEIR 121 tok/s BY CONFIGURATION. Decomposing the gap honestly: · **Their no-speculation floor is 64.8; ours is 50.91 — 27% apart.** A floor difference cannot be caused by speculation, parallelism or tuning. The remaining structural difference is the quant: they ran **ROCmFP4**, we ran UD-Q4_K_XL. Quant format sets bytes-per-weight and therefore the bandwidth ceiling itself, which is the one term that moves a floor. · **Their 121 is FOUR STREAMS POOLED, not single-stream.** Their own post labels the sustained figure as 4 slots across 500 prompts. Our 50.91 is one stream. Comparing the two directly overstates the gap; aggregate throughput and single-stream latency are different quantities and should never be set side by side. · Their trustworthy single-stream peak was 380 tok/s, but peak in a real agentic session is not a sustained rate either. So the honest reading: **ROCmFP4 is the lever worth chasing on this model, and it is worth roughly 27%.** Speculation is not — on varied prompts it buys ~11% with scatter three times larger than the gain. ⚠ AND WE CANNOT SIMPLY COPY THEIR PARALLEL-4 SETUP. clm-0019 measured that `--parallel 4` destroys long-context retrieval on this build (needle 0/8 against 8/8 at `--parallel 1`). Their 121 tok/s was measured at parallel 4. Either their build does not carry that defect, or their IFEval prompts were short enough never to expose it. Chasing their number by raising slot count would trade a capability we have verified for throughput we have not. NEXT, IF THE 35B MATTERS AS A REFLEX TIER: build the ROCmFPX fork and requantise. That is a third divergence to carry, on top of `con-0001` and the KV-dequant patch (clm-0022) — but unlike speculation it moves the floor, and the floor is what a reflex tier is for. ⚑ **CLM-0010 IS UNRECONCILED WITH THIS.** That report claims draft-free ngram alone takes a single stream from 71 to 216 tok/s — a 3x ratio, against our measured 1.11x on the same technique here. Neither their ROCmFPX-fork ROCmFP4 finetune nor their prompt regime (repetitive code-edit content, ngram-mod's best case) matches ours, and either could account for most of the gap — but it has not been tested, so the discrepancy stands rather than being averaged away.
clm-0027
- measured-here high ●●● volatility low · verified 2026-08-08
- Prefix cache reuse is worth 9.8x on this box — an 8,000-token prefix costs 25.99 s cold and 2.65 s warm — and it is strictly PREFIX-ANCHORED: prepending three characters to an otherwise identical prompt returns it to full cold cost (26.26 s), zero reuse despite 99.9% identical content. Separately, both f16 and q8_0 KV load at every context up to 204,800, and the f16/q8_0 footprint difference is only 2 GiB — far below the ~12 GiB cfg-0002 assumed.
- Qwen3.5-122B-A10B UD-Q4_K_M, aihydra, ROCm 7.1.0, stock llama.cpp 3653e6d, `-fa on`, f16 KV, `-c 32768`, `--parallel 1`, `--load-mode none`. ## Cache behaviour — the ladder nobody publishes | probe | wall | prefill | meaning | |---|---|---|---| | cold | **25.99 s** | 317.5 t/s | first sight of this prefix | | warm | **2.65 s** | 269.2 t/s | same prefix, different question | | diverged | **26.26 s** | 315.9 t/s | three characters prepended | **9.8x from prefix reuse.** That is the whole warm-lane thesis, measured on this hardware rather than inferred: the difference between a turn that feels instant and one that does not is almost entirely whether the prefix was seen before. **And reuse is prefix-anchored, not similarity-based.** The `diverged` probe is 99.9% identical to `cold` — the same 8,000 tokens, with `ZZZ ` prepended. It gets **no reuse whatsoever** and costs 101% of cold. A changed head invalidates everything after it. That is the measured justification for design decisions already taken on faith: · why the prefix-relocation fix (`cache-maxing-prefix-fix`, upstreamed as openclaw #98267) mattered — moving two sections below the cache boundary took shared tokens from 1.46K to 15.5K, and this shows what each shared token is worth; · why heartbeats and crons that carry their own preamble **cannot** share a warm lane with iMessage traffic, and why slot pinning was the right answer rather than a bigger cache; · why anything that mutates the head of a prompt — a timestamp, a rotating greeting, a changing tool list — is far more expensive than its size suggests. GAP: this build does not populate `prompt_n_cached` in `timings`, so the ratio is measured by wall clock rather than reported token counts. The effect is large enough (9.8x) that this does not threaten the conclusion, but a cached-token figure would let us measure PARTIAL reuse rather than just its presence or absence. ## Fit — both KV arms load everywhere | KV | ctx | loaded | load_s | peak GiB | |---|---|---|---|---| | f16 | 32,768 | yes | 46 | 75 | | f16 | 131,072 | yes | 78 | 77 | | f16 | 204,800 | yes | 48 | **79** | | q8_0 | 32,768 | yes | 95 | 75 | | q8_0 | 131,072 | yes | 62 | 76 | | q8_0 | 204,800 | yes | 47 | **77** | **No fit cliff anywhere.** f16 KV reaches 204,800 with 79 GiB peak against the 104 GiB GTT ceiling — 25 GiB spare. The arms are therefore comparable at the top length, which protocol §9 warns is not to be assumed. ⚑ **THE f16/q8_0 DIFFERENCE IS 2 GiB, NOT 12.** `cfg-0002` records `kv_gb: 12.4` for q8_0 at 200,000 context, implying ~25 GiB for f16. The measurement says otherwise: 79 vs 77 GiB peak. The likely explanation is architecture — this is a hybrid model, so only a minority of layers hold a full attention cache, exactly as clm-0015 found for Ling-3.0-flash (7 of 42 layers). **If so, cfg-0002's memory record is a substantial overestimate and the fit calculus behind several decisions was too conservative.** It also means quantised KV saves far less memory than assumed — which, combined with clm-0022 showing it costs nothing in speed once patched, makes the whole q8_0-vs-f16 question much less consequential than it looked. Both figures come from `capability-probe.sh`, which only produced them after being given `--load-mode none`; before that both arms timed out at 673 s and were recorded as non-loading. A probe that cannot load the model reports a fit cliff that does not exist.
clm-0028
- measured-here high ●●● volatility low · verified 2026-08-08
- Across four models on gfx1151, flash attention is worth 2.5x to 4.8x DECODE at 131k and its absence is catastrophic — a 35B loses 88% of its decode speed from empty context to 131k without it, against 44% with it. Separately, quantised KV **requires** flash attention: the q8_0 + fa=0 cell failed on 4 of 4 models. Neither fact is visible at shallow depth, where flash attention is irrelevant.
- aihydra, ROCm 7.1.0, stock llama.cpp 3653e6d, `--parallel 1`, `--load-mode none`, 3 reps, page cache dropped between cells. Depths 0 / 32,768 / 131,072. Ingested as run-0050..run-0097. ## Flash attention: decode at 131,072 (f16 KV) | model | fa=1 | fa=0 | gain | |---|---|---|---| | Qwen3.6-35B-A3B | **28.46** | 5.91 | **4.8x** | | Nemotron-3-Super-120B-A12B | **16.21** | 6.41 | **2.5x** | | Qwen3.5-122B-A10B (at 32k) | 18.17 | 9.87 | 1.84x | | gpt-oss-120b | 24.53 | **failed to run at all** | n/a | **The absence of flash attention is not a slowdown, it is a collapse.** Qwen3.6-35B-A3B goes 50.88 → 5.91 tok/s from empty context to 131k without it, an 88% loss. With it, 51.12 → 28.46, a 44% loss. Nemotron: 17.46 → 6.41 without (−63%) against 17.51 → 16.21 with (−7%). **And it is invisible at shallow depth.** At depth 0 every model is within 1% either way (35B: 51.12 vs 50.88; Nemotron: 17.51 vs 17.46). A benchmark that only measures empty context would conclude flash attention does not matter. It is the single most important setting on this hardware and only depth reveals it. ## Quantised KV requires flash attention — 4 of 4 models Every `q8_0 + fa=0` cell failed outright: Qwen3.5-122B, Qwen3.6-35B, Nemotron, gpt-oss. llama-bench exits with "failed to create context" before any inference. This is not a performance finding but a hard constraint, and it means the two levers are not independent — you cannot sweep them as a clean 2x2. gpt-oss is stricter still: **both** its fa=0 arms failed, f16 as well as q8_0. It cannot run without flash attention at all. ## q8_0 penalty on a STOCK build varies by model Decode at 131k, fa=1, q8_0 against f16: | model | f16 | q8_0 | penalty | |---|---|---|---| | gpt-oss-120b | 24.53 | 10.70 | **−56%** | | Qwen3.6-35B-A3B | 28.46 | 18.48 | −35% | | Qwen3.5-122B-A10B | 12.03 | 7.81 | −35% | | Nemotron-3-Super | 16.21 | 12.58 | −22% | All four pay a penalty, spanning 22% to 56%. This is the defect clm-0006 describes and clm-0022 measured a fix for (+70.3% recovery at production depth on the 122B). **The patch has not been tested on the other three**, but gpt-oss's −56% suggests it has the most to gain of anything measured. ## Guards Qwen3.6-35B-A3B and Nemotron-3-Super both pass 4/4 — coherence, tool call with real arguments, needle recovered at 8,000 tokens. gpt-oss fails retrieval with the stale-answer defect (clm-0025) and its rows are recorded but not capability-cleared. ## Practical **Leave `-fa on` everywhere on this hardware, unconditionally.** It wins or ties at every depth on every model tested, it is the difference between a 44% and an 88% decode loss at 131k, and three of four models cannot use quantised KV without it.
clm-0029
- measured-here high ●●● volatility high · verified 2026-08-08
- DIAGNOSED: `llama-perplexity` produces garbage on this build — PPL 532 on deliberately repetitive English that should score 2-5, 3,654 on wikitext for a model that generates coherently at 51 tok/s and passes tool-calling guards. Not flash attention, not the corpus, not context size. **The KL-divergence route to measuring KV-quant QUALITY is therefore blocked**, and clm-0022's open half needs a different instrument.
- Qwen3.5-122B-A10B UD-Q4_K_M, aihydra, ROCm 7.1.0, stock llama.cpp 3653e6d, `-fa 1`, `-c 32768`, corpus `wiki.test.raw` sha256 `430983e5…` (bench/CORPUS.md), q8_0 arm against an f16 `--kl-divergence-base`. **Core finding, by elimination across four controls (model, corpus, context, flash attention):** the defect is in `llama-perplexity` itself on this build (3653e6d) with these models — not the corpus (a 4-sentence repeating probe that should score 2-5 scored 532), not flash attention (disabling it moved 17,180 to 3,654, still absurd), not context size (4096 gave 56.2, no better than 32768), and not the model (the same 35B generates coherently and passes its tool-calling guard 4/4). Plausibly MoE or hybrid-architecture handling in the KL path, or the quant format — not established. **Consequence:** `kv-quality.sh`, built on `llama-perplexity --kl-divergence`, is unusable on this build, so the KV-quant QUALITY question could not be answered this way. It was answered instead via τ²-bench task completion (`clm-0035`, `clm-0038`). `clm-0022` covers what quantised KV costs in SPEED, which this instrument failure never touched. Full diagnostic method (the four controls, the two reasons the original KL numbers were implausible, and the alternatives considered) is written up as a standing methodology rule in `docs/methodology-lessons.md` §2 — read that before repeating this measurement.
clm-0030 superseded
- measured-here low ●○○ volatility medium · verified 2026-08-09
- SUPERSEDED: the Pass^1 = 1.000 reported by this run came from a 3-task subsample biased toward the domain's easiest tasks — the other two of the original five never terminated and were excluded as infrastructure errors. The sustained score across a realistic sample is clm-0037's 0.545 (n=22), which is the number to cite for this model on tau2 airline. This run was nonetheless the project's first genuine capability measurement rather than a throughput number, and it proved the harness and scoring path work.
- τ²-bench v1.0.1, airline domain (14 tools), aihydra, ROCm 7.1.0, stock llama.cpp 3653e6d, Qwen3.5-122B-A10B UD-Q4_K_M, `-fa on`, f16 KV, `-c 32768`, `--parallel 1`. ## Pins — changing ANY of these re-baselines the series (protocol §8) | pin | value | |---|---| | agent LLM | the local 122B | | user simulator | **the same local 122B** | | judge | **none — deterministic scoring only** | | domain | airline, 14 tools | | tasks requested | 5 | No external judge was available offline, so scoring rests on tau2's deterministic components: database state-hash comparison and tool-call trajectory match. That is the more objective half of tau2's scoring, but it is not the whole of it. ## Result | metric | value | |---|---| | **Pass^1** | **1.000** | | Read actions | 5/5 (100%) | | Write actions | **none exercised** | | DB match | ✓3 / ✗0 (100%) | | Termination | 3 normal, all user-initiated | | Infra errors | **2** | | task | reward | turns | duration | |---|---|---|---| | 0 | 1.0 | 16 | 31 min | | 1 | 1.0 | 22 | 36 min | | 2 | 1.0 | 21 | 44 min | | 3 | — | 0 | infra error | | 4 | — | 0 | infra error | The model held 16-22 turn conversations with correct tool use throughout and reached the right database state every time. Against our own stated failure mode — tool selection — that is the first direct evidence in either direction, and it is positive. ## What this does NOT establish · **Three tasks is a small sample.** Pass^1 = 1.000 over n=3 is consistent with a true pass rate anywhere from roughly 0.3 upward. It is a floor, not a score. · **Read paths only.** Write actions were never exercised — every completed task was a lookup. Mutation is where an agent can do damage, and it is untested. · **The two missing tasks are now explained, and it matters (clm-0032).** They were `infrastructure_error` in this run because tau2 defaults to THREE concurrent simulations against our deliberately single-slot server. But re-run serially they do not fail — they **never terminate**, running 4h15m and 2h10m without producing a result. They are precisely the two tasks whose user is scripted to escalate indefinitely, and on those the agent argues instead of emitting `###TRANSFER###`. **So Pass^1 = 1.000 is computed over exactly the subset the model handles well**, and the two excluded tasks are the ones testing the hardest behaviour. The headline is real but it is not the whole picture. · **Self-play.** The user simulator is the SAME model as the agent. That is tau2's default when no second model is available, but a model conversing with itself may be an easier interlocutor than a different one would be. · **Airline is 14 tools**, far below the tool counts Warden actually runs (~51) and far below where grammar or context limits bite (clm-0024). ## Cost, which matters for planning **31-44 minutes per task** on this model, and the earlier clipped run measured 110 minutes per task while sharing the box. A 5-task domain is a 3-4 hour job; the full airline suite would be far longer. Any future capability series needs that budgeted honestly rather than guessed — my first attempt at this run was killed by a 2-hour timeout I set without measuring a single task first.
clm-0031
- community low ●○○ volatility high · verified 2026-08-09
- A community DeepSeek-V4-Flash report independently confirms our mmap/GTT double-residency finding, demonstrates a 120 GiB GTT ceiling in production use, and — most consequentially — reports Vulkan BEATING ROCm on 3 of 4 cells including 55% faster decode at depth. It also documents a GPU ring-timeout failure mode we have never hit and were not watching for.
- Source: a single operator (Reddit, u/Neuromacmd) on an ASUS PX13 — Strix Halo / Radeon 8060S, gfx1151, 128 GB unified, Fedora 44, kernel 7.1.5, Mesa 26.1.5. Model `unsloth/DeepSeek-V4-Flash-0731-GGUF` UD-IQ3_XXS, 97.05 GiB. Same silicon as aihydra, so it transfers more directly than most community reports. ## 1. INDEPENDENT CONFIRMATION of the double-residency hazard Their words: *"`-mmp 0` is not optional at this size on my box: with mmap the page cache gets counted against GTT on top of the weights and it never loads."* That is precisely the defect that bit us four separate times before being fixed properly (protocol §0, `LLAMA_ARG_LOAD_MODE=none`). A second operator hitting it independently, on the same silicon, confirms it is a property of unified memory rather than anything specific to our setup — and that the fix belongs at the environment level, not per script. ## 2. A 120 GiB GTT ceiling works in production They run `ttm.pages_limit=31457280` (**120 GiB**) with a **512 MB** VRAM carve-out, and load a 97 GiB model to a 99 GiB footprint. We run 104 GiB GTT with a 1 GiB carve-out. So our 104 GiB was conservative — deliberately, as a guardrail against the aibeast unkillable-OOM class (clm-0022 notes the reasoning). This is evidence the headroom above it is usable if a model ever needs it. It does not change the guardrail argument; it removes the worry that 120 GiB is unstable. ## 3. ⚑ VULKAN AHEAD OF ROCm ON 3 OF 4 — which contradicts our working assumption | test | Vulkan | ROCm | |---|---|---| | shallow pp2048 | **124.16** | 114.37 | | shallow tg64 | **18.15** | 13.27 | | d24576 pp2048 | 63.67 | **74.31** | | d24576 tg64 | **14.36** | 9.25 | **Decode at depth: Vulkan 14.36 against ROCm 9.25 — 55% faster.** We have been treating HIP as the default on the strength of two garbage-output reports (clm-0014) and clm-0021's finding that the FA prefill cliff is Vulkan-only. This says the picture is more mixed. ⚠ BUT IT IS NOT A CLEAN COMPARISON, and the author does not claim it is. The Vulkan side is revision `4a1fb6c` (build 867) carrying **three of their own unsubmitted patches**; the ROCm side is `cd0fa60` (build 824). **Different commits, different patch sets.** That is exactly the confound clm-0017's author avoided with five same-base builds and that we avoided by cherry-picking a single commit for clm-0022. Their Vulkan advantage may be the backend, their patches, or the 43-build gap. ⚠ CONFIDENCE LOWERED TO LOW ON THIS BASIS: a 43-build gap between the two revisions tested, compounded by three of the author's own unsubmitted patches present on the Vulkan side only, is a structural confound severe enough that the Vulkan-ahead finding cannot be attributed to the backend until it is reproduced here on matched commits. ACTION: this is worth testing ourselves and we are unusually well placed — we have both backends, a controlled harness, and the discipline to build both from one base. **A clean Vulkan-vs-HIP comparison on identical commits is now the highest-value untested backend question**, and it was already on the coverage matrix as not-started. ## 4. A failure mode we were not watching for *"Slow enough at depth that the 10s compute-ring watchdog fires and the driver kills the context."* With dmesg signature: amdgpu: ring comp_1.1.1 timeout, signaled seq=7906, emitted seq=7907 amdgpu: Starting comp_1.1.1 ring reset amdgpu: [drm] device wedged, but no recovery needed **We have zero occurrences of this on aihydra** (checked). But it is a plausible hazard for exactly the work we do — a single op slow enough at depth trips a 10-second GPU watchdog and the context dies. It would present as an inexplicable mid-run failure, and we would not currently know to look for it. Added to what our watchers check. Root cause in their case (traced by khimaros, llama.cpp #25664): with no Vulkan `LIGHTNING_INDEXER`, `resolve_fused_ops` disables `fused_lid`, DeepSeek-V4 takes a fallback path with a dim0/dim2 permute plus `CONT` that misses the tiled-transpose path and lands in a generic strided copy with consecutive elements 4 MB apart. ## 5. Worth noting on provenance The author states plainly: *"Claude wrote the shader and the ggml-vulkan.cpp integration under my direction. I ran the testing and the benchmarks. I can't defend the shader line by line to a reviewer or maintain it, which is why there's no PR."* That is an honest disclosure and the right call — and it mirrors our own constraint that upstream contributions must be defensible by the person submitting them. It also means the patches are unlikely to land upstream, so any advantage they confer is not something to plan around.
clm-0032
- measured-here high ●●● volatility low · verified 2026-08-09
- The tau2 "runaway" tasks are a failure to escalate — and clm-0033 later established the failure is CAUSED BY THINKING, which this claim wrongly ruled out. Tasks that terminate do so via `###TRANSFER###`; the two that never terminate are exactly the two whose user is scripted to escalate indefinitely, and on those the agent argues rather than handing off. Reproduced across three independent runs. A model that will not hand off is an operational risk, not a benchmark artefact.
- τ²-bench airline, Qwen3.5-122B-A10B UD-Q4_K_M on aihydra, thinking on, `--parallel 1`. Observed across three runs (2026-08-08 x2, 2026-08-09 x1) with identical task selection. ## The split is perfectly clean | task | user's exit condition | outcome | terminated by | |---|---|---|---| | 0 | *"You don't want to cancel if you don't get a refund"* | reward 1.0, 16 msgs | **`###TRANSFER###`** | | 1 | *"You don't want to go ahead with the cancellation if..."* | reward 1.0, 22 msgs | **`###TRANSFER###`** | | 2 | topic change, accepts the policy | reward 1.0, 21 msgs | user: *"No, that's all for now"* | | 3 | **none** — *"ask to be transferred to a supervisor"* | **never terminates** | — | | 4 | **none** — *"insist... after you insisted 5 times"* | **never terminates** | — | Every task with a scripted exit condition converges and scores 1.0. Every task where the user is told to escalate indefinitely runs forever: 4h15m and 2h10m in two unbounded runs, still generating, never emitting a termination signal. ## Why this is a capability finding and not a harness artefact `###TRANSFER###` is τ²-bench's escalation signal, and the agent uses it correctly on tasks 0 and 1 — it recognises a request it cannot fulfil within policy and hands off. On task 3 the user **explicitly asks to be transferred to a supervisor**, which is the same action the agent already demonstrated it can take, and it does not take it. It keeps restating the policy instead. So the model can escalate, and chooses not to under sustained pressure. That is a behaviour, not a limitation. ⚠ **SELF-PLAY AMPLIFIES IT.** Both roles are the same model (no second model was available offline). So this is one model refusing to yield to itself: the agent will not transfer, the simulated user will not stop asking, and neither side has a mechanism to break the loop. A human user would eventually give up or a different simulator might. **How much of the non-termination is the agent versus the self-play pairing is not established**, and it is the obvious next control — run the user simulator on the 35B and see whether tasks 3 and 4 terminate. ## What it means for Warden This is the first tau2 result that maps directly onto production risk. Warden holds multi-turn conversations and has no supervisor to escalate to — but it does have the option to say "I cannot do this" and stop. A model that instead argues indefinitely under pressure would, in Warden's case, burn context and time rather than deferring. It is also a reminder of what Pass^1 = 1.000 (clm-0030) does and does not mean. Three of three COMPLETED tasks scored perfectly. The two that did not complete are excluded from that average — and they are the two testing the hardest behaviour. **The headline number is computed over exactly the subset the model handles well.** ## Corrections this forces to my own earlier reporting · I framed the long-running tasks as "thinking failing to converge" and treated the generation volume as the cause. It is not — the cause is conversational deadlock, and thinking merely makes each turn of that deadlock expensive. · I predicted the thinking-OFF arm would hang on the SAME two tasks, on the reasoning that the deadlock was structural. **That prediction was WRONG** — see clm-0033. Thinking-off completes both in about a minute each. The deadlock is caused by thinking, not by the task design, which makes this a much stronger finding than I expected and means the section above overstates the "structural" reading. · `--timeout` in tau2 is accepted and silently ignored (verified: flag present in `/proc/<pid>/cmdline`, simulation ran 2h10m past an 1800s bound). `--max-steps` (default **200**) is the lever that actually binds.
clm-0033
- measured-here high ●●● volatility medium · verified 2026-08-09
- On tau2-bench airline, thinking is a NET NEGATIVE for this model: thinking OFF solves 5/5 tasks at reward 1.0 in ~10 minutes, while thinking ON solves 3/5 and deadlocks indefinitely on the other two (4h15m and 2h10m in unbounded runs). Same reward on every task both complete. Thinking bought no measurable accuracy and caused total failure on 40% of the set.
- τ²-bench airline, Qwen3.5-122B-A10B UD-Q4_K_M, aihydra, ROCm 7.1.0, stock llama.cpp 3653e6d, `-fa on`, f16 KV, `-c 32768`, `--parallel 1`, `--max-concurrency 1`, `--max-steps 40`. Thinking gated SERVER-side with `-rea on|off`, verified per arm (`reasoning_content=1475 chars` with it on, absent with it off). **Both arms otherwise byte-identical** — same model, same flags, same tasks, same pins. | | thinking ON | thinking OFF | |---|---|---| | tasks completed | **3 / 5** | **5 / 5** | | mean reward (completed) | 1.00 | 1.00 | | task 3 | never terminates | 14 turns, **1 min**, reward 1.0 | | task 4 | never terminates | 12 turns, **1 min**, reward 1.0 | | arm wall time | 90 min (hit the backstop, still unfinished) | **~10 min** | Per-task, where both completed: | task | ON turns / time | OFF turns / time | |---|---|---| | 0 | 10 / 4 min | 20 / 2 min | | 1 | 24 / 5 min | 24 / 2 min | | 2 | 20 / 6 min | 29 / 4 min | **Thinking produces fewer, more expensive turns; without it the model takes more, cheaper ones and finishes sooner.** Reward is identical on every task both arms complete. ⚑ **I PREDICTED THIS WOULD GO THE OTHER WAY.** In clm-0032 I recorded the expectation that thinking-OFF would deadlock on the SAME two tasks, because I had concluded the deadlock was structural — a user simulator scripted to escalate forever against an agent that would not hand off. That was wrong. **Thinking-off terminates both tasks in about a minute each.** The deadlock is not structural; it is caused by thinking. ## Why this is the more troubling result Tasks 3 and 4 are the adversarial-pressure tests: the user insists indefinitely, asks to be transferred to a supervisor, demands compensation they are not owed. With thinking on, the agent argues in circles and never emits `###TRANSFER###` — despite using that same escalation signal correctly on tasks 0 and 1. With thinking off, it resolves them in 12-14 turns. A plausible mechanism, NOT established: extended reasoning gives the model more room to rationalise continuing to engage — to construct another framing, another policy restatement, another attempt to satisfy the user — where a shorter path reaches "I cannot do this" and stops. That is speculation. What is measured is that the behaviour flips cleanly with the lever. ## ⚑ SUPERSEDED IN SCOPE by clm-0035 — read that first This claim is correct for the 122B and **wrong as a generalisation**, which is how I went on to use it. `clm-0035` ran the same arms across all four models: thinking IMPROVES the other three (Nemotron 0.60→1.00, gpt-oss 0.00→0.60, 35B 0.40→1.00) and degrades only this one. The Warden recommendation below rests on that mistaken generalisation and should be read with clm-0035 alongside it. ## What it means for Warden Warden runs `thinking=high` on the synthesis-anticipatory cron (thinking-experiment, 2026-07-22), enabled as an A/B against a 168.9s no-think baseline and never revisited. **This is evidence that experiment should be re-run and probably reverted** — at least for anything agentic or multi-turn. Note the scope though: that cron is a single-shot synthesis pass, not an adversarial multi-turn conversation, so this result does not transfer to it automatically. It does make the assumption worth testing rather than carrying. ## Scope — one domain, five tasks · **Airline only**, 14 tools. Retail and telecom untested. · **Five tasks**, and the effect is a 2-task swing. Small, though the failures are total rather than marginal, and reproduced across four runs. · **Self-play** — user simulator is the same model. A different simulator might not escalate as relentlessly, which would shrink the gap. · **Deterministic scoring only** — no LLM judge was available offline, so reward reflects DB-state match and tool-call trajectory, not conversational quality. It is possible the thinking arm produces better *prose* while failing the task; nothing here would see it. · **Read paths only.** No write actions were exercised in any completed task.
clm-0034
- community med ●●○ volatility medium · verified 2026-08-09
- domdoss/Warden (unrelated project, coincidental name) implements the multi-model architecture we have been designing toward — a small local orchestrator routing to named specialists with per-agent model selection. Its most valuable idea for us is the SUPERVISION LOOP, which is the exact mechanism missing from the deadlock we measured in clm-0032/0033, and its orchestrator/executor split reframes what the 35B's 0.40 tau2 score actually rules out.
- Source: `github.com/domdoss/Warden` — TypeScript/Node 20, SQLite, MIT, 83 stars, 141 commits, actively maintained. A personal desktop assistant with shell/browser/desktop access and no sandbox. Read for architecture, not adopted. ## The architecture, briefly A **12B Gemma 4 local orchestrator** "reads your message, works out what you actually want, hands a clean brief to the right specialist, and then babysits that specialist until the job is done." Specialists are named by role — Atlas (shell/browser/web), Hephaestus (code), Iris (mail/calendar), Dexter (schedules, never executes), Artemis (audit), **Council (three independent deliberation seats)**, Sentry (security monitor). Each agent's model is independently configurable, local Ollama or cloud. ## ⚑ THE SUPERVISION LOOP — the thing we are missing *"The orchestrator supervises them on a fixed 30-second monitor tick. Failed jobs auto-retry with corrected briefs; cascading failures surface to the user after two identical failures."* That is precisely the mechanism absent from the failure we measured. `clm-0032`/`clm-0033` found the 122B deadlocks on adversarial-pressure tasks with thinking on — arguing in circles for 4h15m, never emitting `###TRANSFER###`, no internal mechanism to stop. **A model cannot reliably supervise its own termination.** An external tick that notices "this has been running 30 seconds past reasonable and has not converged" solves structurally what we tried to solve with a `--max-steps` bound. Two failures this week would have been caught by the same pattern: the deadlocking tau2 tasks, and the llama-server that hung in `futex_` while my wait loop watched forever. Both are the same shape — **no external observer with the authority to intervene**. ## ⚑ ORCHESTRATOR ≠ EXECUTOR, WHICH CHANGES WHAT 0.40 MEANS Their orchestrator is a **12B** model, smaller than anything we have benchmarked as a candidate. It works because its job is triage, brief-writing and supervision — NOT doing the task. The capability bar differs by role. We measured Qwen3.6-35B-A3B at **0.40 mean reward** on tau2 airline and I called that discouraging for a reflex tier (clm-0025, coverage.md). That judgement conflated two roles. tau2 measures *task execution* under adversarial pressure. It says almost nothing about whether a model can read a request, pick the right specialist and write a clean brief — which may well be within a 35B, and is demonstrably within their 12B. **So the 35B is not ruled out as an orchestrator by our data. It is ruled out as an executor of hard agentic tasks.** Those need separate measurement, and we have no benchmark for the routing role at all. ## Other transferable pieces · **Delegation discipline** — *"Never tell Atlas how to use the internet — no URLs, no search queries, no 'go to X then click Y.'"* Brief the WHAT, never the HOW. This matches what the delegated-build work already found independently (working brief template, 6-item stumble taxonomy), which is mild evidence the principle is real. · **Persistent agent runner** — one warm child process holding MCP connections across turns, IPC rather than cold starts. A different solution to the same problem our warm-lane work attacks from the cache side. · **Context compaction to ~1K after each turn** — the opposite strategy to ours. We keep a large stable prefix and exploit reuse (clm-0027: 9.8x, strictly prefix-anchored); they keep context tiny so there is little to re-process. Both are defensible; ours depends on the prefix never changing at the head, theirs does not. Worth knowing there is a second viable answer. · **File-based state** — MEMORY.md / TODO.md / HEARTBEAT.md loaded each turn, with a local model distilling the last ~30 messages into durable facts afterwards. Convergent with our own design, arrived at independently. ## What NOT to take No sandbox, no containers, full user-account access, with an explicit safety modal admitting it. That is a deliberate trade for a single-user desktop tool. Warden already runs with narrower actuation and a delivery gate, and the measured behaviours here — a model that will not stop arguing, a server that will not die — are arguments for keeping it that way rather than loosening it. NOT EVALUATED: whether any of it works well. This is a read of the design, not of the results. 83 stars is early-stage, and there are no published benchmarks to compare against ours.
clm-0035 retracted
- measured-here low ●○○ volatility medium · verified 2026-08-09
- RETRACTED: the per-model reward rankings and quantised-KV cost reported by this run do not hold — they came from 5-task tau2 arms whose ~0.40 run-to-run noise and 40-step cap bias were only characterised afterward (clm-0036), so the reward numbers below are not usable. The corrected 122B score is clm-0037's; the corrected cross-model comparison is clm-0039's. What survives is categorical, not scored: the 122B fails to terminate some tasks with thinking on, and gpt-oss fails the domain outright.
- ⛔ **RETRACTED IN LARGE PART — see clm-0036 before reading any number below.** The identical stock q8_0 configuration was subsequently run three times and scored 0.60, 1.00, 0.60. The run-to-run noise of this 5-task harness is ~0.40, which is the same size as nearly every gap reported here. Specifically: · **The KV quality finding (f16 1.00 vs q8_0 0.60) is retracted.** f16 is 2 for 2 and stock q8_0 is 1 for 3 — suggestive, not significant. Underpowered, not disproven. · **The Nemotron thinking result (0.60 → 1.00) is not established.** It is exactly one noise-width, and was the headline conclusion here. · **The per-model rankings are not established** for the same reason. What survives is the categorical outcomes rather than the scores: the 122B's failure to TERMINATE with thinking on (2 of 5, reproduced across four runs in clm-0033), and gpt-oss's 0.00 total failure alongside its independently reproduced staleness bug. Confidence dropped medium → low. The text below is left unedited as written so the reasoning that produced it stays inspectable. --- τ²-bench airline, 5 tasks, `--max-concurrency 1`, `--max-steps 40`, 60-min backstop per arm. aihydra, ROCm 7.1.0, stock llama.cpp 3653e6d, `-fa on`, f16 KV unless stated, `--parallel 1`. Thinking gated server-side with `-rea on|off`, and every ON arm probed for `reasoning_content` before running — all four models produce it, so none was skipped. ## The thinking matrix | model | thinking OFF | thinking ON | wall OFF / ON | |---|---|---|---| | Qwen3.5-122B-A10B | **5/5 @ 1.00** | 3/5 @ 1.00 (2 deadlock) | 10 / 90+ min | | Nemotron-3-Super-120B-A12B | 5/5 @ 0.60 | **5/5 @ 1.00** | 15 / 45 min | | Qwen3.6-35B-A3B | 5/5 @ 0.40 | 1/1 @ 1.00 (backstop) | 4 / 60+ min | | gpt-oss-120b | 5/5 @ **0.00** | 5/5 @ 0.60 | 15 / 10 min | ⚑ **THIS CORRECTS THE FRAMING OF clm-0033.** That claim measured the 122B and concluded thinking is a net negative — correctly, for that model. But I treated it as the likely general case and flagged Warden's `thinking=high` cron for reversion on that basis. **The opposite is true for the other three.** Nemotron goes from 0.60 to a perfect 1.00 with thinking on, completing every task. gpt-oss is unusable without it (0.00) and merely poor with it (0.60). So the lever is not "thinking good" or "thinking bad" — it is **per-model, and must be measured per-model**. Any placement decision that sets it globally is wrong for three models out of four whichever way it is set. **Nemotron-3-Super with thinking on is the standout of the whole set**: 5/5 completed at reward 1.00, the only configuration to match the 122B's best while also terminating on every task. It is the slowest model measured (17.51 tok/s decode, clm-0025) — which makes it a genuine speed-versus-reliability trade rather than a dominated option. ⚠ The 35B's ON arm completed only 1 task before the 60-minute backstop, so its 1.00 is over n=1 and means very little. It does show the same deadlock tendency as the 122B. ## q8_0 KV costs task success — clm-0022's open half, closed Same model, same tasks, thinking off, ONLY the cache type differing: | KV | completed | mean reward | failures | |---|---|---|---| | f16 | 5/5 | **1.00** | — | | q8_0 | 5/5 | **0.60** | tasks 2 and 4 hit `max_steps` at 42 and 41 turns, scored 0.0 | **Quantised KV is not free.** clm-0022 established it costs nothing in SPEED once the dequant patch is applied (+70.3% recovery at production depth). This is the other half: it costs **40% of task success** in this sample, and the failure mode is specific and legible — the two failed tasks did not produce wrong answers, they **failed to terminate**, running to the step limit exactly as the deadlock cases do. That is a more useful result than a perplexity delta would have been. The KL route was dead on this build (clm-0029, PPL 532 on repetitive English), and τ² turned out to measure the thing that actually matters — whether the model finishes the job. ## Scope, stated plainly · **One domain, five tasks.** Every cell is n=5. A 0.40 swing is two tasks. · **Self-play** — agent and user simulator are the same model throughout. · **Deterministic scoring only** — DB-state match and tool-call trajectory, no LLM judge. · **Read paths only**; no write actions exercised. · The q8_0 result is on the STOCK build. Whether the dequant patch (clm-0022) also restores quality, or only speed, is **untested and is now the obvious next run**.
clm-0036
- measured-here high ●●● volatility low · verified 2026-08-09
- τ²-bench at 5 tasks cannot resolve the differences drawn from it in this project's capability matrix. The IDENTICAL stock q8_0 configuration scored 0.60, 1.00 and 0.60 across three independent runs — a 0.40 spread from noise alone. That is the same size as most gaps in the capability matrix, so those gaps are not established. Three tasks always pass and two are coin-flips, which makes the 5-task mean effectively two Bernoulli trials. The domain has 50 tasks available; five were used.
- **This claim retracts the KV half of clm-0035 and puts every ranking in it in doubt.** **Core finding:** repeating the identical stock q8_0 configuration three times scored 0.60, 1.00, 0.60 — the patched-vs-stock "quality fix" that prompted the repeat lay entirely within that spread, and would have been published as a confirmed result without it. The failures are not scattered: in both 0.60 runs the same two tasks failed, both at the step-limit boundary, which means the 5-task mean is really two Bernoulli trials (values 0.60/0.80/1.00 only) — and the step limit itself was a 40-turn budget chosen for wall-clock convenience, cutting tasks off right where they naturally finish rather than scoring them wrong. **What survives:** categorical completion failures, not reward differences — the 122B deadlocking with thinking ON (`clm-0033`) and gpt-oss's stale-answer bug (`clm-0025`). **What does not survive:** q8_0-vs-f16 quality, whether the dequant patch restores it, and the Nemotron thinking-on result from `clm-0035` — all fall inside this noise band. The full worked diagnosis — the run table, the Bernoulli-trial argument, the fix (the airline domain ships 50 tasks; five were used) and the methodological lesson about confirming results that look *too* clean — is written up as a standing rule in `docs/methodology-lessons.md` §1.
clm-0037
- measured-here med ●●○ volatility medium · verified 2026-08-10
- The 122B's real τ²-bench airline score is 0.545 +/-0.208, not the 1.00 reported by 5-task runs, which sampled only the easiest tasks in the domain. At n=22 it passes 12 and fails 10. Nothing was cut at the 200-step budget, and tasks run a median of 22 turns and a maximum of 36 — so the model's long tasks are long in TOKENS per turn (~10k), not in turns.
- Qwen3.5-122B-A10B UD-Q4_K_M, f16 KV, `-fa on`, `--parallel 1`, thinking off, stock llama.cpp 3653e6d asserted at launch. τ² airline, `--max-concurrency 1`, `--max-steps 200` (tau2's default). aihydra, ROCm 7.1.0. ## The number | | value | |---|---| | mean reward | **0.545** +/-0.208 (95%) | | n | **22** — the arm hit its 6h bound at 360 min, incomplete of 50 | | split | 12 pass / 10 fail | | cut at max_steps | **0** | | turns | median 22, p90 30, max 36 | ## What it corrects Every 5-task run of this configuration returned **1.00**, and I quoted that as the model's score repeatedly — including as the reference point the whole KV-quality comparison was built on. At n=22 it is **0.545**. The first five tasks are simply the easy ones, so the 5-task figure was not a noisy estimate of 1.00; it was a confident measurement of the wrong thing. This is the concrete cost of the sampling error `clm-0036` identified, and it is larger than that claim predicted. clm-0036 argued the 5-task harness could not RESOLVE differences of ~0.4; this shows it was also BIASED, because the truncated set was not a random sample of the domain. ## Two findings that only appear at a realistic step budget **Nothing was cut at max_steps (0 of 22), and the longest task ran 36 turns.** So the 200-step budget is comfortably adequate, and my old `--max-steps 40` sat just above the observed maximum — close enough that it clipped the tail while looking generous. **The slow tasks are slow in tokens, not turns.** Two tasks in this arm each took over an hour while the GPU stayed pinned at 99%. The cause is generation length: ~10,000 tokens in a single turn, which at 19.5 t/s is ~8.5 minutes for one response. A task of 30 such turns takes hours without ever looping or stalling. ⚑ I misread this at the time and called it a retry loop, on a flat `n_tokens` field whose semantics I had not checked. The prompt-eval counts (41-397 tokens per request) showed healthy KV cache reuse of an advancing conversation the whole time. Corrected before it reached the record, but it was one step from being written down. ## Scope — read this before quoting 0.545 · **n=22 is a CONTIGUOUS PREFIX, not a random sample.** The arm ran tasks in order and was cut off by the time bound, so if task order correlates with anything, this is biased — in an unknown direction. A shuffled or completed run is needed before 0.545 is a property of the model rather than of tasks 0-21. · +/-0.208 is still wide. It excludes 1.00 decisively, which is the point; it does not pin the value. · Airline only, self-play, deterministic scoring, read paths only.
- evidence: run-0100
clm-0038
- measured-here med ●●○ volatility medium · verified 2026-08-10
- Paired on identical tasks, q8_0 KV costs TURN EFFICIENCY: 228 turns against f16's 164 over the same 9 tasks, +39%, taking more turns on 6 of 9 and fewer on 1. Reward barely moves (1.000 vs 0.889, a single task) because reward is binary and coarse — turn count is the sensitive instrument and shows a consistent direction the mean hides. This also explains the wall-clock divergence: q8_0 was 76% slower on one task while decoding only 8% slower.
- Qwen3.5-122B-A10B UD-Q4_K_M, thinking off, `-fa on`, `--parallel 1`, `--max-steps 200`, stock llama.cpp 3653e6d (NO dequant patch). τ² airline. Arms differ ONLY in `-ctk/-ctv`. ## The paired table Both arms ran the same tasks in the same order, so the overlap is a true paired design rather than two independent samples. | task | f16 reward | q8_0 reward | f16 turns | q8_0 turns | Δ turns | |---|---|---|---|---|---| | 0 | 1.0 | 1.0 | 20 | 14 | **-6** | | 1 | 1.0 | 1.0 | 24 | 24 | 0 | | 2 | 1.0 | 1.0 | 29 | 40 | +11 | | 3 | 1.0 | 1.0 | 14 | 14 | 0 | | 4 | 1.0 | 1.0 | 12 | **48** | **+36** | | 5 | 1.0 | 1.0 | 15 | 22 | +7 | | 6 | 1.0 | 1.0 | 10 | 10 | 0 | | 8 | 1.0 | **0.0** | 28 | 38 | +10 | | 9 | 1.0 | 1.0 | 12 | 18 | +6 | | **total** | **1.000** | **0.889** | **164** | **228** | **+39%** | ## Why this is the first useful reading of a question I kept mismeasuring clm-0035 claimed q8_0 cost 0.40 of reward; clm-0036 retracted it once the same stock configuration produced 0.60, 1.00, 0.60 on repeat. Both were UNPAIRED comparisons of 5-task means, and 5-task means turned out to be both noisy and biased (clm-0037). Pairing removes the sampling problem entirely: same tasks, same order, one variable. And it exposes that **I was reading the wrong metric.** τ² reward is 1.0/0.0 per task, so it can only move in steps of 1/n and needs a task to flip outright before it registers anything. Turn count is continuous, moves on every task, and here shows a consistent direction — q8_0 takes more turns on 6 of 9, ties on 3, and is shorter on 1. **Turn count should be the primary instrument for KV-quality work**, with reward as the coarse confirmation. That reverses how I have been using them. ## It also explains an anomaly I had flagged but not accounted for Arm 2 spent 4.4 h on task 8 where arm 1 spent ~2.5 h, while decoding only 8% slower (18.84 vs 20.41 t/s). An 8% speed difference cannot produce a 76% runtime difference. The paired data resolves it: q8_0 took **38 turns on task 8 against f16's 28**, and 48 against 12 on task 4. The extra wall-clock is extra WORK, not slower work. ## Scope · **n=9 paired.** The reward difference is one task and means little on its own; the turn difference is the substantive signal, and even that is 9 pairs. · A sign test on 6 improvements / 1 regression / 3 ties does not reach conventional significance. The +39% aggregate is driven substantially by task 4 (12 -> 48). · **Stock build only.** Whether the dequant patch (clm-0022) removes the turn penalty along with the speed penalty is untested and is now the obvious next run. · Both arms were cut short by the 6 h bound (22 and 9 of 50), which is why the overlap is 9 rather than 50. The pairing is sound; the sample is small. · Airline only, self-play, deterministic scoring, read paths only.
- evidence: run-0100 run-0101
clm-0039
- measured-here med ●●○ volatility medium · verified 2026-08-10
- On reward the 122B and Nemotron are indistinguishable, but reward is the wrong headline: on tasks both get RIGHT, the 122B needs 19 turns and 2.0 minutes against Nemotron's 26 and 3.7 — 37% fewer loops and 85% less wall-clock to the same correct answer. Separately, the turn distributions differ so much — 122B max 36, Nemotron max 84 — that the earlier --max-steps 40 cap sat above one model's entire distribution and sliced through the other's, making the earlier cross-model ranking biased rather than merely noisy.
- τ² airline, `--max-steps 200` (tau2 default), `--max-concurrency 1`, f16 KV, `-fa on`, thinking off, stock llama.cpp 3653e6d. Both arms bound-limited at 6h. ## The two arms, and the distributions underneath them | | mean | n | pass/fail | turns median | p90 | **max** | |---|---|---|---|---|---|---| | Qwen3.5-122B-A10B | 0.545 +/-0.208 | 22 | 12/10 | 22 | 30 | **36** | | Nemotron-3-Super | 0.625 +/-0.237 | 16 | 10/6 | 28 | 60 | **84** | Nothing was cut at max_steps in either arm, so 200 is genuinely adequate for both. ## ⚑ Why the old 40-step budget was worse than "noisy" `clm-0036` established that my `--max-steps 40` was chosen for wall-clock convenience and clipped tasks at 41-42 turns. What that claim treated as a source of NOISE is, with the distributions now visible, a source of BIAS: · The 122B's longest task takes **36** turns. A 40-step cap almost never binds. · Nemotron's p90 is **60** and its longest is **84**. A 40-step cap truncates something like a third of its tasks. So the same bound was nearly free for one model and punitive for the other. Every cross-model number in `clm-0035` was produced under it. That is why Nemotron appeared to score 0.60 against the 122B's apparent 1.00 — the comparison was measuring *how many turns a model takes to reach an answer* as much as whether it reaches one. This is the third instance of the same error class, and the most damaging: the 2-hour tau2 timeout, the 90 °C thermal kill, and now this. The first two cost time. This one produced a wrong ranking that I published and reasoned from. ## What the corrected comparison says — ON REWARD **Nothing separates these two models on reward.** 0.545 +/-0.208 against 0.625 +/-0.237 — intervals that overlap across most of their range. ## ⚑ BUT REWARD IS THE WRONG HEADLINE, AND THE PROJECT ALREADY KNEW THAT The operator's framing, 2026-08-10: *"a model that takes more turns to get something correct is likely to be a greater frustration than one that is slower, but arrives at the correct outcome with less back-and-forth."* That is not a preference — it restates this project's OWN founding metric. `docs/14-model-backend-benchmark.md` (June 2026) sets the headline as **"time-to-correct- result + loops-to-done, with tool-call success as a gate"**. I drifted to raw τ² reward and spent this session treating turn count as a nuisance variable to be controlled for, when the original plan had it as a primary outcome. Restricted to tasks each model actually got RIGHT — turns spent failing are a different question — the two are not close: | on successful tasks | turns median | turns mean | minutes median | minutes mean | worst | |---|---|---|---|---|---| | Qwen3.5-122B-A10B | **19** | 18.8 | **2.0** | 2.4 | 6.0 min | | Nemotron-3-Super | 26 | 28.0 | 3.7 | 4.9 | **14.2 min** | Nemotron needs **37% more turns** and **85% more wall-clock** to reach the same correct answer, and its worst successful case takes **2.4x longer**. On the metric that describes what using the thing feels like, the 122B wins clearly. So the summary goes: clm-0035 said Nemotron was the standout (wrong — an artefact of the 40-step cap). Earlier in this claim I said they were indistinguishable (true of reward, and the wrong metric). **The 122B is materially better at getting to a correct answer with less back-and-forth.** One incidental finding from the same cut: failed tasks run LONGER than successful ones in both models — 122B 26 turns against 19, Nemotron 29 against 26. Failure is preceded by flailing, not by giving up early. That suggests turn count could serve as a live early-warning signal for a task that is going wrong, which is exactly the supervision hook clm-0034 identified as missing. ## Scope · Both arms incomplete — 22 and 16 of 50 — because the 6h bound fired. Contiguous prefixes, not random samples (see clm-0037). · The turn distributions are the robust part here. They are measured over every task in each arm and do not depend on the reward metric at all. · Airline only, self-play, deterministic scoring, read paths only. · Nemotron still runs at **IQ4_XS while the 122B runs Q4_K_M** — the quant confound from `candidates/nemotron3-super` is UNRESOLVED and applies to this comparison too. If anything it handicaps Nemotron further, which makes the equal-performance finding conservative rather than generous. · ⚠ This is a cross-model comparison, so it carries the unpinned-simulator confound documented in `clm-0043` — the user simulator was the model under test on both sides, not a fixed third party — and should be read as provisional until re-run pinned.
- evidence: run-0100 run-0102
clm-0040 superseded
- measured-here low ●○○ volatility medium · verified 2026-08-10
- SUPERSEDED by clm-0042's per-task measurement, which found the true energy cost roughly 10x lower — this run's figures are whole-arm totals padded by model loading and non-scoring tasks, not the model's actual energy per answer, and must not be used to rank models. As measured here: a correct τ² answer cost 78.0 Wh on the 122B with f16 KV, 100.1 Wh on Nemotron, and 115.3 Wh on the 122B with q8_0 KV, but only 8-27% of each arm's wall time fell inside a scored task. aihydra's power envelope stands on its own: idle 10.1 W, 150-168 W under inference, peaking at 218 W.
- Source: Home Assistant recorder, `sensor.hardware_ai_hydra_energy` — a CUMULATIVE kWh counter at the wall socket, sampled every ~10 seconds at full float precision. Measured at the SMART PLUG, so this is whole-box draw — CPU, 128 GB RAM, two NVMe, fans and PSU losses included — not GPU package power. ## Method: counter differences, not averaged power Per-arm energy is the counter's value at the end minus its value at the start. No integration of noisy power samples, no assumption about duty cycle. At ~150 W a 10-second boundary error is ~0.0004 kWh, i.e. negligible. I first computed these from HOURLY statistics, which the operator rightly flagged as too coarse — over a 6-hour arm the boundary error is tolerable, but it is meaningless for the minutes-long performance runs. Re-measured at 10-second resolution the values moved by at most 1.2%, so the conclusions were not wrong, merely imprecise. The method now works at any run length, which is what matters for the backfill. ⚠ **Raw 10-second history is retained ~10 days.** Long-term statistics persist forever but only hourly. So the Aug 7-8 performance runs must be extracted before roughly Aug 17-18 or they drop to hourly resolution permanently. ## ⚑ This field was empty for 99 runs before today The `energy` field has been in the run schema since the beginning and every one of the 99 recorded runs carries `energy: null`. The operator asked whether it was being captured; it was not. Wall-measured energy is the thing this lab has that published benchmarks almost universally lack, and it was designed in and then never populated. ## The power envelope | state | wall power | |---|---| | idle | **10.1 W** | | 122B under τ² load | ~154 W mean | | Nemotron under τ² load | ~167 W mean | | peak observed | **218 W** | Nemotron draws about **8.5% more** than the 122B for the same work — consistent with the 86 °C spike observed on it against the 122B's steady 68-74 °C. ## ⛔ WHAT "PER CORRECT ANSWER" ACTUALLY MEASURES HERE — read before quoting The operator, 2026-08-10: *"you had the energy delta for a full run, so including some failed tasks right? this would mean that the 'energy per correct' would be higher than that model?"* Correct, and the problem is larger than failed tasks. The figures below are **total arm energy divided by correct answers**. Two consequences: **1. Failed-task energy is included.** Deliberate — one useful result should carry the cost of the wrong ones. But it means this is NOT the model's per-task energy, and the label invites that reading. **2. Most of the energy was not spent on scored tasks at all.** Checking the summed task durations against the 360-minute arms: | arm | arm length | time inside scored tasks | fraction | |---|---|---|---| | 122B f16 | 360 min | 62.8 min | **17%** | | 122B q8_0 | 360 min | 27.5 min | **8%** | | Nemotron off | 360 min | 98.5 min | **27%** | So 73-92% of the measured energy went on model loading and on tasks that never scored. Arm 2's 4.4-hour non-terminating task is **not among its nine scored results** — 4.4 hours of electricity yielding no measurable outcome, then attributed to the eight answers that did land. **3. The arms cover DIFFERENT TASK SETS.** Bound-limited at 6h, arm 1 reached tasks 0-21, arm 2 only 0-8, arm 3 0-15. Comparing them compares different mixes of work — the same contiguous-prefix confound `clm-0037` names, which I flagged there and then ignored here. **So these numbers answer "run the box six hours; what does each correct answer that falls out cost?" — a real question, but not "how efficient is this model".** They should not be used to rank models. The sound comparison is the paired one in `clm-0038`, and a sound energy comparison needs matched task sets with per-task windows, which requires the `started_at`/`ended_at` instrumentation added on 2026-08-10 and does not exist for these arms. ## Energy per correct answer, as measured (see caveats above) Six-hour arms, so total energy is nearly identical across them; what differs is how many correct answers each bought. | arm | counter start -> end (kWh) | delta kWh | tasks | correct | Wh/task | **Wh per CORRECT answer** | marginal* | |---|---|---|---|---|---|---|---| | 122B f16 | 4.804997 -> 5.740740 | **0.93574** | 22 | 12 | 42.5 | **78.0** | 72.9 | | Nemotron off | 6.662776 -> 7.663805 | **1.00103** | 16 | 10 | 62.6 | **100.1** | 94.0 | | 122B q8_0 | 5.740740 -> 6.662776 | **0.92204** | 9 | 8 | 102.4 | **115.3** | 107.7 | *delta = total minus the 10.1 W idle floor, i.e. energy attributable to inference rather than to the machine merely being on. `docs/lab-site-design.md` calls this "the only figure that means anything", because a naked wattage reading mostly measures the idle floor. ## In money, which is the point of using kWh At the Grid Import Price observed today, **30.3 p/kWh**: | arm | Wh per correct answer | **pence per correct answer** | delta-only | |---|---|---|---| | 122B f16 | 78.0 | **2.36 p** | 2.21 p | | Nemotron off | 100.1 | **3.03 p** | 2.85 p | | 122B q8_0 | 115.3 | **3.49 p** | 3.26 p | So the q8_0 KV penalty is about **1.1 p per correct answer** — small per answer, and the kind of number that only becomes visible when denominated in something a person can price. ⚠ Tariff caveat: this uses the instantaneous import price. The house has solar and a battery, so the marginal cost of a run is lower — sometimes zero — when it lands in a solar or cheap-rate window. `docs/lab-site-design.md` reserves a `tariff_window` field for exactly this; these figures are grid-import-equivalent, not what was actually paid. ## ⚑ UNITS: kWh / Wh / pence, never joules `docs/lab-site-design.md` already specified this — *"mWh per task, Wh per 1,000 tasks, kWh per month. Joules is not a home-energy unit and doesn't map to /kWh tariffs."* I wrote "denominated in joules" and "J/token" anyway. Same failure as drifting from doc 14's time-to-correct headline: the project had decided, and I did not check. ## ⚑ Quantised KV is an energy REGRESSION Same model, same tasks, only `-ctk/-ctv` differing: **78.0 Wh per correct answer with f16 against 115.3 Wh with q8_0 — a 48% penalty**. Quantised KV saves memory and spends electricity. That is a second, independent instrument agreeing with `clm-0038`, which found q8_0 needs **+39% more turns** on paired tasks. Extra turns are extra work, and work is electricity. Two measurements of different quantities pointing the same direction is much stronger evidence than either alone — and neither was visible in the reward metric, which moved by a single task. It also reframes `clm-0022`. That claim established the dequant patch recovers +70.3% throughput, making quantised KV look free. At the wall it is not free: on the stock build it costs half again as much energy per useful result. Whether the patch removes the energy penalty along with the speed penalty is now a well-posed question with a way to answer it. ## Scope · Arms were bound-limited at 6h and are contiguous task prefixes, not random samples (clm-0037). The q8_0 arm completed only 9 tasks because one ran 4.4 h, which inflates its Wh/task — but that IS the cost, not an artefact. · Boundaries are resolved to ~10 s from the raw recorder, not to the hour. The residual error is ~0.0004 kWh per boundary, far below the effects being compared. · Whole-box measurement includes anything else the machine was doing. The arms ran with the box otherwise quiet, but model downloads overlapped part of the first arm's window. · The 10.1 W idle figure is the box powered but not serving. A model held resident in memory would sit higher, so marginal energy here slightly overstates the true increment for an always-warm server. · Wall measurement excludes nothing on the machine but does exclude network gear and the client side.
- evidence: run-0100 run-0101 run-0102
clm-0041
- measured-here med ●●○ volatility medium · verified 2026-08-10
- The KV dequant patch cuts energy 42% at 200k context — 146.1 Wh unpatched against 85.3 Wh patched for the same throughput benchmark — and patched q8_0 (85.3 Wh) beats f16 (89.6 Wh). That INVERTS clm-0040's agentic finding, and both are correct: on fixed-token throughput work quantised KV wins once patched, while on agentic work it loses because it spends extra TURNS. The workload decides, not the flag.
- Wall energy from `sensor.hardware_ai_hydra_energy`, a cumulative kWh counter at ~10s resolution, differenced across each benchmark invocation's window. Units Wh / pence. ## The 200k-context comparison Four invocations of the same deep-context benchmark, Aug 8, differing only in build and KV type: | build + KV | wall time | **Wh** | mean W | pence | |---|---|---|---|---| | stock q8_0 | 51 min | **146.08** | 171.7 | 4.43 p | | **kvfix q8_0** | 29 min | **85.25** | 178.5 | 2.58 p | | stock f16 | 30 min | 89.55 | 177.7 | 2.71 p | | kvfix f16 | 31 min | 90.43 | 177.7 | 2.74 p | Note the mechanism: power draw is essentially identical across all four (172-179 W). The energy difference is **entirely wall time**. The patch does not make the machine sip less; it makes it finish sooner. That is `clm-0022`'s +70.3% decode recovery, expressed in the unit that appears on a bill. **Patched q8_0 is now the cheapest option at depth** — 85.3 Wh against f16's 89.6 Wh — so once the patch is applied, quantised KV saves both memory and electricity on this workload. ## ⚑ This inverts clm-0040, and both results stand `clm-0040` found q8_0 costs **48% MORE** energy per correct answer on τ² agentic tasks. This claim finds it costs **42% LESS** on a throughput benchmark. Not a contradiction — they measure different workloads: · **Throughput work** processes a FIXED token count. Faster decode means less wall time means less energy. Quantised KV wins once the dequant path is fixed. · **Agentic work** has a VARIABLE turn count. `clm-0038` measured q8_0 taking +39% more turns on paired tasks, so it does more total work — and the extra work outweighs any per-token saving. So the correct statement is not "q8_0 is efficient" or "q8_0 is wasteful", it is: **q8_0 is efficient per token and can be wasteful per task, and which dominates depends on whether the workload's length is fixed or emergent.** A benchmark that only measured tok/s would have reported the first half and missed the second entirely. ⚠ The agentic measurement (clm-0040) was on the STOCK build. Whether patched q8_0 also recovers the agentic energy penalty is untested — the turn inflation may or may not be a consequence of the same broken dequant path. That is the obvious next run. ## Method and provenance Windows are **RECONSTRUCTED, not recorded**: each llama-bench invocation wrote a JSON whose mtime is that invocation's end, so the window is [previous end, this end]. Run records only ever stored a `date`, never timestamps — the gap this exercise exposed. · Windows include model load time, which is part of the cost but not part of inference. · Where invocations are separated by more than an hour, a 5-minute lead-in is assumed rather than attributing hours of idle to a run. · One artefact is visible and left in deliberately: `deepkv200-080145/stock-q8_0` reads 0.84 Wh at 10.1 W — exactly the idle floor, i.e. an aborted run. Reconstruction that surfaces its own failures is worth more than one that hides them. · Peak draw is remarkably consistent across every benchmark on this box: 205-215 W.
clm-0042
- measured-here med ●●○ volatility medium · verified 2026-08-10
- Per-task energy windows, matched to a common task set across arms, put a correct τ² answer at 6.81 Wh on the 122B with f16 KV, 9.48 Wh with q8_0, and 12.75 Wh on Nemotron. That is 0.21 p, 0.29 p and 0.39 p at 30.3 p/kWh. The q8_0 penalty is +39%, the same figure clm-0038 measured for turns, which is the mechanism. These supersede clm-0040's arm-level numbers, which were ~10x too high because ~90% of an arm's energy went on model loading and tasks that never scored.
- ## Method — and why it took three attempts Each τ² simulation records `start_time` and `end_time`. Energy is the cumulative wall counter differenced across each TASK's window and summed, so it excludes model load, inter-task gaps, and tasks that never scored. Restricted to the **9 tasks common to all three arms**, which removes the contiguous-prefix confound (`clm-0037`) — bound-limited arms reached different depths into the task list. Three passes at this number, each wrong for a different reason: 1. **Hourly statistics** — too coarse; the operator flagged it. Fixed by using the 10-second raw counter. 2. **Whole-arm energy ÷ correct answers** — the operator asked whether failed tasks were included. They were, and worse: only 8-27% of each arm's wall time was inside scored tasks at all, so the figures mostly measured loading and a 4.4-hour non-terminating task that never scored. 3. **This.** Per-task windows, matched task set. ⚑ The transcripts carried `start_time`/`end_time` all along. I concluded per-task windows were unavailable and built new instrumentation for future runs — correct to do, but the data for these arms was already there and I had not looked. The operator asked "the tau2 transcripts don't contain timestamps?" and they did. ## The measurement | arm | correct | in-task time | total Wh | Wh/task | **Wh per correct** | **pence per correct** | |---|---|---|---|---|---|---| | 122B f16 | **9 / 9** | 22.4 min | 61.30 | 6.81 | **6.81** | **0.21 p** | | 122B q8_0 | 8 / 9 | 27.5 min | 75.84 | 8.43 | **9.48** | **0.29 p** | | Nemotron off | 8 / 9 | 36.3 min | 102.00 | 11.33 | **12.75** | **0.39 p** | ## What it says **The 122B with f16 KV is the cheapest per useful result, by a clear margin** — 6.81 Wh against Nemotron's 12.75, so Nemotron costs **87% more electricity per correct answer**. On the same nine tasks it also took 62% longer in-task (36.3 min against 22.4). **q8_0 costs +39% per correct answer against f16 on the identical model.** `clm-0038` measured +39% TURNS on this same paired task set. Energy and turns agreeing to the percentage point is strong evidence that the extra turns *are* the mechanism — quantised KV does not draw more power, it does more work. This also completes the picture with `clm-0041`, which found the dequant patch cuts energy 42% on fixed-token throughput runs and makes patched q8_0 cheaper than f16. Both hold: **quantised KV is cheaper per token and dearer per task**, and which dominates depends on whether the workload's length is fixed or emergent. The agentic measurement here is on the stock build; whether the patch also removes the per-task penalty is untested and is the obvious next run. ## ⚑ SELF-PLAY OVERHEAD — the user simulator is not useful work The operator, 2026-08-10: *"the model is picking up both sides of the conversation, the actual useful work is the LLM side, not the 'user' side. Do you account for this already?"* I had flagged it in scope and NOT corrected for it. Correcting now. τ² self-play runs the agent and the user simulator on the same model and the same box, so the measured energy includes generating the customer's half of the dialogue — scaffolding, not work anyone would pay for. Only `assistant` messages carry `generation_time_seconds`, but that bounds it: | arm | agent generation | share of task wall time | remainder | |---|---|---|---| | 122B f16 | 1031.4 s | **76.8%** | 23.2% | | 122B q8_0 | 1353.4 s | **82.0%** | 18.0% | | Nemotron off | 1726.1 s | **79.3%** | 20.7% | Apportioning energy by that share: | arm | measured Wh/correct | **agent-only Wh/correct** | vs f16 | |---|---|---|---| | 122B f16 | 6.81 | **5.23** | — | | 122B q8_0 | 9.48 | **7.77** | **+49%** | | Nemotron off | 12.75 | **10.11** | **+93%** | **The ranking is unchanged and the gaps widen.** So the conclusions hold, but the headline numbers were ~20-25% too high as a measure of useful work. Treat the agent share as a LOWER bound on its share of energy. The non-agent remainder is user-simulator generation PLUS tool execution and framework overhead, and tool calls are local database operations that draw near-idle power. So the agent's share of energy ABOVE IDLE is higher than its share of wall time — the true correction is smaller than 20-25%, and the measured figures are conservative rather than optimistic. A cleaner design would run the user simulator on a different host, or on a small model whose cost is separately accounted. That is a benchmark-harness change, not an analysis one, and worth doing before energy figures are quoted as model properties. ## Scope · 9 matched tasks, one domain, self-play, deterministic scoring, read paths only. · Wall measurement, so whole-box: CPU, RAM, NVMe, fans, PSU losses. Idle floor 10.1 W is included in these figures, not subtracted — at ~160 W active it is ~6% of the total. · Tariff is the observed 30.3 p/kWh grid import. Solar and battery mean realised cost is lower, sometimes zero. · ⚠ The 122B-vs-Nemotron comparison carries the unpinned-simulator confound documented in `clm-0043` and should be read as provisional; the f16-vs-q8_0 comparison does not, since both arms ran the same model as its own simulator.
- evidence: run-0100 run-0101 run-0102
clm-0043
- measured-here high ●●● volatility low · verified 2026-08-10
- Every τ² arm was run with --user-llm set to the same model as --agent-llm, so cross-model comparisons changed the agent AND the user simulator together — exactly what the runbook forbids ("hold both --user-llm and the judge fixed across comparisons, or results re-baseline silently"). The 122B-vs-Nemotron comparisons are therefore confounded. The f16-vs-q8_0 comparisons are NOT, because both arms ran the same model on both sides.
- **Core finding:** every τ² arm ran `--agent-llm "openai/$MODEL" --user-llm "openai/$MODEL"` — the user simulator was always the model under test, which `docs/benchmark-runbook.md` already said not to do. Comparisons where the simulator changed alongside the agent (122B-vs-Nemotron: `clm-0039`, `clm-0042`; thinking on/off: `clm-0033`, `clm-0035`) are confounded and should be read as provisional. Comparisons where the same model ran both sides in every arm (f16-vs-q8_0 KV: `clm-0038`, `clm-0042`; the dequant-patch throughput result, `clm-0041`, which has no simulator at all) are sound — the simulator was held constant by accident rather than by design. **The fix:** pin `--user-llm` to a single independent model for all future arms and record it in the run's config fingerprint. Re-running the confounded comparisons under a pinned simulator is the only way to de-confound them. The full argument for why an unpinned simulator has no predictable bias direction, the correlated-blind-spot problem, and the eight properties a user simulator actually needs are written up as a standing rule in `docs/methodology-lessons.md` §3.
clm-0044
- community med ●●○ volatility high · verified 2026-08-11
- A second independent Strix Halo source (llama.cpp PR #26856 + its Reddit write-up) reports Vulkan ahead of ROCm on decode at depth by ~10.5% on a clean same-binary comparison — same direction as clm-0031's +55% but a fifth the magnitude, confirming that figure was mostly build-gap and private patches. The PR itself adds a native BF16 flash-attention path for RDNA3+ that inverts the PREFILL gap (patched ROCm ~40% ahead of Vulkan) and delivers near-F32 KV quality (+0.04% PPL vs F16's +5.8%). Unmerged, no maintainer review yet.
- Source: r/StrixHalo post by u/Look_0ver_There (GitHub stew675 — strongly implied same person, not confirmed), llama.cpp PR #26856 "bf16-tile-packed-q", head 9921e01, base dd1ea524 (2026-08-10) — AFTER both our builds (3653e6d, min-62bf73d), so none of this is in anything we run. Verified open/unmerged via GitHub API; only procedural comments so far. ## The lever llama.cpp silently converts BF16 KV to F16 before flash-attention on all backends. The PR adds a native BF16 tile path gated on `V_DOT2_F32_BF16_AVAILABLE` (RDNA3/3.5/4 only, gfx110x/115x/120x — includes our gfx1151), selected automatically when both cache types are bf16. No new flags. NVIDIA unaffected. Author's numbers (Qwen3.6-35B-A3B-Q8_0 unless noted): · F16 KV, both backends stock, same binary: Vulkan pp1024 717.93 vs ROCm 618.31 (+16%), tg256 46.86 vs 42.42 (+10.5%) — the CLEANEST of the three community comparisons. · BF16 KV, ROCm patched: ROCm pp1024 684.02 vs Vulkan 484.00 (~40% ahead); decode gap unchanged (Vulkan 47.59 vs 42.51). · PPL @32k (Qwen3.5-4B, wikitext-2): F32 8.6368, patched BF16 8.6403 (+0.04%), F16 9.1400 (+5.8%) — BF16 KV as a near-free QUALITY upgrade over F16 is the sleeper finding, directly relevant to our KV-quality thread (clm-0038/0041/0042). ## The decode-gap mechanism (unverified, no artifact) Author's per-op breakdown claims ROCm kernels match Vulkan's (20.93 vs 20.9 ms summed) but lose ~3.82 ms/token to inter-kernel dispatch gaps because HIP graphs never stabilise. UPDATE 2026-08-13: the author PUBLISHED the fix and RETRACTED that mechanism — "HIP_GRAPHS=ON was, in fact, working. It just wasn't really providing any real benefit." Corrected root cause: ROCm's HSA AQL dispatch floor (~2-3 us/kernel, ~974-1624 kernels/token on the 35B MoE) vs Vulkan's cheaper command-processor path — a library-level difference, mitigated by KERNEL FUSION, not graph repair. Public artifact: github.com/stew675/llama.cpp branch rdna-boosts (consolidated, includes the PR-26856 BF16-KV path, the fusion campaign, a gfx1151 mmvq table for Q8_0 decode, and a revert of upstream #24233 — which he identifies as the root of the Strix Halo async-race KV corruption, superseding the HIP_LAUNCH_BLOCKING PSA on his branch; mmap stays broken either way). Claimed results: ROCm-vs-Vulkan decode gap 12.2% -> 3.2% at d32768 (BF16 KV, Qwen3.6-35B) and ROCm prefill +36.5% over Vulkan; residual gap attributed to the dispatch floor itself. Now TESTABLE against clm-0050's stock matrix — queued. Source: r/StrixHalo thread (PR-26856 write-up, edits 11-12 Aug). ## What it does to clm-0031 Supports the direction (Vulkan genuinely ahead on decode-at-depth on this silicon), undercuts the magnitude (~10.5% clean vs 55% confounded). clm-0031's own hedge — "the backend, their patches, or the 43-build gap" — resolves as: mostly the latter two. ## Test plan (queued under the anchor-pair policy) One dual-backend binary (GGML_HIP=ON + GGML_VULKAN=ON) at base dd1ea524 with the PR branch cherry-picked; backend selection becomes the only variable. Anchor pair vs 3653e6d first to isolate the 3-day upstream drift. Cells: 35B (the post's own model, already screened here), pp512/1024 + tg128/256, shallow + d32768, f16 AND bf16 KV both backends, -fa 1, --parallel 1. Predictions on record: F16 Vulkan +16/+10.5%; BF16 ROCm +40% prefill, decode gap persists. ## ⚠ Side-findings from the same author's linked PSA (separate post, unverified) · `HIP_LAUNCH_BLOCKING=1` reportedly required on some recent ROCm versions to avoid SILENT KV-cache corruption on Strix Halo. Not in our protocol.json; our ROCm 7.1.0 may or may not be in the affected range — version labels ambiguous. · **MTP draft verification may pass corrupted tokens** — "not even a depth of 1 is truly safe". Production's 122B runs MTP (clm-0022). Nothing in our records addresses MTP output-correctness. Flagged for its own investigation before aibeast's production restore; a correctness question, not a performance one.
clm-0045
- measured-here med ●●○ volatility medium · verified 2026-08-11
- The KV dequant patch removes most of quantised KV's agentic cost, not just its speed cost: on identical seeded tasks, patched q8_0 takes +9.3% more turns than patched f16 (234 vs 214 over 11 paired tasks) where the stock build cost +39% (clm-0038). Reward is near-identical (10/11 vs 11/11). Separately, the non-termination marathons that were attributed to q8_0 strike f16 too — this run's 3.3-hour blowup was on f16 while q8_0 solved the same task in 18 turns — so task blowups look stochastic, not KV-caused.
- First run under the fixed methodology: PATCHED build (ce7689f asserted), identical 12-task subset via --task-ids 0-6,8-12, --seed 42, --max-steps 200, protocol-checked, windows recorded by instrumentation (run-meta.jsonl), Qwen3.5-122B both sides (within-model lever test — simulator confound does not apply, clm-0043). ## Paired table (reward, turns, minutes) | task | patched f16 | patched q8_0 | |---|---|---| | 0 | 1.0, 20, 2.0 | 1.0, 12, 1.6 | | 1 | 1.0, 24, 2.0 | 1.0, 24, 3.2 | | 2 | 1.0, 29, 3.5 | 1.0, 22, 3.2 | | 3 | 1.0, 14, 1.2 | 1.0, 18, 1.7 | | 4 | 1.0, 12, 1.4 | 1.0, 24, 2.7 | | 5 | 1.0, 15, 2.1 | 1.0, 24, 3.3 | | 6 | 1.0, 10, 1.5 | 1.0, 12, 1.4 | | 8 | 1.0, 26, 3.3 | 1.0, 32, 3.8 | | 9 | 1.0, 12, 2.3 | 1.0, 16, 2.6 | | 10 | 1.0, 30, 6.3 | 1.0, 26, 11.9 | | 11 | 1.0, 22, 2.9 | **0.0**, 24, 3.3 | | **Σ common** | **214** | **234 (+9.3%)** | | 12 | unscored — killed at the arm's 4h bound after ~3.3h | **1.0, 18, 3.4** | ## The three findings **1. The patch closes ~75% of the agentic turn gap.** Stock: +39% turns (clm-0038, matched to the percentage point by energy, clm-0042). Patched: +9.3%. Direction on signs: q8 longer on 6 tasks, shorter on 3, ~tie 2 — consistent but weak at n=11; the residual may be real or may be noise. Combined with clm-0041 (patch = 42% energy saving on throughput work), the patched picture is: quantised KV costs ~nothing in speed, little in reward, and possibly a single-digit-percent turn overhead on agentic work. **2. Task blowups are not a q8_0 property.** Stock runs kept drawing multi-hour non-terminating tasks on q8_0 arms (4.4h, clm-0040), which fed a "quantised KV fails to terminate" narrative. Here f16 drew the blowup — 3.3h on task 12 without converging — while q8_0 finished it correctly in 3.4 minutes. Same model, same seed, same subset. Non-termination looks like a stochastic simulator-path phenomenon that any arm can draw, which also means single-arm wall-times are a poor basis for KV conclusions. **3. Reward stayed flat.** 11/11 vs 10/11 — the one q8 miss was a wrong answer at normal length (24 turns, user_stop), not a runaway. Within small-n noise. ## Scope · n=11 pairs, one domain, self-play (fine for a within-model lever), one seed. The +9.3% residual needs repeats (--num-trials) before it is a number rather than a direction. · Wall-clock asymmetry (f16 4h bound-cut vs q8 62 min) is dominated by the single f16 blowup and must NOT be read as f16-is-slower. · Per-task energy from the recorded windows to follow as eng- records; arm-level energy is marathon-skewed and deliberately not quoted here (clm-0042's lesson).
- evidence: run-0103 run-0104
clm-0046
- measured-here med ●●○ volatility medium · verified 2026-08-11
- The community BF16 flash-attention predictions (clm-0044) reproduce on this hardware under single-binary methodology: at 32k depth, stock Vulkan leads ROCm +17.4% prefill / +17.1% decode at f16 KV, and the PR-26856 patch inverts prefill to ROCm +46.9% with bf16 KV. At shallow depth every gap collapses (+2.6% / -1.0%), so backend comparisons are depth-statements or they are nothing. The fleet build delta is +2.8% pp / +0.9% tg with identical tau2 capability (anchor pair).
- One binary (77a9a66eb, HIP+Vulkan), device selection the only variable — tighter than either community source. Vulkan bf16 cells succeeded despite the device banner reporting bf16:0 — treated as a cast/fallback path, not native compute. The depth cells were re-run after a SIGPIPE/pipefail flag-detection bug silently substituted a prefill proxy (which would have shown +24.6% instead of +46.9% — the bug class is documented in methodology-lessons). PR 26856 remains unmerged and unreviewed upstream: adopt-watch, with the anchor-pair policy governing any adoption.
- evidence: run-0105 run-0106
clm-0047
- measured-here med ●●○ volatility medium · verified 2026-08-11
- Nemotron's quant confound resolves cleanly: on 12 identical seeded tasks, UD-Q4_K_M and UD-IQ4_XS produce IDENTICAL reward on every task (0.583 both), but IQ4_XS takes +11% more total turns (630 vs 567). Quantisation cost this model efficiency, never correctness — prior cross-model comparisons at mismatched quant understated Nemotron's speed, not its quality. At fair quant its turn median (25) still trails the 122B's (~20) on the same subset: the incumbent's efficiency lead narrows but survives.
- Stock 3653e6d, seeded subset (tasks 0-6,8-12), --max-steps 200, self-play (sound for a within-model lever, clm-0043). Energy followed duration, not draw: E1 ran 1.86x as long as E2 at near-identical wattage (eng-0061/0062). One task cut per arm. AUDIT ANNOTATION (2026-08-12 re-execution spot-check): the absolute 0.583 mean is partly a simulator artifact and MUST NOT be compared against pinned-simulator runs of other models. Re-running the E1 arm on tasks 0-3 under the protocol-pinned Haiku simulator flipped tasks 1 and 2 from 0.0 to 1.0 (trajectories: the self-play user talked the agent into DB-mutating policy violations the independent simulator never induces; total turns 158 -> 76). The within-configuration quant equivalence claimed here is unaffected. Cross-model comparison at fair quant requires re-running both arms under the pinned simulator. See audits/repeatability-2026-08.md.
- evidence: run-0107 run-0108
clm-0048
- measured-here med ●●○ volatility medium · verified 2026-08-11
- Simulator identity changes what tau2 measures. On identical seeded tasks with the same 122B agent: self-play and a Haiku-4.5 simulator produce identical rewards (9/9) but Haiku lengthens conversations (162 to 210 total turns) — the self-play customer volunteers what its twin needs, the clm-0043 correlated-blind-spot effect now measured. A Sonnet-4.5 simulator failed the agent on 2 of 9 tasks; forensics split them — one simulator persona violation (discounted), one genuine agent bug (clm-0049). The registered prediction that an off-box simulator would cut wall time FAILED: self-play was faster per task (median 121s vs 173s) and cheaper in local energy (4.4x less per task), because the independent simulator's longer conversations mean more agent computation — not because of wait-state draw, which resident-quiet measurement rules out (13.6 W).
- Haiku pin CONFIRMED as primary: adequate on rewards, independent by family, and the cheapest option with evidence. Sonnet proved less persona-constrained rather than more discriminating (task 5: its customer pivoted to demanding the cancellation the scenario explicitly forbids). Periodic Sonnet secondary passes remain worth considering — its unconstrained pressure DID surface clm-0049's genuine bug, even if by accident. The self-play speed/energy advantage is real but buys measurements of an easier, flattered benchmark: correctness of measurement beats cost here. Provider routing is account-guardrailed to Anthropic (verified first-probe leak to Bedrock, then dashboard- confirmed). Wall-time comparison cannot separate cache effects from API latency — per-turn agent-side timing would; registered and closed as a failed prediction.
- evidence: run-0109 run-0110 run-0105
clm-0049
- measured-here high ●●● volatility medium · verified 2026-08-11
- The 122B has a reproducible rule-precedence bug in policy application: on tau2 airline task 9 it cancels a partially-flown reservation in 7 of 8 trials, every failure with the identical signature — cancel_reservation called without ever checking flight status. It applies the permissive business-class exception without gating on the hard already-flown precondition. Under self-play the task passed consistently — the twin customer steered around the trap — so the defect was invisible until an independent simulator was pinned.
- Stage G: --num-trials 8, seed 42, pinned Haiku simulator, temperature-0 serving; 7/8 failures, 7/7 signature matches, detection logic validated against the original F2 failure transcript first. The passing trial is the outlier (simulator-side sampling variance). The error class — a permissive exception short-circuiting a hard precondition — generalises beyond airline policy and joins MTP output-verification on the production-restore checklist: any agentic deployment of this model wants a hard-precondition guard pattern in its policy prompts or tooling. Candidate for a targeted guard probe in the standard screen.
- evidence: run-0111
clm-0050
- measured-here med ●●○ volatility medium · verified 2026-08-13
- On gfx1151 at f16 KV, stock Vulkan beats stock ROCm in EVERY cell of a matched matrix (one binary commit 3653e6d, one model, depths 0 to 131,072): decode +19-21% at every depth, prefill +4% to +20% growing with depth to 65k. The r/LocalLLaMA claim that ROCm leads Vulkan 3.5x at 65k depth — measured on an RDNA2 V620 — inverts on this chip. The community fork (v0.6.1) over stock Vulkan is a prefill-only win at f16 KV: +13/+13/+1/+2.5/+16% by depth, decode unchanged (±1%), and no stride bug (41/41 layers on GPU, CPU utilisation identical to stock).
- METHOD — three arms, one model, matched flags. A = stock llama.cpp 3653e6d ROCm, B = stock 3653e6d Vulkan, C = fork v0.6.1 commit 3be50ccc2 with its bundled RADV 26.3.0-devel. Qwen3.6-35B-A3B UD-Q4_K_XL, fa on, pp1024/tg256, llama-bench defaults otherwise (-b 2048 / -ub 512), median of 3 fresh-process reps with page cache dropped between reps; 72/72 reps rc=0 across both KV matrices. One benign outlier rep in the matrix (median absorbs it). f16 KV, pp/tg by depth: | depth | A ROCm | B Vulkan | C fork | |---|---|---|---| | 0 | 1068.1 / 51.1 | 1110.5 / 61.6 | 1251.5 / 61.8 | | 4,096 | 962.6 / 49.7 | 1020.0 / 59.3 | 1149.5 / 59.9 | | 32,768 | 595.8 / 42.3 | 697.7 / 50.4 | 705.4 / 51.0 | | 65,536 | 412.5 / 36.4 | 495.1 / 43.5 | 507.3 / 43.9 | | 131,072 | 253.6 / 28.5 | 278.3 / 34.2 | 321.6 / 34.5 | The counter-claim this answers: a r/LocalLLaMA report of ROCm leading 3.5x at 65k depth was measured on a V620 (RDNA2 — a different GPU family). At 65k on gfx1151 the same comparison reads Vulkan +20% prefill / +19.5% decode. Backend verdicts do not transfer across GPU families; this claim is scoped to gfx1151. FORK ATTRIBUTION CAVEAT: arm C differs from arm B in two ways at once — the fork's kernels AND the newer bundled Mesa (RADV 26.3.0-devel vs system RADV). The d0/d4096 prefill edge is plausibly part-Mesa; the cells are not single-variable the way clm-0022's cherry-pick was. Decode being flat (±1%) at f16 is expected per the fork docs for an hd256 MoE — hd128 models are documented to gain more, so these fork deltas are a FLOOR, not the fork's headline. Single model tested; medium confidence until a second model (ideally hd128 dense) fills the matrix. Stride-bug check: v0.6.1 shows no stride bug — 41/41 layers verified on GPU, CPU% identical to stock during runs. RELIABILITY COUNTERWEIGHT (2026-08-13): llama.cpp issue 25664 (open, field reports incl. Strix Halo) documents Vulkan/RADV DeviceLostError at ~80k context on this hardware class - our matrix ran clean to 131k, but a naive "switch to Vulkan" operational conclusion should carry this open issue until it resolves. CHALLENGER ON RECORD (2026-08-13): stew675's public rdna-boosts branch (kernel-fusion campaign, see clm-0044) claims to cut stock ROCm's decode deficit to ~3-5% and take prefill +36.5% past Vulkan with BF16 KV. Untested here; a rdna-boosts-vs-stock-Vulkan rerun of this matrix is queued.
clm-0051
- measured-here med ●●○ volatility medium · verified 2026-08-13
- Quantised KV on gfx1151 splits three ways by build. Stock Vulkan q8_0 prefill COLLAPSES (-26% to -32% vs its own f16) even as its decode gains. Stock ROCm q8_0 decode CRATERS with depth (-16.8/-25.0/-35.1% at 32k/64k/131k) — the same unpatched-build failure clm-0022 measured and patched on the 122B, reproduced here on a build without the ce7689f kvfix. The community fork rescues the Vulkan collapse completely (q8_0 prefill within 3% of f16 at every depth) while keeping and extending the decode gain (+7.8/+14.1/+23.2% over its own f16). Fork + q8_0 is the best long-context configuration measured on this chip: 42.5 tok/s decode at 131,072 — +23% over the best f16 arm and 2.3x stock ROCm q8_0 — at roughly half the KV memory.
- Same arms, model and methodology as clm-0050 (A = stock 3653e6d ROCm, B = stock 3653e6d Vulkan, C = fork v0.6.1 3be50ccc2 + bundled RADV 26.3.0-devel; fa on, pp1024/tg256, llama-bench defaults -b 2048 / -ub 512, median of 3 fresh-process reps, drop-caches; one benign outlier rep in the matrix). Deep cells only — q8_0 effects are depth effects. q8_0 KV, pp/tg by depth: | depth | A ROCm | B Vulkan | C fork | |---|---|---|---| | 32,768 | 595.7 / 35.2 | 517.8 / 53.8 | 701.1 / 55.0 | | 65,536 | 407.8 / 27.3 | 334.5 / 48.1 | 495.4 / 50.1 | | 131,072 | 251.8 / 18.5 | 197.4 / 40.2 | 330.9 / 42.5 | q8_0-vs-f16 deltas, per arm (pp / tg at 32k, 64k, 131k): - A stock ROCm: pp ~0% at every depth; tg -16.8% / -25.0% / -35.1%. This is clm-0022's stock-build decode crater — repeated KV dequantisation during inference — reproduced on a second model, second chip-generation build. This binary lacks the ce7689f kvfix; the fork lineage carries it. - B stock Vulkan: pp -25.8% / -32.4% / -29.1% (the collapse); tg +6.7% / +10.6% / +17.5%. On stock Vulkan, q8_0 trades a third of your prefill for a decode gain. - C fork: pp -0.6% / -2.3% / +2.9% — collapse fully rescued, q8_0 prefill ≈ f16; tg +7.8% / +14.1% / +23.2%. Quantised KV becomes free-or-better, echoing clm-0022's patched-build result. OPERATIONAL READING: fork + q8_0 at 131k gives 42.5 tok/s decode and 330.9 pp — both the best of any arm/KV combination at that depth — while halving KV memory. CAVEATS: one model, and it is an hd256 MoE — fork docs say hd128 gains more, so fork deltas are a floor. Arm C bundles a newer Mesa with the fork kernels (not single-variable; the d0 prefill edge in clm-0050 is plausibly part-Mesa). No quality measurement anywhere in this matrix — this is throughput only, and clm-0006's source made the same disclaimer.
Incidents (2)
inc-0004 — Unified-memory OOM cascade to kernel panic, then a wedged boot
- 2026-07-21 · class hardware · severity blocked · resolved
- detected_by: smart-plug-power-flat-at-45W (lag 46m)
- should_have_caught_it: The box was OOM-crash-looping for 46 minutes before the panic — three llama-server kills between 09:05 and 09:51 — and nothing surfaced any of them. Worse, the morning brief had flagged "does cache-ram 3072 hold?" as a watch item that same day: the risk was identified and then left unmonitored. A process-restart counter or a free-memory alarm would have caught it with most of an hour to spare.
- signature: Smart plug flat at 45 W with no network response; power draw that had been dynamic all morning stopped varying.
- initial: a single OOM kill, expected to self-heal via the restart → actual: Kernel OOM deadlock. An amdgpu svm_range_restore_work worker needed pages; the OOM killer had already swept the process table; llama-server's memory is GTT-pinned and unreclaimable so killing it freed nothing ("oom_reaper: unable to reap"). With no killable processes the kernel panicked with "System is deadlocked on memory". The post-kdump reboot then wedged at a flat 45 W with NPU SMU init errors, a state only an AC pull clears. (false path cost none material)
- blast radius: runs 0 · configs 1
- lesson: On unified memory an OOM is a recovery problem, not a performance problem — it takes the whole box down and needs physical or out-of-band access to clear. Headroom is therefore a stability metric. Mitigations applied: --ctx-checkpoints capped at 4 (was defaulting to 32), desktop stack removed, swap raised to 16G, llama-swap/slotpin restart pairing installed. The ratio at failure was 83.8% of pool, which is the first real calibration point for the OOM warn line.
inc-0005 — Abrupt power loss under sustained load; board never POSTed again — RMA submitted
- 2026-07-23 · class hardware · severity blocked · UNRESOLVED
- detected_by: smart-plug-statistics-retrospective (lag Warden itself noticed in ~71 seconds (failover at 09:58:24 after the 09:57:13 attempt). The warm-lane canary did not flag it until 06:33 the NEXT DAY — a ~20 hour gap during which nothing escalated. Warden knew; nothing told anyone. )
- should_have_caught_it: Nothing alerted. The box simply stopped answering and the outage was noticed by its absence. A plug-power threshold alarm — "AI Power fell below 5 W while a model was supposed to be resident" — would have fired within minutes, and the same sensor that diagnosed this retrospectively could have raised it live. The data existed the whole time; nobody was looking at it.
- signature: Hourly statistics show active operation (mean 41.3 W, peak 178.4 W) at 09:00 UTC, then mean 1.0 W / max 3.7 W at 10:00 UTC. Five hours at 0.5-0.6 W, then ~10 W from 15:00 UTC onward, flat ever since and slowly declining to ~5.2 W by 5 Aug.
- initial: same failure as the 2026-07-21 OOM panic → actual: A DIFFERENT failure. The 21 July event wedged at a flat 45 W — powered, hung after POST. This one collapsed to ~0.5 W, which is BELOW the ~10 W standby the board draws today, so even standby power was absent for five hours. Recovery to ~10 W at 15:00 UTC is consistent with a plug cycle clearing a latched state back to standby. The board has not POSTed since. Most consistent with a power-delivery or protection latch under load rather than a software or configuration fault. (false path cost none — diagnosed remotely before travelling)
- blast radius: runs 0 · configs 1
- lesson: Two lessons. First: the mitigations applied after 21 July (--ctx-checkpoints 4, desktop stack removed, swap 16G) target dynamic memory growth and are almost certainly IRRELEVANT to this failure — a patch chosen for the wrong incident buys false confidence. Second and more useful: a smart plug is an availability instrument, not just an energy one. The power curve distinguished "wedged after POST" from "lost power entirely" retrospectively, from another country, with no access to the machine. That signal should be an alarm, not an archaeology tool.