<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>HALOBENCH — log</title>
    <link>https://halobench.com/log/</link>
    <atom:link href="https://halobench.com/feed.xml" rel="self" type="application/rss+xml" />
    <description>The lab notebook, generated from the record: claims landed, gate changes, run series, incidents and contributions — retractions badged inline.</description>
    <language>en-gb</language>
    <lastBuildDate>Thu, 13 Aug 2026 00:00:00 GMT</lastBuildDate>
    <item>
      <title>[claim] On gfx1151 at f16 KV, stock Vulkan beats stock ROCm in EVERY cell of a matched matrix (one binary commit 3653e6d, one model, depths 0 to 131,072): decode +19-21% at every depth, prefill +4% to +20% growing with depth to 65k.</title>
      <link>https://halobench.com/log/#e-2026-08-13-1</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-13-1</guid>
      <pubDate>Thu, 13 Aug 2026 00:00:00 GMT</pubDate>
      <description>On gfx1151 at f16 KV, stock Vulkan beats stock ROCm in EVERY cell of a matched matrix (one binary commit 3653e6d, one model, depths 0 to 131,072): decode +19-21% at every depth, prefill +4% to +20% growing with depth to 65k. — records: clm-0050</description>
      <category>claim</category>
    </item>
    <item>
      <title>[claim] Quantised KV on gfx1151 splits three ways by build.</title>
      <link>https://halobench.com/log/#e-2026-08-13-2</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-13-2</guid>
      <pubDate>Thu, 13 Aug 2026 00:00:00 GMT</pubDate>
      <description>Quantised KV on gfx1151 splits three ways by build. — records: clm-0051</description>
      <category>claim</category>
    </item>
    <item>
      <title>[gate] 7 candidates → screened: gemma4-12b, glm-45-air, ling-30-flash, llama4-scout, maple-preview, nemotron35-lightning-30b, qwen36-27b-mtp</title>
      <link>https://halobench.com/log/#e-2026-08-13-3</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-13-3</guid>
      <pubDate>Thu, 13 Aug 2026 00:00:00 GMT</pubDate>
      <description>7 candidates → screened: gemma4-12b, glm-45-air, ling-30-flash, llama4-scout, maple-preview, nemotron35-lightning-30b, qwen36-27b-mtp</description>
      <category>gate</category>
    </item>
    <item>
      <title>[gate] nemotron35-lightning-30b: listed → acquired — Official ggml-org Q4_K_M downloaded to aihydra (25,430,738,944 bytes, exactly the listed size) and sha256-verified against the HF LFS oid (6110e2e2e6cd324e6ee69ddced5a6b34fad6c94ca9827222a1e420fb92e3c90b).</title>
      <link>https://halobench.com/log/#e-2026-08-13-4</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-13-4</guid>
      <pubDate>Thu, 13 Aug 2026 00:00:00 GMT</pubDate>
      <description>nemotron35-lightning-30b: listed → acquired — Official ggml-org Q4_K_M downloaded to aihydra (25,430,738,944 bytes, exactly the listed size) and sha256-verified against the HF LFS oid (6110e2e2e6cd324e6ee69ddced5a6b34fad6c94ca9827222a1e420fb92e3c90b).</description>
      <category>gate</category>
    </item>
    <item>
      <title>[gate] deepseek-v4-flash: listed → listed — Identity now concrete via the r/LocalLLaMA Strix Halo guide thread: DeepSeek V4 Flash 0731, deepseek4 arch, 256 experts/6 active + 1 shared, MIT licence, 1M native context.</title>
      <link>https://halobench.com/log/#e-2026-08-12-1</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-12-1</guid>
      <pubDate>Wed, 12 Aug 2026 00:00:00 GMT</pubDate>
      <description>deepseek-v4-flash: listed → listed — Identity now concrete via the r/LocalLLaMA Strix Halo guide thread: DeepSeek V4 Flash 0731, deepseek4 arch, 256 experts/6 active + 1 shared, MIT licence, 1M native context. — records: clm-0031</description>
      <category>gate</category>
    </item>
    <item>
      <title>[gate] nemotron35-lightning-30b: listed — Surfaced via an operator-shared r/AIDeveloperNews link (&quot;NVIDIA has launched Nemotron 3.5 Lightning&quot;) plus a follow-up r/StrixHalo post on a community ROCmFP4 requant with hardware-matched Strix Halo numbers.</title>
      <link>https://halobench.com/log/#e-2026-08-12-2</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-12-2</guid>
      <pubDate>Wed, 12 Aug 2026 00:00:00 GMT</pubDate>
      <description>nemotron35-lightning-30b: listed — Surfaced via an operator-shared r/AIDeveloperNews link (&quot;NVIDIA has launched Nemotron 3.5 Lightning&quot;) plus a follow-up r/StrixHalo post on a community ROCmFP4 requant with hardware-matched Strix Halo numbers.</description>
      <category>gate</category>
    </item>
    <item>
      <title>[claim] A second independent Strix Halo source (llama.cpp PR #26856 + its Reddit write-up) reports Vulkan ahead of ROCm on decode at depth by ~10.5% on a clean same-binary comparison — same direction as clm-0031's +55% but a fifth the magnitude, confirming that figure was mostly build-gap and private patches.</title>
      <link>https://halobench.com/log/#e-2026-08-11-1</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-11-1</guid>
      <pubDate>Tue, 11 Aug 2026 00:00:00 GMT</pubDate>
      <description>A second independent Strix Halo source (llama.cpp PR #26856 + its Reddit write-up) reports Vulkan ahead of ROCm on decode at depth by ~10.5% on a clean same-binary comparison — same direction as clm-0031's +55% but a fifth the magnitude, confirming that figure was mostly build-gap and private patches. — records: clm-0044</description>
      <category>claim</category>
    </item>
    <item>
      <title>[claim] The KV dequant patch removes most of quantised KV's agentic cost, not just its speed cost: on identical seeded tasks, patched q8_0 takes +9.3% more turns than patched f16 (234 vs 214 over 11 paired tasks) where the stock build cost +39% (clm-0038).</title>
      <link>https://halobench.com/log/#e-2026-08-11-2</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-11-2</guid>
      <pubDate>Tue, 11 Aug 2026 00:00:00 GMT</pubDate>
      <description>The KV dequant patch removes most of quantised KV's agentic cost, not just its speed cost: on identical seeded tasks, patched q8_0 takes +9.3% more turns than patched f16 (234 vs 214 over 11 paired tasks) where the stock build cost +39% (clm-0038). — records: clm-0045</description>
      <category>claim</category>
    </item>
    <item>
      <title>[claim] The community BF16 flash-attention predictions (clm-0044) reproduce on this hardware under single-binary methodology: at 32k depth, stock Vulkan leads ROCm +17.4% prefill / +17.1% decode at f16 KV, and the PR-26856 patch inverts prefill to ROCm +46.9% with bf16 KV.</title>
      <link>https://halobench.com/log/#e-2026-08-11-3</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-11-3</guid>
      <pubDate>Tue, 11 Aug 2026 00:00:00 GMT</pubDate>
      <description>The community BF16 flash-attention predictions (clm-0044) reproduce on this hardware under single-binary methodology: at 32k depth, stock Vulkan leads ROCm +17.4% prefill / +17.1% decode at f16 KV, and the PR-26856 patch inverts prefill to ROCm +46.9% with bf16 KV. — records: clm-0046</description>
      <category>claim</category>
    </item>
    <item>
      <title>[claim] Nemotron's quant confound resolves cleanly: on 12 identical seeded tasks, UD-Q4_K_M and UD-IQ4_XS produce IDENTICAL reward on every task (0.583 both), but IQ4_XS takes +11% more total turns (630 vs 567).</title>
      <link>https://halobench.com/log/#e-2026-08-11-4</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-11-4</guid>
      <pubDate>Tue, 11 Aug 2026 00:00:00 GMT</pubDate>
      <description>Nemotron's quant confound resolves cleanly: on 12 identical seeded tasks, UD-Q4_K_M and UD-IQ4_XS produce IDENTICAL reward on every task (0.583 both), but IQ4_XS takes +11% more total turns (630 vs 567). — records: clm-0047</description>
      <category>claim</category>
    </item>
    <item>
      <title>[claim] Simulator identity changes what tau2 measures.</title>
      <link>https://halobench.com/log/#e-2026-08-11-5</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-11-5</guid>
      <pubDate>Tue, 11 Aug 2026 00:00:00 GMT</pubDate>
      <description>Simulator identity changes what tau2 measures. — records: clm-0048</description>
      <category>claim</category>
    </item>
    <item>
      <title>[claim] The 122B has a reproducible rule-precedence bug in policy application: on tau2 airline task 9 it cancels a partially-flown reservation in 7 of 8 trials, every failure with the identical signature — cancel_reservation called without ever checking flight status.</title>
      <link>https://halobench.com/log/#e-2026-08-11-6</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-11-6</guid>
      <pubDate>Tue, 11 Aug 2026 00:00:00 GMT</pubDate>
      <description>The 122B has a reproducible rule-precedence bug in policy application: on tau2 airline task 9 it cancels a partially-flown reservation in 7 of 8 trials, every failure with the identical signature — cancel_reservation called without ever checking flight status. — records: clm-0049</description>
      <category>claim</category>
    </item>
    <item>
      <title>[gate] muse-glimmer-30b: blocked → screened — Screened on min-62bf73d (its minimum build, anchor-calibrated: +2.8%/+0.9% vs fleet baseline, identical tau2 capability).</title>
      <link>https://halobench.com/log/#e-2026-08-11-7</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-11-7</guid>
      <pubDate>Tue, 11 Aug 2026 00:00:00 GMT</pubDate>
      <description>muse-glimmer-30b: blocked → screened — Screened on min-62bf73d (its minimum build, anchor-calibrated: +2.8%/+0.9% vs fleet baseline, identical tau2 capability).</description>
      <category>gate</category>
    </item>
    <item>
      <title>[runs] 9 runs landed on cfg-0029, cfg-0030, cfg-0025, cfg-0031, cfg-0028, cfg-0027 (tau2-bench-airline)</title>
      <link>https://halobench.com/log/#e-2026-08-11-8</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-11-8</guid>
      <pubDate>Tue, 11 Aug 2026 00:00:00 GMT</pubDate>
      <description>9 runs landed on cfg-0029, cfg-0030, cfg-0025, cfg-0031, cfg-0028, cfg-0027 (tau2-bench-airline) — records: cfg-0029, cfg-0030, cfg-0025, cfg-0031, cfg-0028, cfg-0027</description>
      <category>runs</category>
    </item>
    <item>
      <title>[claim] The 122B's real τ²-bench airline score is 0.545 +/-0.208, not the 1.00 reported by 5-task runs, which sampled only the easiest tasks in the domain.</title>
      <link>https://halobench.com/log/#e-2026-08-10-1</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-10-1</guid>
      <pubDate>Mon, 10 Aug 2026 00:00:00 GMT</pubDate>
      <description>The 122B's real τ²-bench airline score is 0.545 +/-0.208, not the 1.00 reported by 5-task runs, which sampled only the easiest tasks in the domain. — records: clm-0037</description>
      <category>claim</category>
    </item>
    <item>
      <title>[claim] Paired on identical tasks, q8_0 KV costs TURN EFFICIENCY: 228 turns against f16's 164 over the same 9 tasks, +39%, taking more turns on 6 of 9 and fewer on 1.</title>
      <link>https://halobench.com/log/#e-2026-08-10-2</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-10-2</guid>
      <pubDate>Mon, 10 Aug 2026 00:00:00 GMT</pubDate>
      <description>Paired on identical tasks, q8_0 KV costs TURN EFFICIENCY: 228 turns against f16's 164 over the same 9 tasks, +39%, taking more turns on 6 of 9 and fewer on 1. — records: clm-0038</description>
      <category>claim</category>
    </item>
    <item>
      <title>[claim] On reward the 122B and Nemotron are indistinguishable, but reward is the wrong headline: on tasks both get RIGHT, the 122B needs 19 turns and 2.0 minutes against Nemotron's 26 and 3.7 — 37% fewer loops and 85% less wall-clock to the same correct answer.</title>
      <link>https://halobench.com/log/#e-2026-08-10-3</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-10-3</guid>
      <pubDate>Mon, 10 Aug 2026 00:00:00 GMT</pubDate>
      <description>On reward the 122B and Nemotron are indistinguishable, but reward is the wrong headline: on tasks both get RIGHT, the 122B needs 19 turns and 2.0 minutes against Nemotron's 26 and 3.7 — 37% fewer loops and 85% less wall-clock to the same correct answer. — records: clm-0039</description>
      <category>claim</category>
    </item>
    <item>
      <title>[superseded] by clm-0042's per-task measurement, which found the true energy cost roughly 10x lower — this run's figures are whole-arm totals padded by model loading and non-scoring tasks, not the model's actual energy per answer, and must not be used to rank models.</title>
      <link>https://halobench.com/log/#e-2026-08-10-4</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-10-4</guid>
      <pubDate>Mon, 10 Aug 2026 00:00:00 GMT</pubDate>
      <description>by clm-0042's per-task measurement, which found the true energy cost roughly 10x lower — this run's figures are whole-arm totals padded by model loading and non-scoring tasks, not the model's actual energy per answer, and must not be used to rank models. — records: clm-0040, clm-0042</description>
      <category>superseded</category>
    </item>
    <item>
      <title>[claim] The KV dequant patch cuts energy 42% at 200k context — 146.1 Wh unpatched against 85.3 Wh patched for the same throughput benchmark — and patched q8_0 (85.3 Wh) beats f16 (89.6 Wh).</title>
      <link>https://halobench.com/log/#e-2026-08-10-5</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-10-5</guid>
      <pubDate>Mon, 10 Aug 2026 00:00:00 GMT</pubDate>
      <description>The KV dequant patch cuts energy 42% at 200k context — 146.1 Wh unpatched against 85.3 Wh patched for the same throughput benchmark — and patched q8_0 (85.3 Wh) beats f16 (89.6 Wh). — records: clm-0041</description>
      <category>claim</category>
    </item>
    <item>
      <title>[claim] Per-task energy windows, matched to a common task set across arms, put a correct τ² answer at 6.81 Wh on the 122B with f16 KV, 9.48 Wh with q8_0, and 12.75 Wh on Nemotron.</title>
      <link>https://halobench.com/log/#e-2026-08-10-6</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-10-6</guid>
      <pubDate>Mon, 10 Aug 2026 00:00:00 GMT</pubDate>
      <description>Per-task energy windows, matched to a common task set across arms, put a correct τ² answer at 6.81 Wh on the 122B with f16 KV, 9.48 Wh with q8_0, and 12.75 Wh on Nemotron. — records: clm-0042</description>
      <category>claim</category>
    </item>
    <item>
      <title>[claim] Every τ² arm was run with --user-llm set to the same model as --agent-llm, so cross-model comparisons changed the agent AND the user simulator together — exactly what the runbook forbids (&quot;hold both --user-llm and the judge fixed across comparisons, or results re-baseline silently&quot;).</title>
      <link>https://halobench.com/log/#e-2026-08-10-7</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-10-7</guid>
      <pubDate>Mon, 10 Aug 2026 00:00:00 GMT</pubDate>
      <description>Every τ² arm was run with --user-llm set to the same model as --agent-llm, so cross-model comparisons changed the agent AND the user simulator together — exactly what the runbook forbids (&quot;hold both --user-llm and the judge fixed across comparisons, or results re-baseline silently&quot;). — records: clm-0043</description>
      <category>claim</category>
    </item>
    <item>
      <title>[gate] 7 candidates → acquired: cascade2-30b, gemma4-26b, glm-47-flash, laguna-s-21, lfm2-24b, muse-glimmer-30b, qwen36-27b-mtp</title>
      <link>https://halobench.com/log/#e-2026-08-10-8</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-10-8</guid>
      <pubDate>Mon, 10 Aug 2026 00:00:00 GMT</pubDate>
      <description>7 candidates → acquired: cascade2-30b, gemma4-26b, glm-47-flash, laguna-s-21, lfm2-24b, muse-glimmer-30b, qwen36-27b-mtp</description>
      <category>gate</category>
    </item>
    <item>
      <title>[gate] 5 candidates → screened: cascade2-30b, gemma4-26b, glm-47-flash, laguna-s-21, lfm2-24b</title>
      <link>https://halobench.com/log/#e-2026-08-10-9</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-10-9</guid>
      <pubDate>Mon, 10 Aug 2026 00:00:00 GMT</pubDate>
      <description>5 candidates → screened: cascade2-30b, gemma4-26b, glm-47-flash, laguna-s-21, lfm2-24b</description>
      <category>gate</category>
    </item>
    <item>
      <title>[gate] 16 candidates → listed: deepseek-v4-flash, gemma4-12b, glm-45-air, ling-30-flash, llama4-scout, maple-preview, muse-glimmer-30b, npu-embeddinggemma, npu-embeddinggemma, npu-lfm2, npu-qwen3-4b-thinking, npu-qwen3-4b-thinking, npu-whisper, npu-whisper, qwen38-27b, qwen38-27b</title>
      <link>https://halobench.com/log/#e-2026-08-10-10</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-10-10</guid>
      <pubDate>Mon, 10 Aug 2026 00:00:00 GMT</pubDate>
      <description>16 candidates → listed: deepseek-v4-flash, gemma4-12b, glm-45-air, ling-30-flash, llama4-scout, maple-preview, muse-glimmer-30b, npu-embeddinggemma, npu-embeddinggemma, npu-lfm2, npu-qwen3-4b-thinking, npu-qwen3-4b-thinking, npu-whisper, npu-whisper, qwen38-27b, qwen38-27b — records: clm-0031, clm-0034, clm-0015, clm-0018</description>
      <category>gate</category>
    </item>
    <item>
      <title>[gate] 4 candidates → blocked: laguna-s-21, laguna-s-21, muse-glimmer-30b, muse-glimmer-30b</title>
      <link>https://halobench.com/log/#e-2026-08-10-11</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-10-11</guid>
      <pubDate>Mon, 10 Aug 2026 00:00:00 GMT</pubDate>
      <description>4 candidates → blocked: laguna-s-21, laguna-s-21, muse-glimmer-30b, muse-glimmer-30b</description>
      <category>gate</category>
    </item>
    <item>
      <title>[runs] 2 runs landed on cfg-0026, cfg-0027 (tau2-bench-airline)</title>
      <link>https://halobench.com/log/#e-2026-08-10-12</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-10-12</guid>
      <pubDate>Mon, 10 Aug 2026 00:00:00 GMT</pubDate>
      <description>2 runs landed on cfg-0026, cfg-0027 (tau2-bench-airline) — records: cfg-0026, cfg-0027</description>
      <category>runs</category>
    </item>
    <item>
      <title>[superseded] the Pass^1 = 1.000 reported by this run came from a 3-task subsample biased toward the domain's easiest tasks — the other two of the original five never terminated and were excluded as infrastructure errors.</title>
      <link>https://halobench.com/log/#e-2026-08-09-1</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-09-1</guid>
      <pubDate>Sun, 09 Aug 2026 00:00:00 GMT</pubDate>
      <description>the Pass^1 = 1.000 reported by this run came from a 3-task subsample biased toward the domain's easiest tasks — the other two of the original five never terminated and were excluded as infrastructure errors. — records: clm-0030, clm-0037</description>
      <category>superseded</category>
    </item>
    <item>
      <title>[claim] A community DeepSeek-V4-Flash report independently confirms our mmap/GTT double-residency finding, demonstrates a 120 GiB GTT ceiling in production use, and — most consequentially — reports Vulkan BEATING ROCm on 3 of 4 cells including 55% faster decode at depth.</title>
      <link>https://halobench.com/log/#e-2026-08-09-2</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-09-2</guid>
      <pubDate>Sun, 09 Aug 2026 00:00:00 GMT</pubDate>
      <description>A community DeepSeek-V4-Flash report independently confirms our mmap/GTT double-residency finding, demonstrates a 120 GiB GTT ceiling in production use, and — most consequentially — reports Vulkan BEATING ROCm on 3 of 4 cells including 55% faster decode at depth. — records: clm-0031</description>
      <category>claim</category>
    </item>
    <item>
      <title>[claim] The tau2 &quot;runaway&quot; tasks are a failure to escalate — and clm-0033 later established the failure is CAUSED BY THINKING, which this claim wrongly ruled out.</title>
      <link>https://halobench.com/log/#e-2026-08-09-3</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-09-3</guid>
      <pubDate>Sun, 09 Aug 2026 00:00:00 GMT</pubDate>
      <description>The tau2 &quot;runaway&quot; tasks are a failure to escalate — and clm-0033 later established the failure is CAUSED BY THINKING, which this claim wrongly ruled out. — records: clm-0032</description>
      <category>claim</category>
    </item>
    <item>
      <title>[claim] On tau2-bench airline, thinking is a NET NEGATIVE for this model: thinking OFF solves 5/5 tasks at reward 1.0 in ~10 minutes, while thinking ON solves 3/5 and deadlocks indefinitely on the other two (4h15m and 2h10m in unbounded runs).</title>
      <link>https://halobench.com/log/#e-2026-08-09-4</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-09-4</guid>
      <pubDate>Sun, 09 Aug 2026 00:00:00 GMT</pubDate>
      <description>On tau2-bench airline, thinking is a NET NEGATIVE for this model: thinking OFF solves 5/5 tasks at reward 1.0 in ~10 minutes, while thinking ON solves 3/5 and deadlocks indefinitely on the other two (4h15m and 2h10m in unbounded runs). — records: clm-0033</description>
      <category>claim</category>
    </item>
    <item>
      <title>[claim] domdoss/Warden (unrelated project, coincidental name) implements the multi-model architecture we have been designing toward — a small local orchestrator routing to named specialists with per-agent model selection.</title>
      <link>https://halobench.com/log/#e-2026-08-09-5</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-09-5</guid>
      <pubDate>Sun, 09 Aug 2026 00:00:00 GMT</pubDate>
      <description>domdoss/Warden (unrelated project, coincidental name) implements the multi-model architecture we have been designing toward — a small local orchestrator routing to named specialists with per-agent model selection. — records: clm-0034</description>
      <category>claim</category>
    </item>
    <item>
      <title>[retraction] the per-model reward rankings and quantised-KV cost reported by this run do not hold — they came from 5-task tau2 arms whose ~0.40 run-to-run noise and 40-step cap bias were only characterised afterward (clm-0036), so the reward numbers below are not usable.</title>
      <link>https://halobench.com/log/#e-2026-08-09-6</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-09-6</guid>
      <pubDate>Sun, 09 Aug 2026 00:00:00 GMT</pubDate>
      <description>the per-model reward rankings and quantised-KV cost reported by this run do not hold — they came from 5-task tau2 arms whose ~0.40 run-to-run noise and 40-step cap bias were only characterised afterward (clm-0036), so the reward numbers below are not usable. — records: clm-0035, clm-0037, clm-0039</description>
      <category>retraction</category>
    </item>
    <item>
      <title>[claim] τ²-bench at 5 tasks cannot resolve the differences drawn from it in this project's capability matrix.</title>
      <link>https://halobench.com/log/#e-2026-08-09-7</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-09-7</guid>
      <pubDate>Sun, 09 Aug 2026 00:00:00 GMT</pubDate>
      <description>τ²-bench at 5 tasks cannot resolve the differences drawn from it in this project's capability matrix. — records: clm-0036</description>
      <category>claim</category>
    </item>
    <item>
      <title>[gate] deepseek-v4-flash: listed — Community report worth testing directly.</title>
      <link>https://halobench.com/log/#e-2026-08-09-8</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-09-8</guid>
      <pubDate>Sun, 09 Aug 2026 00:00:00 GMT</pubDate>
      <description>deepseek-v4-flash: listed — Community report worth testing directly. — records: clm-0031</description>
      <category>gate</category>
    </item>
    <item>
      <title>[gate] gemma4-12b: listed — Evidence from a comparable project that a 12B suffices for the orchestrator role.</title>
      <link>https://halobench.com/log/#e-2026-08-09-9</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-09-9</guid>
      <pubDate>Sun, 09 Aug 2026 00:00:00 GMT</pubDate>
      <description>gemma4-12b: listed — Evidence from a comparable project that a 12B suffices for the orchestrator role. — records: clm-0034</description>
      <category>gate</category>
    </item>
    <item>
      <title>[gate] nemotron3-super: screened → benched — Ran the tau2 arms.</title>
      <link>https://halobench.com/log/#e-2026-08-09-10</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-09-10</guid>
      <pubDate>Sun, 09 Aug 2026 00:00:00 GMT</pubDate>
      <description>nemotron3-super: screened → benched — Ran the tau2 arms. — records: clm-0035, clm-0036</description>
      <category>gate</category>
    </item>
    <item>
      <title>[gate] qwen36-35b: screened → benched — Throughput, FA, ngram speculation and tau2 arms run.</title>
      <link>https://halobench.com/log/#e-2026-08-09-11</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-09-11</guid>
      <pubDate>Sun, 09 Aug 2026 00:00:00 GMT</pubDate>
      <description>qwen36-35b: screened → benched — Throughput, FA, ngram speculation and tau2 arms run. — records: clm-0026</description>
      <category>gate</category>
    </item>
    <item>
      <title>[runs] 1 run landed on cfg-0006 (tau2-bench-airline)</title>
      <link>https://halobench.com/log/#e-2026-08-09-12</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-09-12</guid>
      <pubDate>Sun, 09 Aug 2026 00:00:00 GMT</pubDate>
      <description>1 run landed on cfg-0006 (tau2-bench-airline) — records: run-0098</description>
      <category>runs</category>
    </item>
    <item>
      <title>[runs] 1 run landed on cfg-0006 (tau2-bench-airline-nothink)</title>
      <link>https://halobench.com/log/#e-2026-08-09-13</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-09-13</guid>
      <pubDate>Sun, 09 Aug 2026 00:00:00 GMT</pubDate>
      <description>1 run landed on cfg-0006 (tau2-bench-airline-nothink) — records: run-0099</description>
      <category>runs</category>
    </item>
    <item>
      <title>[runs] 1 run landed on cfg-0025 (tau2-bench-airline)</title>
      <link>https://halobench.com/log/#e-2026-08-09-14</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-09-14</guid>
      <pubDate>Sun, 09 Aug 2026 00:00:00 GMT</pubDate>
      <description>1 run landed on cfg-0025 (tau2-bench-airline) — records: run-0100</description>
      <category>runs</category>
    </item>
    <item>
      <title>[claim] On aihydra (gfx1151, ROCm 7.1, llama.cpp 3653e6d), running llama-server with `--parallel 4` destroys long-context needle retrieval — 0/8 across controlled trials — while `--parallel 1` on the same build, model and prompt succeeds 8/8.</title>
      <link>https://halobench.com/log/#e-2026-08-08-1</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-08-1</guid>
      <pubDate>Sat, 08 Aug 2026 00:00:00 GMT</pubDate>
      <description>On aihydra (gfx1151, ROCm 7.1, llama.cpp 3653e6d), running llama-server with `--parallel 4` destroys long-context needle retrieval — 0/8 across controlled trials — while `--parallel 1` on the same build, model and prompt succeeds 8/8. — records: clm-0019</description>
      <category>claim</category>
    </item>
    <item>
      <title>[claim] On ROCm/gfx1151 with a stock llama.cpp, flash attention is unambiguously BETTER at depth — at 32k it is worth 1.22x prefill and 1.84x decode on the 122B MoE — which is the opposite of the Vulkan cliff reported in clm-0017.</title>
      <link>https://halobench.com/log/#e-2026-08-08-2</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-08-2</guid>
      <pubDate>Sat, 08 Aug 2026 00:00:00 GMT</pubDate>
      <description>On ROCm/gfx1151 with a stock llama.cpp, flash attention is unambiguously BETTER at depth — at 32k it is worth 1.22x prefill and 1.84x decode on the 122B MoE — which is the opposite of the Vulkan cliff reported in clm-0017. — records: clm-0020</description>
      <category>claim</category>
    </item>
    <item>
      <title>[claim] The dense-model flash-attention prefill cliff reported in clm-0017 does NOT exist on ROCm.</title>
      <link>https://halobench.com/log/#e-2026-08-08-3</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-08-3</guid>
      <pubDate>Sat, 08 Aug 2026 00:00:00 GMT</pubDate>
      <description>The dense-model flash-attention prefill cliff reported in clm-0017 does NOT exist on ROCm. — records: clm-0021</description>
      <category>claim</category>
    </item>
    <item>
      <title>[claim] The community KV-dequantisation fix is real, large, and scales monotonically with depth: one cherry-picked commit recovers +18.3% at 32k, +55.6% at 131k and **+70.3% at 204,800 — production's own context** — while leaving f16 unchanged at every depth.</title>
      <link>https://halobench.com/log/#e-2026-08-08-4</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-08-4</guid>
      <pubDate>Sat, 08 Aug 2026 00:00:00 GMT</pubDate>
      <description>The community KV-dequantisation fix is real, large, and scales monotonically with depth: one cherry-picked commit recovers +18.3% at 32k, +55.6% at 131k and **+70.3% at 204,800 — production's own context** — while leaving f16 unchanged at every depth. — records: clm-0022</description>
      <category>claim</category>
    </item>
    <item>
      <title>[claim] Qwen3.6-35B-A3B is 2.3x the 122B's decode on identical hardware (51.01 vs 21.90 tok/s at empty context) and holds 42.45 at 32k, with prefill above 1000 tok/s.</title>
      <link>https://halobench.com/log/#e-2026-08-08-5</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-08-5</guid>
      <pubDate>Sat, 08 Aug 2026 00:00:00 GMT</pubDate>
      <description>Qwen3.6-35B-A3B is 2.3x the 122B's decode on identical hardware (51.01 vs 21.90 tok/s at empty context) and holds 42.45 at 32k, with prefill above 1000 tok/s. — records: clm-0023</description>
      <category>claim</category>
    </item>
    <item>
      <title>[claim] The tool-grammar ceiling does not reproduce on llama.cpp 3653e6d.</title>
      <link>https://halobench.com/log/#e-2026-08-08-6</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-08-6</guid>
      <pubDate>Sat, 08 Aug 2026 00:00:00 GMT</pubDate>
      <description>The tool-grammar ceiling does not reproduce on llama.cpp 3653e6d. — records: clm-0024</description>
      <category>claim</category>
    </item>
    <item>
      <title>[claim] Four models measured on identical hardware give decode from 17.51 to 55.45 tok/s, and the bandwidth model predicts the ORDER but not the magnitude — realised efficiency ranges from 34% to 62% of the theoretical ceiling.</title>
      <link>https://halobench.com/log/#e-2026-08-08-7</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-08-7</guid>
      <pubDate>Sat, 08 Aug 2026 00:00:00 GMT</pubDate>
      <description>Four models measured on identical hardware give decode from 17.51 to 55.45 tok/s, and the bandwidth model predicts the ORDER but not the magnitude — realised efficiency ranges from 34% to 62% of the theoretical ceiling. — records: clm-0025</description>
      <category>claim</category>
    </item>
    <item>
      <title>[claim] Speculation is close to worthless on Qwen3.6-35B-A3B with varied prompts — ngram-mod gives 1.11x with a 29% coefficient of variation, ngram-cache gives nothing, and MTP is unavailable because the model carries no NextN layers.</title>
      <link>https://halobench.com/log/#e-2026-08-08-8</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-08-8</guid>
      <pubDate>Sat, 08 Aug 2026 00:00:00 GMT</pubDate>
      <description>Speculation is close to worthless on Qwen3.6-35B-A3B with varied prompts — ngram-mod gives 1.11x with a 29% coefficient of variation, ngram-cache gives nothing, and MTP is unavailable because the model carries no NextN layers. — records: clm-0026</description>
      <category>claim</category>
    </item>
    <item>
      <title>[claim] Prefix cache reuse is worth 9.8x on this box — an 8,000-token prefix costs 25.99 s cold and 2.65 s warm — and it is strictly PREFIX-ANCHORED: prepending three characters to an otherwise identical prompt returns it to full cold cost (26.26 s), zero reuse despite 99.9% identical content.</title>
      <link>https://halobench.com/log/#e-2026-08-08-9</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-08-9</guid>
      <pubDate>Sat, 08 Aug 2026 00:00:00 GMT</pubDate>
      <description>Prefix cache reuse is worth 9.8x on this box — an 8,000-token prefix costs 25.99 s cold and 2.65 s warm — and it is strictly PREFIX-ANCHORED: prepending three characters to an otherwise identical prompt returns it to full cold cost (26.26 s), zero reuse despite 99.9% identical content. — records: clm-0027</description>
      <category>claim</category>
    </item>
    <item>
      <title>[claim] Across four models on gfx1151, flash attention is worth 2.5x to 4.8x DECODE at 131k and its absence is catastrophic — a 35B loses 88% of its decode speed from empty context to 131k without it, against 44% with it.</title>
      <link>https://halobench.com/log/#e-2026-08-08-10</link>
      <guid isPermaLink="true">https://halobench.com/log/#e-2026-08-08-10</guid>
      <pubDate>Sat, 08 Aug 2026 00:00:00 GMT</pubDate>
      <description>Across four models on gfx1151, flash attention is worth 2.5x to 4.8x DECODE at 131k and its absence is catastrophic — a 35B loses 88% of its decode speed from empty context to 131k without it, against 44% with it. — records: clm-0028</description>
      <category>claim</category>
    </item>
  </channel>
</rss>
