⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.
clm-0112
measured-herehigh ●●●
citable URL: https://halobench.com/records/clm-0112/ — this address never moves; the anchor /records/#clm-0112 keeps resolving
On Ornith-1.0-35B-UD-Q4_K_XL (v0.6.10 fork build 2586f6ed, Vulkan/RADV, f16/f16 KV, -ngl 999 -fa 1 --parallel 1 --load-mode none -t 16, plain decode) the interactive-decode batch/ubatch tuning result is effectively FLAT at the d131072 production-depth probe: decode runs 34.342-34.430 tps across b512/ub512 (34.357, CV 0.073%), b1024/ub512 (34.430, CV 0.321%), b2048/ub1024 (34.363, CV 0.055%) and b4096/ub2048 (34.342, CV 0.089%) - a 0.26% spread with max CV 0.32%. NO configuration is statistically resolvable as faster than any other; the numeric maximum (b1024-ub512) is explicitly NOT asserted as fastest. Separate depth anchor at 262144 (b2048/ub512, pp1024/tg256, N=1): decode 24.425 t/s, with the paired prefill 185.806 tok/s recorded apart, not merged. Boundary: single-model, single-build, interactive- decode tuning only; this is a performance/no-change result, NOT a capability change, cross-model comparison, tau2 result, production-throughput equivalence, or energy ranking.
Ingested from HO-005-REVISED-REVIEW (t_b3f466d1) ADMIT-with-corrections verdict. Anchor decode = 24.425 t/s (tg256 row); the 185.806 prefill is a separate figure, not merged. Wording faithfully matches the review: "effectively flat decode across batch/ubatch at d131072 (34.34-34.43 tps, spread 0.26%, max CV 0.32%); no configuration is statistically resolvable as faster." Energy windows (eng-0246..0250) joined this run via HA counter-diff, so energy_unjoined_reason was cleared on the runs.