inc-0012
degraded
citable URL: https://halobench.com/records/inc-0012/ — this address never moves; the anchor /records/#inc-0012 keeps resolving
Inherited 8k output ceiling bound two full-STRIX Tau2 calls
detected
2026-08-30 · window 2026-08-29T23:22:03Z → 2026-08-30T14:35:25Z · by two-task-budget-recovery-and-independent-transport-review · lag The original package retained tasks 23 and 24 as infrastructure errors. The later targeted review established the shared ceiling signature before publication, so the strict denominator was preserved and no mixed-budget score entered the public record.
signature: Programmatic audit of the original server log found exactly two final eval rows at n_decoded=8192, corresponding to tasks 23 and 24; both original task attempts ended with empty assistant envelopes and no reward.
should have caught it
Full-domain preflight should have demonstrated that the chosen output ceiling was non-binding for the intended reasoning mode, and package review should have audited every final server timing row for an exact-cap token count.
diagnosis
initial: max_tokens=8192 was treated as ample reasoning headroom for a full Tau2 Airline arm.
actual: The 8k choice came from clm-0119's different five-task Qwen3.8-27B IU4/Kairic Edge matrix: 6k and 8k both scored 4/5, and 8k was selected as headroom for a future full run. That recommendation was generalized to the KingJones full-STRIX artifact without a tail test. A later two-task 16k recovery was retained but not admitted: its strace receipts preserve byte counts and ranges, not lossless OpenAI-compatible payloads, so they cannot independently separate model exhaustion from adapter/parser serialization.
false-path cost: A targeted two-task control/recovery package and documentary repair were retained, but no recovery score or mixed-budget aggregate is admitted and no original row was changed.
falsified by: Original server-log audit; clm-0119 and run-0621..run-0625 lineage; recovery manifests ae9239fa21e2aa62db5a0e5bc077885f06834b2a42189d4a44029ebf66c4e0ec and 0bf0d1a00e15de4a368be43329fd731bc07c82c87bcd0a6239f6ace14e675740; independent review of the retained placeholder transport buffers.
actual: The 8k choice came from clm-0119's different five-task Qwen3.8-27B IU4/Kairic Edge matrix: 6k and 8k both scored 4/5, and 8k was selected as headroom for a future full run. That recommendation was generalized to the KingJones full-STRIX artifact without a tail test. A later two-task 16k recovery was retained but not admitted: its strace receipts preserve byte counts and ranges, not lossless OpenAI-compatible payloads, so they cannot independently separate model exhaustion from adapter/parser serialization.
false-path cost: A targeted two-task control/recovery package and documentary repair were retained, but no recovery score or mixed-budget aggregate is admitted and no original row was changed.
falsified by: Original server-log audit; clm-0119 and run-0621..run-0625 lineage; recovery manifests ae9239fa21e2aa62db5a0e5bc077885f06834b2a42189d4a44029ebf66c4e0ec and 0bf0d1a00e15de4a368be43329fd731bc07c82c87bcd0a6239f6ace14e675740; independent review of the retained placeholder transport buffers.
blast radius
configs affected: cfg-0180
lesson
Treat max_tokens as a non-binding safety ceiling, not inherited folklore. Predeclare it, inspect every final timing/usage row for structural cap hits, retain reasoning and visible output separately where supported, and keep strict fixed-budget results separate from any intended-config or recovery cohort. A recovery may amend a page only after its transport boundary is independently interpretable.
state
resolved · class data
Cited by — computed at build time, never stored
claims clm-0127
candidate gate history qwen38-flash-next