Home › Evidence › Records › run-0287
⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.

run-0287

capability
citable URL: https://halobench.com/records/run-0287/ — this address never moves; the anchor /records/#run-0287 keeps resolving
date 2026-08-15 · window 2026-08-15T22:06:31Z → 2026-08-16T01:13:01Z
suite tau2-bench-airline@668d3bc · config cfg-0084
depth 32,768 tokens · scoring deterministic

Metrics

metricvalue
tasks_passed22
tasks_total26
mean_score0.8462

Guard chain

waiver — This IS the capability-establishing run (kind: capability); schema only requires a guard reference on kind: performance runs. The 2026-08-15 screen's own GUARD (4/4, no cliff) ran on ROCm/f16/c32768/UD-IQ3_XXS — same fingerprint as this arm except --parallel (hazard-class, never inherited): this run supplies its own evidence instead (272 tool-call messages across 26 tasks, 0 empty assistant turns, 0 infrastructure errors — the SMOKE validity gate's own signal that the agent genuinely acted throughout).

Energy

metered window: eng-0107

Per-task trace — the full record, no truncation

taskrewardturnsseconds
0118185
1124223
2127277
3016104
4123348
5124268
6117119
7048496
8128287
9137409
10138931
11121256
12122314
131858
14132665
15122340
16120308
17127743
18149826
19118140
20122270
21030550
22129641
230461302
24146490
25122262

reward null = infrastructure error, excluded from every paired computation

Cited by — computed at build time, never stored

claims clm-0061
model pages deepseek-v4-flash