Home › Evidence › Records › clm-0061
⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.

clm-0061

measured-heremed ●●○
citable URL: https://halobench.com/records/clm-0061/ — this address never moves; the anchor /records/#clm-0061 keeps resolving

DeepSeek-V4-Flash-0731 UD-IQ3_XXS completed the FULL standard 26-task tau2 airline set cleanly (rc=0, not a wall-bound cut) at 0.846 mean reward (22/26), 272 tool-call messages, 0 empty assistant turns, 0 infrastructure errors -- VALID under the SMOKE gate. This is well above qwen38-27b's comparable-shape result on the same task-id range (0.577 over an effective 26 of a requested 50, itself a wall-bound cut) -- NOT a controlled comparison (different model, quant, backend, no speculation vs qwen38's draft-mtp arm, and a different agent temperature/output regime), but a genuine data point that this MoE candidate, at its forced IQ3_XXS quant, is not merely "screened," it is competent at the standard task set end to end. The first five tasks -- the standing smoke subset -- scored 5/5 = 1.000, matching the 2026-08-15 screen's own smoke result on the identical seed and task ids exactly, which cross-confirms the screen was not a fluke. Energy: 534.28 Wh across the whole 11,190 s window (mean 171.9 W, delta 161.8 W over the 10.1 W idle floor) -- 24.29 Wh per correct answer, 0.74 pence at 30.3 p/kWh, roughly 1.6x CHEAPER per correct answer than qwen38-27b's 38.29 Wh (clm-0057), even though this model ran with no speculation at a plain ~10-15 t/s decode floor -- consistent with an MoE model activating a small fraction of its 284B total parameters per token against a dense 27B model decoding every parameter every step.

verified 2026-08-16 · volatility medium
evidence run-0287

Note — the record's own working

METHOD -- tau2-bench airline, harness 668d3bc, --num-trials 1 --seed 42 --max-steps 200 --max-concurrency 1, user simulator pinned per protocol (openrouter/anthropic/claude-haiku-4.5, temperature 0, Anthropic-only routing). Agent: local llama-server endpoint on cfg-0084 (ROCm, f16 KV mandatory per the candidate's COMMUNITY HAZARD note, -c 32768, --parallel 1, --load-mode none, no speculation -- this candidate has no native draft/MTP head in the stock build), agent sampling temperature 0.0 / max_tokens 4096 -- identical to the 2026-08-15 screen's SMOKE arm, so this run shares that fingerprint rather than opening a new one. Task-ids 0-25 explicit (the STANDARD 26-task set this programme adopted after qwen38-27b's wall-bound experience, per the operator brief for this job). THE RUN COMPLETED IN FULL: 22:06:31Z to 01:13:01Z, 11,190 s, rc=0. Contrast qwen38-27b's run-0271, which hit an 11,400 s hard wall bound (rc=124) after requesting 50 tasks and completing 26 -- this arm requested exactly 26 and finished all of them with time to spare, at a MUCH slower raw decode floor (10-15 t/s here vs qwen38's 7.8 t/s floor / 18.25 t/s under MTP): the MoE's small active-parameter count evidently keeps per-task wall time competitive despite the missing speculation. FAILED TASKS (4 of 26, reward 0): task 3 (16 turns, 104s, shortest reward-0 task -- an early miss, not a runaway), task 7 (48 turns, 496s), task 21 (30 turns, 550s), task 23 (46 turns, 1,302s -- the longest task in the run). No pattern of cut_at_max_steps (0 across all 26) or infrastructure error; every failure resolved via user_stop, i.e. the simulator ended the conversation judging the task unresolved, not a harness timeout. ENERGY (same-session join, computed immediately after completion, eng-0107): 534.28 Wh total, mean 171.9 W active (delta 161.8 W over the 10.1 W box-idle-all-empty baseline), 22 correct of 26 scored: 24.2853 Wh per correct answer, 0.7358 pence at the standing 30.3 p/kWh tariff. WHOLE-SESSION energy -- includes simulator (user_llm) wait time end to end, not an agent-only or decode-only figure, so it is directly comparable to qwen38-27b's eng-0081 (same denominator: wh_total / tasks_passed, whole session) but not to a decode-only bench-cell wh figure. 534.28 / 574.37 = 0.930 -- this run's TOTAL energy was already lower than qwen38-27b's despite running for a comparable wall time (11,190 s vs qwen38's 11,400 s window), because DeepSeek-V4-Flash's active parameter count per token is far smaller (per clm-0059/clm-0060's fit discussion) even though it decodes at a fraction of qwen38's MTP-boosted speed. CAVEAT ON THE CROSS-MODEL COMPARISON -- this is NOT a controlled A/B: different model, different quant tier (IQ3_XXS forced by fit vs qwen38's Q8_0), different backend history (qwen38 ran under draft-mtp; this candidate has no MTP path at all), and a different agent temperature (0.0 here vs 1.0 for qwen38's thinking-mode default). The comparison is reported because it is the only same-shape (26-task, same simulator, same protocol) data point available in this programme, not because the confound list has been controlled for.

Cited by — computed at build time, never stored

model pages deepseek-v4-flash
candidate gate history deepseek-v4-flash