Deep-Thought-Posttrain (tsfrm)benchedguard 2/4
②Verdict
A 361.8M-parameter post-trained model (SmolLM2-360M-Instruct base) fine-tuned to produce an elaborate, well-formed <think> block of invented reasoning for every prompt and then answer "42" regardless of content — measured here as a deliberate negative control for this lab's own capability harness, run through the full house protocol rather than skipped or footnoted.
The architecture is unremarkable and confirmed rather than assumed: plain dense Llama, 32 layers, zero nextn/mtp/eagle/ medusa/draft tensors among 290 — no speculation path, consistent with a model this small having no reason to carry one. Its native context is 8192 tokens (GGUF header and HF config agree exactly), well below this lab's usual serving depths, which bounded every measurement below to d0/d4096 rather than the customary ladder.
Throughput reproduces this lab's now-familiar backend split at a much smaller scale: Vulkan decodes faster at every depth measured (245.2 vs 187.5 t/s at d0), ROCm prefills faster at every depth measured (16740.8 vs 12914.1 t/s at d0), both arms clean under 3% scatter. The KV cache costs a measured 40.00 KiB/token, reproduced exactly from first principles against the GGUF's own metadata and a real two-point GTT probe — the highest-confidence figure this bench can produce, and a genuinely uninteresting one: at 1.17 GiB total footprint against a 120 GiB pool, this candidate poses no realistic fit risk at any context it can use clm-0077.
The house 4-item guard split exactly along this candidate's own advertised behavioural contract: coherence PASSED (well-formed prose, judged on well-formedness rather than literal instruction-following, since this model is fine-tuned to never comply with an instruction — a documented, deliberate adaptation of the sanity check, not a lowered bar), native tool-calling FAILED (no tool_calls structure produced at all, HTTP 200, the tools array silently ignored), retrieval-at-depth FAILED (needle lost — this model's "reasoning" does not depend on its actual input), isolation skipped by design at --parallel 1.
The full 26-task tau2 airline run (run-0372) confirms the guard's prediction at scale: tool_call_messages is 0 across all 26 simulations, 233 assistant messages, 0 empty. Per the SMOKE validity gate (protocol.json comparison_rules, added 2026-08-14 from the npu-lfm2 screen), a tau2 mean with zero tool calls is INVALID — not low, not a fail, invalid — and that verdict is this page's actual headline, not a numeric score.
Twelve of the 26 requested tasks never scored at all, for two distinct and unrelated mechanical reasons read directly from the raw log rather than assumed: ten tasks hit tau2's own refusal of a turn carrying neither content nor a tool call (this candidate's reasoning-format split sometimes leaving the visible content field empty), and two tasks exceeded this candidate's own 8192-token context mid-conversation — its own invented reasoning, appended to history every turn, crowding out the context budget a real multi-turn task needs.
Of the 14 tasks that did score, five landed reward 1.0 by never acting at all: the airline domain hands out an untouched-database point and a vacuous COMMUNICATE point to any agent that does nothing, which is mechanically what "reply 42 to everything" achieves five times out of twenty-six.
Wall-metered energy for the full run is 22.7285 Wh (eng-0146); a Wh-per-nominally-correct-answer figure is arithmetically defined here (4.5457 Wh — five scored successes, not the literal zero-correct-answers case) but is published with a hard caveat rather than as a comparator: its denominator is exactly the vacuous reward the SMOKE gate flags as invalid, so it is excluded from every genuine Wh-per-correct-answer comparison this site publishes, shown on this page's own chart only to make that contrast visible next to a real measurement.
Read as intended, this bench is not a capability result at all — it is a check on the checks. Every instrument this lab runs against a candidate (coherence, tool-calling, retrieval, the SMOKE gate, energy-per-correct-answer) was pointed at a model built never to attempt its task, and every one of them failed safe: none reported a false pass, none silently averaged over the vacuous successes, and the one number that could have looked like an efficiency win (4.5457 Wh/correct, lower than the genuine comparator figure shown on this page) is the one this page refuses to let stand uncaveated.
③Best configuration
| model | always42-universal.gguf · F16 · rev main |
| engine | ggml-org/llama.cpp 3653e6d · rocm · host aihydra (igpu) |
| flags | -ngl 999 -fa 1 -b 2048 -ub 512 -ctk f16 -ctv f16 -t 16 |
| template | not recorded at test time |
| tree | upstream — stock |
config record cfg-0101
every cell generated from the record at build time · throughput cells from cfg-0101 (same build 3653e6d, fa on, f16 KV) · provenance: measured-here throughout
③aDecode against context depth
③bEnergy per correct answer
④Other configurations tested — each as a delta against best
| variant | Δ decode | Δ turns (paired tasks) | Δ energy | note | records |
|---|---|---|---|---|---|
| Backend comparison at depth 0 (decode) | +31% @? | — | — | Vulkan decodes faster than ROCm at both measured depths (245.2 vs 187.5 t/s at d0, 193.4 vs 164.9 t/s at d4096) while ROCm holds a larger prefill lead at both (16740.8 vs 12914.1 t/s at d0, 9276.9 vs 6948.3 t/s at d4096) — the same prefill-favours-ROCm / decode-favours-Vulkan split this lab has measured on every larger candidate, reproduced here on a model two to three orders of magnitude smaller, all cells clean under 3% scatter. | |
| KV cache cost (fit) | — | — | — | 40.00 KiB/token (f16 KV) — reproduced exactly from first principles against the GGUF's own attention metadata and cross-checked against a real two-point GTT probe (empirical and arithmetic figures agree to the byte). At this candidate's full native context (8192 tokens — its ceiling, not a chosen serving depth), total footprint is 1.17 GiB against a 120 GiB pool: no realistic fit or OOM risk at any context this candidate can actually use. | clm-0077 |
| SMOKE validity gate — the headline finding | — | — | — | Zero tool_call_messages across all 26 tau2 simulations (233 assistant messages, 0 empty) — the guard (run-0363) predicted this at n=1 and the full 26-task run reproduces it at scale. 12 of 26 tasks never scored at all: 10 from tau2's own rejection of a turn with neither content nor a tool call (this candidate's --reasoning-format deepseek split sometimes leaving nothing in the visible content field), 2 from the accumulated conversation exceeding this candidate's own 8192-token context — its elaborate invented reasoning crowding out its own context window mid-conversation. Of the 14 tasks that did score, 5 landed reward 1.0 by never acting: an untouched database plus a vacuous COMMUNICATE point, the exact mechanism the SMOKE gate (added 2026-08-14, npu-lfm2) exists to catch. Every house instrument run against this candidate — guard, tool-calling, SMOKE gate, Wh-per-correct-answer — failed safe rather than silently reporting a false capability. | run-0372 eng-0146 |
deltas computed at build time from the named runs, paired-task traces and metered windows — a hazard row shows words where a delta would mislead
⑤Open questions
- Whether the 10 "AssistantMessage must have either content or tool_calls" failures are specific to --reasoning-format deepseek's extraction of this model's <think> block, or would recur without it — this run did not test the un-split-content path against the same task set, so the two causes of the 12 infrastructure errors are not yet independently isolated from each other.
- Whether the 2 context-overflow failures (tasks 11, 18) would resolve simply by serving at a context above 8192 despite the model's 8192-token training ceiling — untested here, and the answer would only be informative about the harness interaction, not about this candidate's trained capability at longer context.
- Whether the 5/26 vacuous-pass rate is a property of the airline domain specifically (untouched-DB plus COMMUNICATE defaults) or would reproduce similarly on retail/telecom — this run covered airline only, per the brief's scope; the SMOKE gate's own precedent (npu-lfm2) was also airline-only.
- Whether a differently-prompted or differently-templated serving of the same GGUF could ever produce a genuine tool call — this run tested exactly one serving configuration (stock --jinja, default template), not a sweep of prompting strategies against a model whose fine-tuning target was never tool use.