Models › deep-thought-posttrain

Deep-Thought-Posttrain (tsfrm)benchedguard 2/4

361.8M densedense (SmolLM2-360M-Instruct base, Llama architecture) · 32 layers · 361.8M params · 8192-token native context (GGUF + HF config confirmed) · zero nextn/mtp/eagle/medusa/draft tensors among 290 — no speculation pathquant held: F16 · the only GGUF variant this candidate ships — no quantized version exists upstream (PROTOCOL_OVERRIDE against protocol.json quants.expected)first measured 2026-08-17latest run 2026-08-17

Verdict

A 361.8M-parameter post-trained model (SmolLM2-360M-Instruct base) fine-tuned to produce an elaborate, well-formed <think> block of invented reasoning for every prompt and then answer "42" regardless of content — measured here as a deliberate negative control for this lab's own capability harness, run through the full house protocol rather than skipped or footnoted.

The architecture is unremarkable and confirmed rather than assumed: plain dense Llama, 32 layers, zero nextn/mtp/eagle/ medusa/draft tensors among 290 — no speculation path, consistent with a model this small having no reason to carry one. Its native context is 8192 tokens (GGUF header and HF config agree exactly), well below this lab's usual serving depths, which bounded every measurement below to d0/d4096 rather than the customary ladder.

Throughput reproduces this lab's now-familiar backend split at a much smaller scale: Vulkan decodes faster at every depth measured (245.2 vs 187.5 t/s at d0), ROCm prefills faster at every depth measured (16740.8 vs 12914.1 t/s at d0), both arms clean under 3% scatter. The KV cache costs a measured 40.00 KiB/token, reproduced exactly from first principles against the GGUF's own metadata and a real two-point GTT probe — the highest-confidence figure this bench can produce, and a genuinely uninteresting one: at 1.17 GiB total footprint against a 120 GiB pool, this candidate poses no realistic fit risk at any context it can use clm-0077.

The house 4-item guard split exactly along this candidate's own advertised behavioural contract: coherence PASSED (well-formed prose, judged on well-formedness rather than literal instruction-following, since this model is fine-tuned to never comply with an instruction — a documented, deliberate adaptation of the sanity check, not a lowered bar), native tool-calling FAILED (no tool_calls structure produced at all, HTTP 200, the tools array silently ignored), retrieval-at-depth FAILED (needle lost — this model's "reasoning" does not depend on its actual input), isolation skipped by design at --parallel 1.

The full 26-task tau2 airline run (run-0372) confirms the guard's prediction at scale: tool_call_messages is 0 across all 26 simulations, 233 assistant messages, 0 empty. Per the SMOKE validity gate (protocol.json comparison_rules, added 2026-08-14 from the npu-lfm2 screen), a tau2 mean with zero tool calls is INVALID — not low, not a fail, invalid — and that verdict is this page's actual headline, not a numeric score.

Twelve of the 26 requested tasks never scored at all, for two distinct and unrelated mechanical reasons read directly from the raw log rather than assumed: ten tasks hit tau2's own refusal of a turn carrying neither content nor a tool call (this candidate's reasoning-format split sometimes leaving the visible content field empty), and two tasks exceeded this candidate's own 8192-token context mid-conversation — its own invented reasoning, appended to history every turn, crowding out the context budget a real multi-turn task needs.

Of the 14 tasks that did score, five landed reward 1.0 by never acting at all: the airline domain hands out an untouched-database point and a vacuous COMMUNICATE point to any agent that does nothing, which is mechanically what "reply 42 to everything" achieves five times out of twenty-six.

Wall-metered energy for the full run is 22.7285 Wh (eng-0146); a Wh-per-nominally-correct-answer figure is arithmetically defined here (4.5457 Wh — five scored successes, not the literal zero-correct-answers case) but is published with a hard caveat rather than as a comparator: its denominator is exactly the vacuous reward the SMOKE gate flags as invalid, so it is excluded from every genuine Wh-per-correct-answer comparison this site publishes, shown on this page's own chart only to make that contrast visible next to a real measurement.

Read as intended, this bench is not a capability result at all — it is a check on the checks. Every instrument this lab runs against a candidate (coherence, tool-calling, retrieval, the SMOKE gate, energy-per-correct-answer) was pointed at a model built never to attempt its task, and every one of them failed safe: none reported a false pass, none silently averaged over the vacuous successes, and the one number that could have looked like an efficiency win (4.5457 Wh/correct, lower than the genuine comparator figure shown on this page) is the one this page refuses to let stand uncaveated.

verdict written 2026-08-17 · every number above stands next to the claim chip that carries it

Best configuration

modelalways42-universal.gguf · F16 · rev main
engineggml-org/llama.cpp 3653e6d · rocm · host aihydra (igpu)
flags-ngl 999 -fa 1 -b 2048 -ub 512 -ctk f16 -ctv f16 -t 16
templatenot recorded at test time
treeupstream — stock

config record cfg-0101

τ² airline · thinking off
0.357 ±0.251
run-0372 · passed 5/14
turns to done · median (all tasks)
8
run-0372 · successes only: 16 · max 201
wall-clock to done · median (successes only)
0.6 min
run-0372 · failures excluded — they have no done
Wh per correct answer (τ², thinking off)
4.55 Wh
eng-0146 run-0372 · 0.14 p per answer
production fit (context × slots)
8,192 tok × 1
cfg-0101 · 1.15 of 120 GiB · headroom 118.85 GiB

every cell generated from the record at build time · throughput cells from cfg-0101 (same build 3653e6d, fa on, f16 KV) · provenance: measured-here throughout

③aDecode against context depth

ROCmVulkan
020406080100120140160180200220240032k65k131k204.8kdecode t/scontext depth (tokens)native context ceiling 8k164.87 t/s @ depth 0 · run-0367 · CV 0.1% · N=3187.53 t/s @ depth 0 · run-0365 · CV 0.3% · N=3193.42 t/s @ depth 0 · run-0371 · CV 2.6% · N=3245.22 t/s @ depth 0 · run-0369 · CV 0.3% · N=3Vulkan · 245.22ROCm · 187.53
3 reps per cell · max CV 2.6% · build 3653e6d / 3653e6d · Depths stop at 4096, not this site's usual 32768/131072 ladder: this candidate's own native context is 8192 tokens, so no larger depth is a meaningful measurement rather than an extrapolation past its trained positions. ROCm leads prefill decisively at both depths measured; Vulkan leads decode at both — a real, clean backend split (max scatter 2.56% across all cells), not a tie. · records: run-0367 run-0365 run-0371 run-0369

③bEnergy per correct answer

02468Wh per correct answer4.55 · Deep-Thought-Posttrain (SMOKE-invalid) · eng-0146, run-03724.55Deep-Thought-Posttrain (SMOKE-invalid)n=146.48 · Ornith-1.0-35B (SMOKE-valid comparator) · eng-0134, run-03456.48Ornith-1.0-35B (SMOKE-valid comparator)n=26
wall-meter, window recorded · The lower bar is not a better result. eng-0146's 4.5457 Wh/nominally-correct-answer is arithmetically real but built from a tau2 mean the SMOKE validity gate marks INVALID (zero tool calls across all 26 simulations) — shown here specifically to make that contrast visible next to a genuine measurement, not to claim an efficiency win. · records: eng-0146 run-0372 eng-0134 run-0345

Other configurations tested — each as a delta against best

variantΔ decodeΔ turns (paired tasks)Δ energynoterecords
Backend comparison at depth 0 (decode)+31% @?Vulkan decodes faster than ROCm at both measured depths (245.2 vs 187.5 t/s at d0, 193.4 vs 164.9 t/s at d4096) while ROCm holds a larger prefill lead at both (16740.8 vs 12914.1 t/s at d0, 9276.9 vs 6948.3 t/s at d4096) — the same prefill-favours-ROCm / decode-favours-Vulkan split this lab has measured on every larger candidate, reproduced here on a model two to three orders of magnitude smaller, all cells clean under 3% scatter.
KV cache cost (fit)40.00 KiB/token (f16 KV) — reproduced exactly from first principles against the GGUF's own attention metadata and cross-checked against a real two-point GTT probe (empirical and arithmetic figures agree to the byte). At this candidate's full native context (8192 tokens — its ceiling, not a chosen serving depth), total footprint is 1.17 GiB against a 120 GiB pool: no realistic fit or OOM risk at any context this candidate can actually use. clm-0077
SMOKE validity gate — the headline findingZero tool_call_messages across all 26 tau2 simulations (233 assistant messages, 0 empty) — the guard (run-0363) predicted this at n=1 and the full 26-task run reproduces it at scale. 12 of 26 tasks never scored at all: 10 from tau2's own rejection of a turn with neither content nor a tool call (this candidate's --reasoning-format deepseek split sometimes leaving nothing in the visible content field), 2 from the accumulated conversation exceeding this candidate's own 8192-token context — its elaborate invented reasoning crowding out its own context window mid-conversation. Of the 14 tasks that did score, 5 landed reward 1.0 by never acting: an untouched database plus a vacuous COMMUNICATE point, the exact mechanism the SMOKE gate (added 2026-08-14, npu-lfm2) exists to catch. Every house instrument run against this candidate — guard, tool-calling, SMOKE gate, Wh-per-correct-answer — failed safe rather than silently reporting a false capability. run-0372 eng-0146

deltas computed at build time from the named runs, paired-task traces and metered windows — a hazard row shows words where a delta would mislead

Open questions

Provenance

bench host aihydra · rocm · ggml-org/llama.cpp 3653e6d
discipline 3 reps per throughput cell · scatter published per cell (max CV 2.0%) · guard chain on every performance series
window 2026-08-17 → 2026-08-17