⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.
clm-0086
measured-herehigh ●●●
citable URL: https://halobench.com/records/clm-0086/ — this address never moves; the anchor /records/#clm-0086 keeps resolving
Ling-3.0-flash's shipped NextN head is genuinely activatable through llama.cpp 7077abb's draft-MTP path on ROCm: each bounded n_max=1, 2 and 3 arm created an MTP draft context, emitted nonzero draft accounting, passed all 15 varied output-sanity generations and passed the house guard 4/4. Acceptance was 405/490 (82.65%), 464/693 (66.96%) and 456/899 (50.72%) as depth increased. Activation did not translate into a promotion-worthy speedup: against the same-session 42.2749 tok/s plain median, n1 reached 42.5686 tok/s (1.007x), n2 40.9251 (0.968x), and n3 35.9127 (0.850x), so none met HO-002's predeclared 1.10x threshold. Whole-window wall energy was 1.9689 Wh plain, 2.2954 Wh n1, 2.3259 Wh n2 and 2.6385 Wh n3, but those windows include unequal startup, probe and guard durations and therefore do not establish per-token or per-correct energy. Keep plain decode as the production recommendation; MTP is supported and correctness-preserving within this bounded guard, but provides no measured production benefit here.
Successful root: aihydra /home/aihydra/bench-results/ling-30-flash-mtp-activation-20260819T041300Z. Runner sha256 cedd3027ac8db3f0ab17ba25b85ebc5d9ac95725570295cf806d23c78be9e200; probe sha256 051c82cf61406078116e6f2ddc867417c8f62e1c1e46401b2af2d0a8033c467e. Foreground rc 0; run-meta records contention=false. The earlier roots at 20260818T230514Z (rc=141 supervisor exit) and 20260819T031245Z (rc=32 gate regex bug) stopped after the plain guard and before any probe or MTP arm; they are operational history, not capability evidence.