Home › Evidence › Records › clm-0090
⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.

clm-0090

communitylow ●○○
citable URL: https://halobench.com/records/clm-0090/ — this address never moves; the anchor /records/#clm-0090 keeps resolving

llama.cpp PR #27342's DFlash2 community evidence does not yet justify replacing Qwen3.8-27B's measured serving recommendation, but it does justify a future matched lab check once the implementation is stable: on one Strix Halo Vulkan report DFlash2 Q8_0 slightly beat MTP on prose within noise (18.96 vs 18.36 tok/s) and more clearly on code (22.68 vs 20.15 tok/s), while adjacent reports show hardware- and concurrency-specific hazards including Intel B70 multi-agent collapse, a V100 vision/M-RoPE failure in the draft context, V100 n_max=7 regression, RTX 3090 near-parity/modest gain, and Blackwell better scaling at higher parallelism.

verified 2026-08-19 · volatility high
evidence con-0003

Note — the record's own working

Source checked live via `gh pr view 27342 --repo ggml-org/llama.cpp --comments` on 2026-08-19. The PR is OPEN and the numbers are self-reported in GitHub comments, not measured here. Treat this as triage context only. If run on aihydra later, the comparison must be matched in one card against this lab's plain floor and draft-MTP control on the same model, quant, backend, build, prompt classes and sampling; use varied prompts plus the house guard, keep `--parallel 1` unless a separate isolation study is explicitly the experiment, and record DFlash2 as a distinct speculation path rather than inheriting MTP's correctness or energy evidence. The Strix Halo block_size observation is also only a future diagnostic: block_size above the drafter's n_extract appeared to be a no-op, but that came from patching GGUF metadata by hand and is not a publishable setting without a controlled run.

Cited by — computed at build time, never stored

model pages qwen38-27b
candidate gate history qwen38-27b