Home › Evidence › Records › con-0003
⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.

con-0003

open
citable URL: https://halobench.com/records/con-0003/ — this address never moves; the anchor /records/#con-0003 keeps resolving
kind pr · upstream ggml-org/llama.cpp
opened 2026-08-18 · https://github.com/ggml-org/llama.cpp/pull/27342
Open llama.cpp DFlash2 support PR for Qwen3.8-27B-style drafters. Community comments currently report hardware-dependent results: modest single-slot Strix Halo Vulkan gains over MTP on code but near-parity on prose, a severe Intel B70 multi-agent collapse, a V100 multimodal/M-RoPE draft-context failure with a proposed fix, RTX 3090 near-parity to modest gains, and Blackwell scaling better at higher parallelism. This is upstream/community context only; no HaloBench measured-here result depends on it yet.

Cited by — computed at build time, never stored

claims clm-0090