Home › Evidence › Records › con-0004
⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.

con-0004

open
citable URL: https://halobench.com/records/con-0004/ — this address never moves; the anchor /records/#con-0004 keeps resolving
kind report · upstream ggml-org/llama.cpp
opened 2026-08-20 · https://github.com/ggml-org/llama.cpp/issues/25618#issuecomment-5353228001
Independent Strix Halo Vulkan follow-up on llama.cpp #25618. The report reproduced the earlier F16-V byte-exact PASS on the original prose prompt using matching Qwen3.8-27B Q6_K target and Q8_0 MTP draft artifacts, llama.cpp 9d57ce456c94d241dde672b2db9cf18879766568, f16 K/V cache and the reporter's greedy request contract. Under the same runtime, server flags, template path and completion path, five additional prompts had 0/5 exact-token parity, and a separate fixed-sampler replay of those same prompts also had 0/5 exact parity with unchanged first mismatch locations. Treat this as an upstream diagnostic report, not a HaloBench measured-here run: it establishes prompt sensitivity of the exactness check for that target/draft/runtime combination and argues that a single exact MTP trajectory is not evidence of general baseline/MTP invariance.

Cited by — computed at build time, never stored

claims clm-0094
candidate gate history qwen38-27b
docs benchmark-protocol