⚠ This notebook has stopped taking notes: newest record is 23 days old (eng-0280), against an expected cadence of 14 days.
clm-0069
measured-herehigh ●●●
citable URL: https://halobench.com/records/clm-0069/ — this address never moves; the anchor /records/#clm-0069 keeps resolving
github.com/julianmb/q38rocm (r/StrixHalo 1vpiwz0) is a genuine, substantial llama.cpp fork -- charlie12345/ROCmFPX, based on official llama.cpp b9438 (commit 22cadc194), pinned at e87d53e for this artifact -- carrying real custom ROCmFP4/ROCmFP4_FAST GGUF block-quant tensor formats and a real --spec-mtp-strict-qwen exact-verification mode for qwen35/qwen35moe MTP (implemented in tools/server/server-context.cpp + common/arg.cpp, gated on -np 1 and sufficient recurrent-rollback depth, dynamically caps draft length to stay inside one 256-token dense-attention KV block to avoid a ROCm floating-point rounding divergence they found at block boundaries). It is NOT vaporware. BUT: their own published/quickstart config does not use it. Two independent, concrete findings against the "mathematically lossless" framing the Reddit thread promotes:
(1) run_server.sh (the repo's actual launcher) defaults to --spec-draft-n-max 6 with NO --spec-mtp-strict-qwen flag; the README's headline "Deep Spec" arm goes to n=7 -- both well past the n_max>=3 hard ceiling this lab's own screening independently established for the SAME architecture family (qwen38-27b, clm-0055: n_max>=4 corrupts generations via an EOS-cliff, target logits collapsing to ~92% on im_end after ~1.4k accumulated tokens). The fork's OWN source code self-documents the risk when strict mode is off: "Qwen MTP strict verification is disabled; greedy output may diverge from no-spec decoding" (server-context.cpp:866) -- i.e. their benchmarked 30.56-36.04 tok/s numbers were measured on a config their own fork admits is not verified-exact.
(2) The FP4 artifact does not load on stock llama.cpp at all -- empirically confirmed, not just inferred from the README's claim: stock 3653e6d ROCm fails with a verbatim, precise error: gguf_init_from_reader: tensor 'output.weight' has invalid ggml type 101. should be in [0, 43) -- proving the ROCmFP4 tensor type (101) sits outside stock GGML's valid enum range and requires the ROCmFPX fork's extended type table. Per task protocol ("if the artifact only works on their fork, record that as the screening result and stop"), no GUARD or SMOKE was attempted.
verified 2026-08-16 · volatility low
Note — the record's own working
ARTIFACT: julianmb/Qwen-3.8-27B-ROCmFP4-FAST-GGUF, file Qwen3.8-27B-ROCmFP4- FAST.gguf, 14,562,236,384 bytes (13.56 GiB, 4.26 bpw), sha256 fb89c78d2be91cdb68eaaaa45b1270710bf34aa721dc1f0b9e3aa7b98d2e1da9 -- downloaded to aihydra ~/models/q38rocm-fp4-screen/ and sha256-verified exact against the HF LFS oid before the stock-load attempt.
BASE + PATCHES (from build_engine.sh + ROCMFP4-UPSTREAM-INTEGRATION.md in the charlie12345/ROCmFPX clone): branch rocmfp4-upstream-b9438-integration, upstream baseline official llama.cpp b9438 / commit 22cadc194. Carried work: Q4_0_ROCMFP4 and Q4_0_ROCMFP4_FAST custom GGUF tensor formats + quantization tooling; CPU reference, HIP/ROCm and Vulkan runtime support for those types; ROCmFP4 KV-cache FlashAttention handling; an MTP host-path embedding-fetch cleanup; a Vulkan exact-scale-search pruning optimisation; plus unrelated StepFun Step 3.7 Flash conversion/runtime support carried in the same tree. q38rocm itself (the julianmb repo the Reddit thread links) ships NO llama.cpp source -- it clones charlie12345/ROCmFPX fresh at build time and pins commit e87d53e ("213") in its README/Limitations section; all of the above is charlie12345/ROCmFPX's own work, not julianmb's.
--spec-mtp-strict-qwen MECHANISM (read from ROCmFPX source, not executed): gates on general.architecture == qwen35 or qwen35moe AND draft-mtp enabled; requires -np 1 ("Qwen strict MTP requires a single server slot/sequence"); requires llama_n_rs_seq(ctx_tgt) >= spec-draft-n-max ("bounded recurrent rollback covering the full draft"). When active: caps n_draft_max per step to stay within a 256-token dense-attention KV padding block ("A verification batch that straddles a block makes its earlier rows use a wider ROCm reduction than serial decoding and can change greedy output through rounding" -- server-context.cpp ~2469) and forces a full prompt-cache invalidation on a specific cache-hit path to keep MTP boundary state paired (server-context.cpp ~2929, "prompt cache cold fallback: reason=strict-qwen- exact-hit"). This targets a DIFFERENT failure mode than clm-0055's EOS-cliff (a ROCm floating-point rounding divergence at KV block boundaries vs. this lab's target-distribution collapse to im_end after accumulated session volume) -- it is NOT established, and should not be assumed, that enabling --spec-mtp-strict-qwen would prevent the EOS-cliff signature this lab characterised; that is an open question for anyone actually running their fork, out of scope here (task protocol: audit only, no third-party binaries executed).
q38rocm's own npu_sidecar_drafter.py exposes --strict as an opt-in flag (default OFF, same as run_server.sh's silent omission) with the argparse help text "Enable strict lossless greedy equivalence" -- the marketing language ("mathematically lossless") describing a mode the fork COULD run in but that neither of the repo's two launch paths enables by default.
STOCK-LOAD ATTEMPT (empirical, this job): `~/src/llama.cpp/build/bin/ llama-server -m Qwen3.8-27B-ROCmFP4-FAST.gguf -ngl 999 -fa on -c 4096 --load-mode none` on stock 3653e6d ROCm, rc=1, ~50ms to failure. Verbatim: "gguf_init_from_reader: tensor 'output.weight' has invalid ggml type 101. should be in [0, 43)" then "failed to load model from ...". This is the SAME charlie12345/ROCmFPX fork already carried as a blocker in this lab's records for the Nemotron-3.5-Lightning-30B community ROCmFP4 requant (nemotron35-lightning-30b.yaml note) and muse-glimmer-30b -- third confirmed instance of this exact fork-dependency pattern for a ROCmFP4-labelled community artifact.
RECOMMENDED-CONFIG DRAFT DEPTH, for the public-reply / upstream-filing use case this audit feeds: run_server.sh DRAFT_N default = 6, README's own "Deep Spec" row = n=7. Both exceed this lab's independently-measured n_max=3 hard ceiling for stock draft-mtp on the same architecture family (qwen38-27b). Whether ROCmFPX's fork (a materially different draft-mtp implementation from stock, per the above) reproduces, avoids, or differently exhibits the EOS-cliff at n_max 6-7 is NOT established by this audit and would require actually running their fork, which is out of scope.