Home › Evidence › Records › clm-0018

clm-0018

communitymed ●●○
citable URL: https://halobench.com/records/clm-0018/ — this address never moves; the anchor /records/#clm-0018 keeps resolving

DeepGrove Maple-Preview (20.2B-A1.49B ternary, 5.31 GB, MIT) is a credible second-lane candidate for the Mac mini, but not a Warden candidate — its own model card concedes underperformance on agentic benchmarks. The "on-device weight adaptation (dreaming)" feature is undocumented at source, and even if it works it is architecturally opposed to how Warden's state is built.

verified 2026-08-06 · volatility high

Note — the record's own working

PROVENANCE MATTERS HERE MORE THAN USUAL — the claims arrive at three different levels of evidence and the most interesting one is the weakest: · **Model card (primary).** 20B-A1B, 24 layers, 256 experts / 8 active, ternary weights 2-bit packed as {-a, 0, +a} with one a per row, 5.31 GB checkpoint, 131,072 context, 3:1 SWA-512:global attention, MIT licence. Quantisations exist for llama.cpp, Ollama, LM Studio, Jan. Apple Silicon runs via the authors' own MLX fork (deepgrove-ai/mlx-lm-deepgrove). · **Model card, self-reported, NO NUMBERS.** Evaluated on LCBv6, AIME 2026, HMMT 2026, GPQA-D — but published as a *visual comparison chart* with no scores. Claims a new point on the Pareto frontier for memory-to-performance. · **Launch demo on X + aggregator blog only.** The "dreaming" weight adaptation, the ~5.9 GB adaptation peak, and the 200+ tok/s Mac mini figure. The model card says **nothing** about weight adaptation. The company describes adaptive behaviour partly in FUTURE TENSE ("plans to enhance"), so this is a demo and a roadmap, not a released, reproducible capability. 419 downloads in the last month — very low adoption. For scale, the NPU build in clm-0013 had ~2,000. Nothing here has been independently reproduced. ⛔ DISQUALIFYING FOR WARDEN'S PRIMARY ROLE, on the authors' own evidence. The model card acknowledges **underperformance on agentic benchmarks** and notes minimal post-training for agentic tasks. Warden is an agentic system whose stated failure mode is tool selection. A model that is weak at exactly that is not a candidate for the main lane, however good its reasoning scores turn out to be. This is the rare case where the vendor tells you the disqualifying fact themselves. ✅ WHERE IT IS GENUINELY INTERESTING — THE MAC MINI. wardenmac is an M4 Mac mini with 16 GB currently carrying only vision (MLX) and STT; no LLM lane. A 5.31 GB checkpoint at a claimed 200+ tok/s would fit with enormous headroom, and MLX is a path already proven on that box. That makes it a candidate SECOND LANE that does not depend on aibeast at all — which is the point, because the single-slot bottleneck is an aibeast problem. It is a more plausible route than the XDNA2 NPU option (clm-0013), where measured data showed the iGPU winning both phases by ~4x. ON "DREAMING" — THE NAME COLLIDES WITH OURS AND THE MECHANISM IS OPPOSITE. Warden's dreaming consolidates recalls into MEMORY.md: text, dated, diffable, revertible. Maple's dreaming runs a local fine-tune that embeds a preference **into the weights**. The architectural appeal is real — persistent preference that costs no context is a direct attack on our dominant cost, since the entire warm-lane programme exists because prefill of a large persistent prompt is expensive. But it is opposed to the principle the whole system is built on: **legible, auditable, revertible state.** `anticipatory-flags.md` declares the operator as its owner. Every recovery this month depended on being able to read state and diff it — the 2026-08-06 audit found a SESSION-STATE that had asserted a wrong host address for six weeks, and it was fixable precisely because it was text. A preference baked into weights cannot be inspected, diffed, dated, or switched off with a flag; you would discover a wrong one only by noticing bad behaviour, and then have no way to attribute it. You cannot `git diff` a weight delta into an explanation. If we ever pursue this, the defensible split is: weights may hold **stable, low-stakes, stylistic** adaptation (tone, formatting); anything with consequences stays in text. And even then the audit story is poor enough that it should be a deliberate experiment on a lab box, never on the production agent. TERNARY CAVEAT: replacing multiplies with additions is real, but the speedup needs kernels that exploit it. Community llama.cpp quants exist (e.g. stamsam/maple-preview-gguf), but ternary support in llama.cpp is narrower than standard K-quants and the Apple Silicon numbers come from the authors' own MLX fork rather than a neutral runtime. Treat the 200+ tok/s figure as vendor-path until measured on a stock runtime. IF WE TEST IT, the honest first question is not speed but capability: does a 1.49B-active model actually hold up? That is precisely what the capability tier exists for — and per protocol §1a it must be measured, not inherited, since this differs from everything we run by architecture, quant class and runtime simultaneously.

Cited by — computed at build time, never stored

candidate gate history maple-preview