Home › Evidence › Records › clm-0004

clm-0004

measured-heremed ●●○
citable URL: https://halobench.com/records/clm-0004/ — this address never moves; the anchor /records/#clm-0004 keeps resolving

gpt-oss-120b placed last on our blind conceptual set — 3/8, mean 1.38 — against Qwen3.6-35B-A3B's 7/8, mean 2.12.

verified 2026-06-28 · volatility high
evidence run-0003

Note — the record's own working

FINGERPRINT, without which this is a slight rather than a finding. Config cfg-0003: Q4_K_M, ROCm, -c 131072, q8_0 KV, reasoning OFF, `-ngl 999 --no-mmap --parallel 1 -fa 1 --jinja`, on a ~82 GiB budget. Tasks C1-C8, one response per model per task. Judging: a separate, fresh Claude Opus sub-agent per response, given only the task prompt, the rubric and one anonymised answer — no model name, no comparison set, no project context. Every rationale preserved and re-checkable. SCOPE: eight tasks, N=1 per cell, one quant, one date, one llama.cpp build, and a rubric aimed at proactive/generative assistant work. It is not a general capability verdict. The same model was the FASTEST reactive performer in the field (~5 s). Volatility high: this was true of that build on that date and has not been re-tested.

Cited by — computed at build time, never stored