Models › gpt-oss-120b

gpt-oss-120bbenchedguard failed

117B / ~5B activeMoE · harmony chat formatquant held: UD-Q4_K_XLfirst measured 2026-06-27latest run 2026-08-08

Verdict

Excluded. The guard failed: the model returns stale answers from previous requests — the first request on a fresh server answers correctly and every subsequent request returns the first one's answer, reproduced 6/6 at --parallel 1 with prompt caching ruled out clm-0025. A model that occasionally serves a previous answer is far worse than a slow one, so it is not a candidate for any role until the staleness is understood. Its throughput rows stay recorded but not guard-cleared: the fastest decode at empty context in the field, 55.45 tok/s, and those numbers say nothing about whether it generates the right thing clm-0025. Flash attention is mandatory — both fa-off arms failed outright — and q8_0 KV costs 56% of decode at 131k on the stock build, more than anything else measured clm-0025. On the blind eight-task conceptual screen it placed last at 3/8; that result is one date, one quant, one build, N=1 per cell, and is not a general capability verdict clm-0004.

verdict written 2026-08-13 · every number above stands next to the claim chip that carries it

Best configuration

modelgpt-oss-120b-UD-Q4_K_XL-00001-of-00002.gguf · UD-Q4_K_XL
engineggml-org/llama.cpp 3653e6d · rocm · host aihydra (igpu)
flags-ngl 999 -fa 1 -b 2048 -ub 512 -ctk f16 -ctv f16 -t 16
templatenot recorded at test time
treeupstream — stock

backfilled aged evidence — reconstructed from the archive · config record cfg-0016

decode @ 0
55.45 t/s
run-0045 · CV 0.1%
decode @ 32k
41.70 t/s
run-0049 · CV 0%
prefill @ 0
450.83 t/s
run-0044 · CV 0.3%
prefill @ 32k
308.55 t/s
run-0048 · CV 0.7%
Wh per correct answer
unmeasured
open question ↓

every cell generated from the record at build time · throughput cells from cfg-0016 (same build 3653e6d, fa on, f16 KV) · provenance: measured-here throughout

③aDecode against context depth

f16 KV · fa onq8_0 KV · stock
0102030405060032k65k131k204.8kdecode t/scontext depth (tokens)55.43 t/s @ depth 0 · run-0051 · CV 0.1% · N=355.45 t/s @ depth 0 · run-0045 · CV 0.1% · N=352.79 t/s @ depth 4k · run-0047 · CV 0.1% · N=341.67 t/s @ depth 32k · run-0053 · CV 0.1% · N=341.70 t/s @ depth 32k · run-0049 · CV 0% · N=324.53 t/s @ depth 131k · run-0055 · CV 0.1% · N=354.93 t/s @ depth 0 · run-0057 · CV 0.1% · N=334.08 t/s @ depth 32k · run-0059 · CV 0.1% · N=310.70 t/s @ depth 131k · run-0061 · CV 0% · N=3f16 KV · fa on · 24.53q8_0 KV · stock · 10.70
3 reps per cell · max CV 0.1% · build 3653e6d · no flash-attention-off series exists — both fa-off arms failed to run at all; throughput here is recorded, not guard-cleared · records: run-0051 run-0045 run-0047 run-0053 run-0049 run-0055 run-0057 run-0059 run-0061

Other configurations tested — each as a delta against best

variantΔ decodeΔ turns (paired tasks)Δ energynoterecords
q8_0 KV · stock build-56% @131kThe steepest quantised-KV penalty in the field — which also makes this model the largest potential beneficiary of the KV-dequant patch, if it is ever cleared to matter. clm-0025 clm-0022
flash attention off hazardfailed outright — no series existsBoth flash-attention-off arms (f16 and q8_0 KV) failed to produce a run — recorded as failures against their fingerprints, not retry candidates. clm-0025

deltas computed at build time from the named runs, paired-task traces and metered windows — a hazard row shows words where a delta would mislead

Open questions

Provenance

bench host aihydra · rocm · ggml-org/llama.cpp 3653e6d
discipline 3 reps per throughput cell · scatter published per cell (max CV 0.8%) · guard chain on every performance series
window 2026-06-27 → 2026-08-08