Golden Benchmark
A fixed set of 385 deterministic, journal-graded tasks every model is scored on once per release, composed with the arena’s Bradley-Terry rating in a frozen cohort. The arena moves with every vote; this half does not. Version gb-1.3.0/env-1.1.0/grader-1.3.0, arena snapshot 2026-09-06.
Standings
gb-1.3.0/env-1.1.0/grader-1.3.0report 2026-09-07arena snapshot 2026-09-06calibration cohort 14W_MAX 0.4
- Tasks
- 385
- Templates
- 77
- Models
- 26
- Episodes
- 10,010
- Templates at ceiling
- 29
- Report generated
- 2026-09-07
- Arena snapshot
- 2026-09-06 (computed 2026-09-06)
- Cohort
- 14 models · roster ∩ coverage = 1.0 ∩ Bradley-Terry ∩ games ≥ 300
- W_MAX
- 0.4
- μB / σB
- 96.857 / 1.891
- μL / σL
- 0.049 / 0.271
- r(Base, arena logit)
- 0.433
- Balance point W
- 0.435
- Arena floor
- 30 games · cohort ≥ 300 games
- Calibration
- scripts/bench/calibration/gb-1.3.0_env-1.1.0_grader-1.3.0.json
| Rank95% band, point position | Model | Composite95% interval | Basetask-set 95% CI | Coretier | Hardtier | Diversetier | Edgetier | Failepisodes < 1.0 | Severeepisodes ≤ 0.25 | Solvedtemplates | Arena±95%, games, status | kshrinkage |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1–4pt 1 | Claude Opus 5 | 99.2 ±0.695% 98.6–99.8 | 99.9task-set 99.6–≤100 | 100.0 | 100.0 | 100.0 | 99.4 | 0.3%Wilson 0.0–1.5% | 0.0% | 76/77 | 1048 ±411,476 games · settled | 0.84 |
| 1–4pt 2 | Claude Fable 5 | 99.1 ±1.095% 98.1–≤100 | 98.9task-set 97.3–≤100 | 100.0 | 97.9 | 100.0 | 96.9 | 2.1%Wilson 1.1–4.0% | 0.0% | 74/77 | 1079 ±351,585 games · settled | 0.87 |
| 2–9pt 3 | Claude Fable 5.1 | 98.4 ±0.795% 97.7–99.1 | 99.5task-set 98.9–≤100 | 100.0 | 99.4 | 98.3 | 100.0 | 0.8%Wilson 0.3–2.3% | 0.0% | 74/77 | 1006 ±49186 games · young | 0.78 |
| 2–9pt 4 | Gemini 3.7 Flash | 97.9 ±1.795% 96.2–99.6 | 97.5task-set 94.2–99.9 | 99.7 | 91.5 | 100.0 | 96.9 | 4.2%Wilson 2.6–6.6% | 1.3% | 73/77 | 1054 ±281,110 games · settled | 0.92 |
| 2–12pt 5 | GPT-5.6 Sol | 97.9 ±1.895% 96.1–99.6 | 97.5task-set 94.2–99.9 | 100.0 | 89.7 | 100.0 | 98.1 | 3.6%Wilson 2.2–6.0% | 2.3% | 73/77 | 1051 ±252,016 games · settled | 0.93 |
| 3–12pt 6 | Claude Sonnet 5 | 97.6 ±1.295% 96.4–98.8 | 97.7task-set 95.6–99.3 | 99.2 | 96.3 | 97.0 | 96.9 | 4.7%Wilson 3.0–7.3% | 0.3% | 69/77 | 1027 ±35958 games · settled | 0.87 |
| 3–12pt 7 | Gemini 3.8 Flash | 97.6 ±1.995% 95.7–99.4 | 97.5task-set 94.2–≤100 | 100.0 | 91.5 | 100.0 | 96.4 | 3.9%Wilson 2.4–6.3% | 1.6% | 74/77 | 1041 ±71205 games · young | 0.63 |
| 4–15pt 8 | GPT-6 Astra | 97.3 ±1.595% 95.8–98.8 | 98.7task-set 96.2–≤100 | 100.0 | 93.6 | 100.0 | 100.0 | 1.6%Wilson 0.7–3.4% | 1.6% | 75/77 | 924 ±9588 games · young | 0.49 |
| 5–14pt 9 | GPT-5.6 Terra | 97.2 ±1.795% 95.5–98.9 | 97.0task-set 93.9–99.3 | 99.4 | 88.6 | 100.0 | 98.1 | 4.2%Wilson 2.6–6.6% | 2.6% | 71/77 | 1025 ±361,107 games · settled | 0.87 |
| 5–15pt 10 | Qwen3.8 Max | 97.0 ±1.295% 95.7–98.2 | 97.0task-set 95.0–98.8 | 98.9 | 92.8 | 97.5 | 97.5 | 4.9%Wilson 3.2–7.6% | 0.5% | 67/77 | 1007 ±34481 games · settled | 0.88 |
| 5–16pt 11 | GPT-5.6 Luna | 96.8 ±1.895% 95.0–98.6 | 96.3task-set 93.1–98.8 | 98.9 | 87.2 | 100.0 | 96.9 | 5.5%Wilson 3.6–8.2% | 2.3% | 68/77 | 1030 ±29898 games · settled | 0.91 |
| 7–16pt 12 | Muse Spark 1.2 | 96.7 ±1.195% 95.6–97.7 | 98.5task-set 96.8–99.7 | 99.8 | 94.2 | 98.9 | 100.0 | 2.6%Wilson 1.4–4.7% | 0.5% | 70/77 | 918 ±44363 games · settled | 0.82 |
| 6–16pt 13 | Kimi K3 | 96.4 ±1.595% 95.0–97.9 | 95.7task-set 93.2–97.8 | 100.0 | 91.4 | 92.0 | 95.8 | 7.0%Wilson 4.9–10.0% | 1.8% | 62/77 | 1030 ±39553 games · settled | 0.85 |
| 9–16pt 14 | Gemini 3.5 Flash | 96.3 ±1.795% 94.6–98.0 | 96.8task-set 93.9–99.1 | 100.0 | 88.6 | 99.8 | 96.4 | 6.0%Wilson 4.0–8.8% | 0.8% | 69/77 | 967 ±35710 games · settled | 0.87 |
| 8–16pt 15 | GLM 5.3 Flash | 96.3 ±1.395% 95.0–97.5 | 97.4task-set 95.4–99.1 | 99.0 | 95.1 | 98.0 | 96.4 | 4.7%Wilson 3.0–7.3% | 0.3% | 67/77 | 931 ±52182 games · young | 0.76 |
| 10–16pt 16 | Grok 4.6 | 96.2 ±1.895% 94.4–98.0 | 96.9task-set 93.6–99.4 | 100.0 | 88.9 | 100.0 | 96.1 | 5.2%Wilson 3.4–7.9% | 1.3% | 71/77 | 961 ±34472 games · settled | 0.88 |
| 16–18pt 17 | Claude Haiku 4.5 | 94.3 ±2.195% 92.2–96.4 | 93.5task-set 89.9–96.6 | 99.4 | 83.6 | 91.0 | 95.3 | 10.6%Wilson 7.9–14.1% | 2.6% | 58/77 | 968 ±42768 games · settled | 0.83 |
| 17–18pt 18 | Inkling | 93.8 ±2.395% 91.5–96.1 | 92.9task-set 88.9–96.3 | 98.3 | 85.3 | 87.2 | 96.2 | 11.2%Wilson 8.4–14.7% | 4.9% | 56/77 | 954 ±50475 games · settled | 0.77 |
| Rank95% band | Model | Composite95% interval | Basetask-set CI | Arena±95% |
|---|---|---|---|---|
| — | Muse Spark 1.3 | 97.6 ±2.595% 95.1–≤100 | 97.6task-set 94.7–99.6 | 932 ±4003 games · arena pending |
| Rank95% band | Model | Composite95% interval | Basetask-set CI | Arena±95% |
|---|---|---|---|---|
| — | Gemini 3.6 Flash | 97.7 ±1.795% 96.0–99.4 | 97.0task-set 93.9–99.4 | 1061 ±30741 games · settled |
| — | GPT-5.5 | 97.6 ±1.695% 96.1–99.2 | 97.4task-set 94.7–99.6 | 1040 ±34442 games · settled |
| — | GPT-5 | 96.4 ±1.195% 95.3–97.5 | 97.9task-set 96.2–99.3 | 930 ±36249 games · young |
| — | Muse Glimmer 30B | 95.1 ±1.995% 93.2–97.0 | 94.4task-set 91.2–97.1 | 974 ±9552 games · young |
| — | Kimi K2.7 Code | 93.7 ±3.495% 90.4–97.1 | 93.7task-set 90.2–96.9 | 1000 ±4001 games · arena pending |
| — | Gemini 3.1 Flash-Lite | 92.5 ±2.495% 90.2–94.9 | 90.6task-set 86.5–94.2 | 966 ±46254 games · young |
| — | Gemini 3.5 Flash-Lite | 92.0 ±2.395% 89.7–94.4 | 89.8task-set 85.7–93.3 | 1024 ±33438 games · settled |
How to read the table
Composite = clamp(Base + W_MAX·σB·clamp(k·z_arena − z_base, ±3), 0, 100); k = σL²/(σL² + SE_L²); w = W_MAX once the arena rating is Bradley-Terry (≥ 30 games), else 0 (listed, unranked)
Base is 100 × the mean over templates of the template mean over its seeds, every template counting once: a fixed set of deterministic, journal-graded tasks with a known correct outcome, run once per model release with the same tasks, seeds and grader version for every model. The task-set CI is a 95% bootstrap over templates: how Base would move if the task set were resampled.
Composite is Base plus W_MAX·σB times the difference between the shrunk arena z-score and the Base z-score, both standardised in the frozen calibration cohort, with that difference clamped to ±3 and the result to [0, 100]. It is a declared 60/40 blend of two constructs (Base and the arena logit correlate at r = 0.433 in the cohort), not an estimate of one quantity. The interval under each Composite is the published 95% interval, 1.96·SE(C) either side of the point and clipped to [0, 100]; the report’s ± is not in the published JSON and is not rebuilt here from the rounded bounds.
k is the empirical-Bayes shrinkage of the arena estimate at its current precision, σL²/(σL² + SE_L²): a young ±95 rating counts for about half of what it says, a settled ±25 for 93%. The blend weight w is W_MAX once the rating is Bradley-Terry (≥ 30 games) and 0 otherwise, so a model below the floor is listed with Composite = Base and not numbered.
Rank is the 95% interval of the model’s rank under a joint bootstrap (templates resampled in step across models, each arena rating jittered by its SE). Where the data cannot separate two models, the band says so instead of printing 0.1-point order.
Top-pair order sensitivity (the W_MAX at which adjacent top pairs would change order; gap in Composite points): Claude Opus 5 vs Claude Fable 5: gap 0.114, flips at W_MAX 0.454 · Claude Fable 5 vs Claude Fable 5.1: gap 0.654, flips at W_MAX 0.192 · Claude Fable 5.1 vs Gemini 3.7 Flash: gap 0.517, flips at W_MAX 0.537 · Gemini 3.7 Flash vs GPT-5.6 Sol: gap 0.046, flip point not reported · GPT-5.6 Sol vs Claude Sonnet 5: gap 0.250, flips at W_MAX 0.174. Balance point W = 0.435; W_MAX kept at 0.4.
Published as src/data/golden/gb-1.3.0_env-1.1.0_grader-1.3.0.json, copied from the bench report and never recomputed in the app. JSON
Stability
Measured over the arena’s real daily snapshots with Base held fixed: mean daily |Δ| in Base points, mean Kendall τ between consecutive daily rankings, and the share of days on which #1 changed — for the arena alone, for the composite at steady state (k = 1 for every agent), and for the composite as published (each day’s k). 14 agents compared over 30 days.
| Series | Mean daily |Δ|Base points | Kendall τconsecutive days | #1 changedshare of days | Days |
|---|---|---|---|---|
| Arena alonein Base points | 0.287 | 0.883 | 7% | 30 |
| Composite, steady statek = 1 for every agent | 0.111 | 0.951 | 27% | 30 |
| Composite, as publishedeach day’s k | 0.082 | 0.958 | 7% | 30 |
Δcomposite = w·σB·k·Δz_arena per agent-day (Base is fixed): the composite rows are W_MAX × the arena's own churn, damped by k.
Reliability
Split-half Spearman ρ of model Base over all ten 2-vs-3 seed splits — mean, minimum, maximum, the Fisher-z 95% interval and the Spearman-Brown full-length estimate. A tier whose cross-model sd is below 1 Base point is marked uninformative. 26 models used.
| Tier | Mean ρ | Min | Max | Fisher-z 95% | Spearman-Brown | Splits | Cross-model sdBase points | Reads |
|---|---|---|---|---|---|---|---|---|
| All templates | 0.853 | 0.774 | 0.924 | 0.696–0.932 | 0.921 | 10 | 2.536 | informative |
| Core tier | 0.482 | 0.179 | 0.657 | 0.117–0.733 | 0.651 | 10 | 0.927 | uninformative |
| Hard tier | 0.848 | 0.806 | 0.931 | 0.685–0.930 | 0.917 | 10 | 5.515 | informative |
| Diverse tier | 0.861 | 0.800 | 0.927 | 0.710–0.936 | 0.925 | 10 | 5.332 | informative |
| Edge tier | 0.573 | 0.428 | 0.720 | 0.239–0.786 | 0.729 | 10 | 1.706 | informative |
Repeatability
A second full pass on the same tasks: Base on run 1 against run 2, the difference, and the number of episodes whose pass/fail flipped. The task-set CI says how Base moves if the tasks are resampled; this says how it moves when the model is re-run. Implied per-run SE of Base (rms Δ/√2): 0.16 points.
| Model | Episodesper pass | Base, run 1 | Base, run 2 | Δ | Pass flips |
|---|---|---|---|---|---|
| Claude Fable 5 | 220 | 98.1 | 98.1 | 0.05 | 4 |
| Claude Fable 5.1 | 220 | 99.2 | 99.5 | 0.32 | 3 |
| Gemini 3.7 Flash | 385 | 97.5 | 97.5 | 0.00 | 0 |
| GPT-5.6 Sol | 385 | 97.5 | 97.5 | 0.07 | 5 |
| Claude Opus 5 | 220 | 99.8 | 99.5 | -0.27 | 1 |
| Qwen3.8 Max | 385 | 97.0 | 97.4 | 0.38 | 16 |
| Claude Sonnet 5 | 385 | 97.7 | 97.6 | -0.11 | 7 |