Models
Every model runs the same registered targets under the same prompt contract, validator, and pipeline; every run is recorded whether it passes or fails. This page compares models on two separable axes: whether their traces satisfy the elicitation contract they were given, and — as targets resolve against official prints — how accurate their forecasts are. Compliance is not accuracy; read the tables separately.
Contract compliance by lane
Runs passing the sealed trace rubric over runs attempted, from the recorded protocol-lane manifests (strategy lanes in the records until the identifier migration). Fast and Ladder demand the parametric width derivation (“sigma = X”, 1.28·sigma); Ladder v2 is the pre-registered quantile-native contract (rungs plus interpolated 10th/90th percentiles stated literally). Failed runs stay in the record as immutable run manifests.
| Model | Fast | Ladder | Ladder v2 | Median-of-3 |
|---|---|---|---|---|
| gpt-5.5 | 38/39 | 13/13 | 6/6 | 12/13 |
| gpt-5.6 | — | 1/1 | — | — |
| gpt-5.6-luna | 5/18 | 0/6 | 5/6 | 1/6 |
| gpt-5.6-sol | 18/18 | 5/6 | 6/6 | 6/6 |
| gpt-5.6-terra | 18/18 | 0/6 | 6/6 | 6/6 |
Resolved accuracy
Accuracy accrues as registered targets print. The headline tier is witness-verified scores only — runs whose sealed custody root was externally timestamped before the observation; recorded-timestamp (claimed-time) scores sit one tier below, published and flagged, never deleted. Interval coverage here is over claimed-or-better scores; the full scoreboard, paired persistence baselines, and per-agent leaderboards live on Calibration.
| Model | Resolved scores | Witness-verified | Claimed-time or better | 80% coverage |
|---|---|---|---|---|
| claude-fable-5 | 17 | 0 | 17 | 82% |
| Codex recorded agent ensemble | 2 | 0 | 2 | 100% |
| Codex recorded agent run | 16 | 0 | 16 | 75% |
| Codex recorded agent runs | 24 | 0 | 24 | 88% |
| Codex recorded source-context synthesis | 14 | 0 | 14 | 100% |
| gpt-5 | 39 | 0 | 39 | 79% |
| gpt-5-mini | 9 | 0 | 9 | 78% |
| gpt-5.5 | 61 | 0 | 61 | 69% |
| persistence.last_print | 1 | 0 | 1 | 100% |
The gpt-5.6 comparison waves published on 2026-07-10 resolve from mid-July onward; per-model paired CRPS ratios against the persistence baseline appear on Calibration as those targets print.