Models

Every model runs the same registered targets under the same prompt contract, validator, and pipeline; every run is recorded whether it passes or fails. This page compares models on two separable axes: whether their traces satisfy the elicitation contract they were given, and — as targets resolve against official prints — how accurate their forecasts are. Compliance is not accuracy; read the tables separately.

Contract compliance by lane

Runs passing the sealed trace rubric over runs attempted, from the recorded protocol-lane manifests (strategy lanes in the records until the identifier migration). Fast and Ladder demand the parametric width derivation (“sigma = X”, 1.28·sigma); Ladder v2 is the pre-registered quantile-native contract (rungs plus interpolated 10th/90th percentiles stated literally). Failed runs stay in the record as immutable run manifests.

ModelFastLadderLadder v2Median-of-3
gpt-5.538/3913/136/612/13
gpt-5.61/1
gpt-5.6-luna5/180/65/61/6
gpt-5.6-sol18/185/66/66/6
gpt-5.6-terra18/180/66/66/6

Resolved accuracy

Accuracy accrues as registered targets print. The headline tier is witness-verified scores only — runs whose sealed custody root was externally timestamped before the observation; recorded-timestamp (claimed-time) scores sit one tier below, published and flagged, never deleted. Interval coverage here is over claimed-or-better scores; the full scoreboard, paired persistence baselines, and per-agent leaderboards live on Calibration.

ModelResolved scoresWitness-verifiedClaimed-time or better80% coverage
claude-fable-51701782%
Codex recorded agent ensemble202100%
Codex recorded agent run1601675%
Codex recorded agent runs2402488%
Codex recorded source-context synthesis14014100%
gpt-53903979%
gpt-5-mini90978%
gpt-5.56106169%
persistence.last_print101100%

The gpt-5.6 comparison waves published on 2026-07-10 resolve from mid-July onward; per-model paired CRPS ratios against the persistence baseline appear on Calibration as those targets print.