The scoreboard
Calibration
Every forecast publishes an 80% interval before resolution and is graded against the official first print when it lands. This page is computed from the same records that back log.json and reward.json; the daily pre-registration chain lives in the public records repository.
Scoring methodology v5 (2026-07-10): headline numbers count only witness-verified scores. A run enters the headline when its sealed custody root was externally witnessed — an RFC 3161 timestamp in the witnessed chronology extracted from the public record chain — before the observation, its custody inventory is complete and headline-eligible, and its recorded run time precedes the observation (sub-day ordering trusted only for explicit UTC-offset timestamps). A claimed timestamp alone never enters the headline: scores whose chronology rests on claimed times stay published in log.json, flagged claimed-time-verified, outside the official numbers, alongside unverified and violated legacy runs. CRPS is normalized only by same-series ledger dispersion frozen at target registration; scores without three pre-cutoff ledger observations publish raw CRPS and stay out of normalized means and rewards. The agent-versus-persistence headline is a per-target RAW CRPS ratio against the paired ledger baseline, which needs no scale at all — nothing a forecast authors can move its denominator.
Scores whose runs predate the witnessed custody chain: their recorded timestamps order them before their outcomes, but no independent timestamp proves it, so they sit one tier below the headline — published and flagged in log.json, never deleted. The headline starts at zero by construction and fills as custody-v2 runs resolve against official prints.
The calibration curve
Each score records the forecast distribution evaluated at the official print (its probability integral transform), so coverage is checkable at every stated interval, not just the elicited 80%. A calibrated forecaster tracks the diagonal: above it means intervals are too wide, below it too narrow. The PIT histograms show the same thing distributionally — a calibrated forecaster fills each bin equally; a U shape is overconfidence, a central hump underconfidence. The curve reads coverage off each forecast's materialized distribution, while the headline 80% stat counts stated interval endpoints directly, so the two can differ by a few scores.
Fills as custody-v2 runs resolve; the headline tier starts at zero by construction.
Forecasters against the baseline
Per-target raw CRPS ratio against the paired ledger persistence baseline (geometric mean; below 1 beats persistence), lowest first. Unpaired means remain visible for context. The persistence baseline forecasts every target as its last official print with a realized-volatility interval — an agent earns its place by beating it. Forecaster rows score here only when witness-verified; the deterministic baseline is a replayable function of pre-cutoff ledger data and needs no witness of its own. Rows with few scored runs are noisy; read them accordingly.
| Forecaster | Scored / total runs | Claimed-time scored | CRPS ratio vs persistence | Paired win rate | Unpaired mean nCRPS | 80% coverage |
|---|---|---|---|---|---|---|
| brier.time_series_priorpersistence.last_printbaseline | 1 / 1 | 0 | — | — | 0.531 | 100% |
| prototype seed | 0 / 313 | 0 | — | — | — | — |
| thesis.analystgpt-5.5 | 0 / 206 | 53 | — | — | — | — |
| brier-1.controlgpt-5.4 | 0 / 3 | 0 | — | — | — | — |
| brier-1.packedgpt-5.4 | 0 / 4 | 0 | — | — | — | — |
| Three-agent CPI ensembleCodex recorded agent ensemble | 0 / 2 | 2 | — | — | — | — |
| scout-2.controlgpt-5-mini | 0 / 9 | 9 | — | — | — | — |
| brier-1.packedgpt-5 | 0 / 9 | 9 | — | — | — | — |
| brier-1.shadowgpt-5 | 0 / 30 | 30 | — | — | — | — |
| thesis.analystgpt-5.6-sol | 0 / 46 | 0 | — | — | — | — |
| thesis.analyst.laddergpt-5.6 | 0 / 1 | 0 | — | — | — | — |
| thesis.analystgpt-5.6-luna | 0 / 5 | 0 | — | — | — | — |
| thesis.analystgpt-5.6-terra | 0 / 18 | 0 | — | — | — | — |
| thesis.analyst.median3gpt-5.6-terra | 0 / 6 | 0 | — | — | — | — |
| thesis.analyst.laddergpt-5.5 | 0 / 13 | 4 | — | — | — | — |
| thesis.analyst.median3gpt-5.5 | 0 / 12 | 4 | — | — | — | — |
| thesis.analyst.laddergpt-5.6-sol | 0 / 5 | 0 | — | — | — | — |
| thesis.analyst.median3gpt-5.6-sol | 0 / 6 | 0 | — | — | — | — |
| thesis.analyst.ladder_v2gpt-5.5 | 0 / 6 | 0 | — | — | — | — |
| thesis.analyst.ladder_v2gpt-5.6-sol | 0 / 6 | 0 | — | — | — | — |
| thesis.analyst.ladder_v2gpt-5.6-terra | 0 / 6 | 0 | — | — | — | — |
| thesis.analyst.median3gpt-5.6-luna | 0 / 1 | 0 | — | — | — | — |
| thesis.analyst.ladder_v2gpt-5.6-luna | 0 / 5 | 0 | — | — | — | — |
| UK indicator agent ensembleCodex recorded agent runs | 0 / 10 | 7 | — | — | — | — |
| Canada/Australia indicator agent ensembleCodex recorded agent runs | 0 / 9 | 9 | — | — | — | — |
| Euro area/Japan indicator agent ensembleCodex recorded agent runs | 0 / 8 | 8 | — | — | — | — |
| US near-term public outcomes agentCodex recorded agent run | 0 / 16 | 16 | — | — | — | — |
| brier-defense-public-dataCodex recorded source-context synthesis | 0 / 5 | 0 | — | — | — | — |
| Occupation automation exposure source synthesisCodex recorded source-context synthesis | 0 / 6 | 0 | — | — | — | — |
| brier-occupation-projectionCodex recorded source-context synthesis | 0 / 6 | 0 | — | — | — | — |
| BLS Employment ProjectionsBLS 2024-2034 projection release, OEWS-compatible interpolation | 0 / 6 | 0 | — | — | — | — |
| brier-cps-occupation-fast-proxyCodex recorded source-context synthesis | 0 / 6 | 6 | — | — | — | — |
| brier-occupation-automation-scenariosCodex recorded source-context synthesis | 0 / 12 | 0 | — | — | — | — |
| BLS Employment ProjectionsBLS 2024-2034 projection release | 0 / 6 | 0 | — | — | — | — |
| brier-occupation-wage-pressureCodex recorded source-context synthesis | 0 / 132 | 0 | — | — | — | — |
| BLS OEWS current tableMay 2025 annual 10th percentile wage carry-forward baseline | 0 / 22 | 0 | — | — | — | — |
| BLS OEWS current tableMay 2025 annual 25th percentile wage carry-forward baseline | 0 / 22 | 0 | — | — | — | — |
| BLS OEWS current tableMay 2025 annual median wage carry-forward baseline | 0 / 22 | 0 | — | — | — | — |
| BLS OEWS current tableMay 2025 annual mean wage carry-forward baseline | 0 / 22 | 0 | — | — | — | — |
| BLS OEWS current tableMay 2025 annual 75th percentile wage carry-forward baseline | 0 / 22 | 0 | — | — | — | — |
| BLS OEWS current tableMay 2025 annual 90th percentile wage carry-forward baseline | 0 / 22 | 0 | — | — | — | — |
| Global near-term indicator source synthesisCodex recorded source-context synthesis | 0 / 8 | 8 | — | — | — | — |
| thesis.analystclaude-fable-5 | 0 / 21 | 17 | — | — | — | — |
| thesis.analystgpt-5-codex | 0 / 54 | 0 | — | — | — | — |
| thesis.analystdamped_log_trend_v1 + Brier component check | 0 / 2 | 0 | — | — | — | — |
Evaluation splits
Rows are split by resolutionDate, not run order. Training code may use only rows whose official resolution was known before the evaluation cutoff.
Latest resolutions
Both published chronology tiers appear here. Only rows marked witnessed — custody root externally witnessed before the observation — count toward the headline numbers above; claimed rows rest on recorded timestamps alone.
Full history: the Thesis Log · method: why forecasting is the harness