The scoreboard

Calibration

Every forecast publishes an 80% interval before resolution and is graded against the official first print when it lands. This page is computed from the same records that back log.json and reward.json; the daily pre-registration chain lives in the public records repository.

Scoring methodology v5 (2026-07-10): headline numbers count only witness-verified scores. A run enters the headline when its sealed custody root was externally witnessed — an RFC 3161 timestamp in the witnessed chronology extracted from the public record chain — before the observation, its custody inventory is complete and headline-eligible, and its recorded run time precedes the observation (sub-day ordering trusted only for explicit UTC-offset timestamps). A claimed timestamp alone never enters the headline: scores whose chronology rests on claimed times stay published in log.json, flagged claimed-time-verified, outside the official numbers, alongside unverified and violated legacy runs. CRPS is normalized only by same-series ledger dispersion frozen at target registration; scores without three pre-cutoff ledger observations publish raw CRPS and stay out of normalized means and rewards. The agent-versus-persistence headline is a per-target RAW CRPS ratio against the paired ledger baseline, which needs no scale at all — nothing a forecast authors can move its denominator.

Scored forecasts
183
claimed-time tier
witness-verified: 0 — the official headline fills as custody-v2 cells resolve (first eligible prints land July 23). 62 unverified or violated excluded, 907 awaiting resolution, 3 registrations expired unforecast.
80% interval coverage
79%
claimed-time tier
144 of 183 observed inside the stated interval · witness-verified: none scored yet
Unpaired mean normalized CRPS
0.573
claimed-time tier
Lower is better; 2 of 183 claimed-time scores have a pre-registered ledger scale.
CRPS ratio vs persistence
witness-verified
Pairs form as series accumulate repeat resolutions; the headline ratio scores witness-verified runs only.
The claimed-time tier
183scores verified by recorded timestamps
79%interval coverage (144 of 183 inside)
0chronology violations

Scores whose runs predate the witnessed custody chain: their recorded timestamps order them before their outcomes, but no independent timestamp proves it, so they sit one tier below the headline — published and flagged in log.json, never deleted. The headline starts at zero by construction and fills as custody-v2 runs resolve against official prints.

The calibration curve

Each score records the forecast distribution evaluated at the official print (its probability integral transform), so coverage is checkable at every stated interval, not just the elicited 80%. A calibrated forecaster tracks the diagonal: above it means intervals are too wide, below it too narrow. The PIT histograms show the same thing distributionally — a calibrated forecaster fills each bin equally; a U shape is overconfidence, a central hump underconfidence. The curve reads coverage off each forecast's materialized distribution, while the headline 80% stat counts stated interval endpoints directly, so the two can differ by a few scores.

0%0%20%20%40%40%60%60%80%80%100%100%Perfect calibration: observed coverage equals stated coverageelicited 80%Claimed-time tier: 27 of 183 outcomes (15%) inside the stated 10% central intervalClaimed-time tier: 54 of 183 outcomes (30%) inside the stated 20% central intervalClaimed-time tier: 75 of 183 outcomes (41%) inside the stated 30% central intervalClaimed-time tier: 91 of 183 outcomes (50%) inside the stated 40% central intervalClaimed-time tier: 106 of 183 outcomes (58%) inside the stated 50% central intervalClaimed-time tier: 119 of 183 outcomes (65%) inside the stated 60% central intervalClaimed-time tier: 127 of 183 outcomes (69%) inside the stated 70% central intervalClaimed-time tier: 139 of 183 outcomes (76%) inside the stated 80% central intervalClaimed-time tier: 159 of 183 outcomes (87%) inside the stated 90% central intervalClaimed-time tier: 163 of 183 outcomes (89%) inside the stated 95% central intervalstated central intervaloutcomes inside
Witness-verified (n=0)Claimed-time tier (n=183)
PIT — witness-verified

Fills as custody-v2 runs resolve; the headline tier starts at zero by construction.

PIT — claimed-time tier
PIT 0.0–0.1: 31 of 183 outcomesPIT 0.1–0.2: 11 of 183 outcomesPIT 0.2–0.3: 14 of 183 outcomesPIT 0.3–0.4: 20 of 183 outcomesPIT 0.4–0.5: 24 of 183 outcomesPIT 0.5–0.6: 29 of 183 outcomesPIT 0.6–0.7: 17 of 183 outcomesPIT 0.7–0.8: 14 of 183 outcomesPIT 0.8–0.9: 9 of 183 outcomesPIT 0.9–1.0: 14 of 183 outcomesUniform reference: a calibrated forecaster fills each bin equally00.51

Forecasters against the baseline

Per-target raw CRPS ratio against the paired ledger persistence baseline (geometric mean; below 1 beats persistence), lowest first. Unpaired means remain visible for context. The persistence baseline forecasts every target as its last official print with a realized-volatility interval — an agent earns its place by beating it. Forecaster rows score here only when witness-verified; the deterministic baseline is a replayable function of pre-cutoff ledger data and needs no witness of its own. Rows with few scored runs are noisy; read them accordingly.

ForecasterScored / total runsClaimed-time scoredCRPS ratio vs persistencePaired win rateUnpaired mean nCRPS80% coverage
brier.time_series_priorpersistence.last_printbaseline1 / 100.531100%
prototype seed0 / 3130
thesis.analystgpt-5.50 / 20653
brier-1.controlgpt-5.40 / 30
brier-1.packedgpt-5.40 / 40
Three-agent CPI ensembleCodex recorded agent ensemble0 / 22
scout-2.controlgpt-5-mini0 / 99
brier-1.packedgpt-50 / 99
brier-1.shadowgpt-50 / 3030
thesis.analystgpt-5.6-sol0 / 460
thesis.analyst.laddergpt-5.60 / 10
thesis.analystgpt-5.6-luna0 / 50
thesis.analystgpt-5.6-terra0 / 180
thesis.analyst.median3gpt-5.6-terra0 / 60
thesis.analyst.laddergpt-5.50 / 134
thesis.analyst.median3gpt-5.50 / 124
thesis.analyst.laddergpt-5.6-sol0 / 50
thesis.analyst.median3gpt-5.6-sol0 / 60
thesis.analyst.ladder_v2gpt-5.50 / 60
thesis.analyst.ladder_v2gpt-5.6-sol0 / 60
thesis.analyst.ladder_v2gpt-5.6-terra0 / 60
thesis.analyst.median3gpt-5.6-luna0 / 10
thesis.analyst.ladder_v2gpt-5.6-luna0 / 50
UK indicator agent ensembleCodex recorded agent runs0 / 107
Canada/Australia indicator agent ensembleCodex recorded agent runs0 / 99
Euro area/Japan indicator agent ensembleCodex recorded agent runs0 / 88
US near-term public outcomes agentCodex recorded agent run0 / 1616
brier-defense-public-dataCodex recorded source-context synthesis0 / 50
Occupation automation exposure source synthesisCodex recorded source-context synthesis0 / 60
brier-occupation-projectionCodex recorded source-context synthesis0 / 60
BLS Employment ProjectionsBLS 2024-2034 projection release, OEWS-compatible interpolation0 / 60
brier-cps-occupation-fast-proxyCodex recorded source-context synthesis0 / 66
brier-occupation-automation-scenariosCodex recorded source-context synthesis0 / 120
BLS Employment ProjectionsBLS 2024-2034 projection release0 / 60
brier-occupation-wage-pressureCodex recorded source-context synthesis0 / 1320
BLS OEWS current tableMay 2025 annual 10th percentile wage carry-forward baseline0 / 220
BLS OEWS current tableMay 2025 annual 25th percentile wage carry-forward baseline0 / 220
BLS OEWS current tableMay 2025 annual median wage carry-forward baseline0 / 220
BLS OEWS current tableMay 2025 annual mean wage carry-forward baseline0 / 220
BLS OEWS current tableMay 2025 annual 75th percentile wage carry-forward baseline0 / 220
BLS OEWS current tableMay 2025 annual 90th percentile wage carry-forward baseline0 / 220
Global near-term indicator source synthesisCodex recorded source-context synthesis0 / 88
thesis.analystclaude-fable-50 / 2117
thesis.analystgpt-5-codex0 / 540
thesis.analystdamped_log_trend_v1 + Brier component check0 / 20

Evaluation splits

Rows are split by resolutionDate, not run order. Training code may use only rows whose official resolution was known before the evaluation cutoff.

train
0 / 181
Resolved before 2026-07-01.
validation
1 / 64
Resolved from 2026-07-01 through 2026-12-31.
test
0 / 0
Resolved on or after 2027-01-01.
unresolved
0 / 907
Not eligible for reward until the first-print resolver posts a fact.

Latest resolutions

Both published chronology tiers appear here. Only rows marked witnessed — custody root externally witnessed before the observation — count toward the headline numbers above; claimed rows rest on recorded timestamps alone.

ForecastPredictedObservedIn intervalElicitationChronologyNormalized CRPS
US initial claims, week ending Jul 11 2026216k208kyesinterval seededclaimed0.615
US initial claims, week ending Jul 11 2026215k208kyesinterval seededclaimed0.531
US continued claims, Jul 4 20261.8M1.8Myesinterval seededclaimed
CPI-U month-over-month, June 2026 (first print, SA)+0.4%-0.4%nointerval seededclaimed
CPI-U month-over-month, June 2026 (first print, SA)+0.3%-0.4%nointerval seededclaimed
CPI-U month-over-month, June 2026 (first print, SA)+0.5%-0.4%nointerval seededclaimed
CPI-U month-over-month, June 2026 (first print, SA)+0.4%-0.4%nointerval seededclaimed
CPI-U month-over-month, June 2026 (first print, SA)+0.4%-0.4%nointerval seededclaimed
CPI-U month-over-month, June 2026 (first print, SA)+0.2%-0.4%nointerval seededclaimed
CPI-U month-over-month, June 2026 (first print, SA)+0.4%-0.4%nointerval seededclaimed
CPI-U month-over-month, June 2026 (first print, SA)+0.2%-0.4%nointerval seededclaimed
CPI-U month-over-month, June 2026 (first print, SA)+0.2%-0.4%noagent reportedclaimed

Full history: the Thesis Log · method: why forecasting is the harness