Brier Lab

Strategy Lab · baseline discipline

Forecast strategies have to beat simple replayable benchmarks

This page compares deterministic baselines with the actual Brier agent run on resolved target panels. The first slice is SNAP FY2025 state payment error rates: every baseline uses only FY2023 and FY2024 values, while the agent row is the immutable published forecast. Everything here is a retrospective reconstruction — strategy rows are replayed over outcomes that already resolved, carry no chronology verification, and never enter headline calibration, rewards, or leaderboards.

strategies
3
targets
53
resolved
53
score rows
159

LLM Judge Layer

Judge records score forecast process quality from the public trace: base rates, source grounding, resolution clarity, uncertainty, mechanisms, counterarguments, and coherence. They are auxiliary diagnostics; reward still comes only from resolved CRPS.

reward-ineligible process eval
trace judges
1,187
scored judges
4
judge ↔ -nCRPS
0.010
post-resolution reviews
200
bandjudgedscoredjudgenCRPS80% cover
weak
0.0-2.0
00pendingpendingpending
developing
2.0-3.0
422302.729pending87%
strong
3.0-4.0
7651703.5151.74775%
disagreementtargetrunjudgenCRPS80% cover
high judge bad score
initial-claims-week-2026-07-18Ledger persistence baseline3.1602.996outside
high judge bad score
initial-claims-week-2026-07-18Headline3.7802.845outside

Agent Adjustments From Time-Series Prior

Each eligible agent run now gets a last-print persistence prior. The adjustment is the primary Brier point estimate minus that prior, scaled where useful by the prior's 80% interval.

2 prior comparisons
direction
2 up / 0 down
flat
0
median |adj| / interval
0.06x
mean |adj| / interval
0.06x
targetlatestprioragentadjustmentinterval share
us.dol.initial_claims.sa.week_2026-07-11
US initial claims, week ending Jul 11 2026
215k
2026-07-04
215k216k+1k+0.06x
us.dol.initial_claims.sa.week_2026-07-18
US initial claims, week ending July 18
215k
2026-07-04
215k216k+1k+0.06x

SNAP FY2025 state payment error panel

Resolved state and territory payment error rate forecasts scored against simple baselines and the recorded Brier agent run.

53 resolved targets
FY2023 median
10.29pp
FY2024 median
9.52pp
median change
-0.11pp
p10 change
-7.98pp
p90 change
+1.97pp

Strategy Scoreboard

lower nCRPS and MAE are better
strategymoderowsnCRPSMAEvs persistencebias80% cover
Last-print persistence
Hold the most recent official value flat and score it as a replayable benchmark.
historical replay
deterministic baseline
interval seeded
53pending1.38pp0.00pp+0.01pp87%
Panel shrinkage trend
Blend 25% of the entity's last annual change with 75% of the panel median change.
historical replay
deterministic baseline
interval seeded
53pending1.79pp+0.41pp+0.66pp75%
Brier primary agent
The public Thesis forecast run actually recorded before resolution.
forward only
agent forward
interval seeded
53pending1.81pp+0.43pp+0.72pp72%
persistence
1.38pp
mean absolute error
Brier primary
1.81pp
MAE delta +0.43pp

Largest Agent nCRPS Misses Against Persistence

top 0
targetactualpersistenceagentnCRPS deltaMAE delta