Strategy Lab · baseline discipline
Forecast strategies have to beat simple replayable benchmarks
This page compares deterministic baselines with the actual Brier agent run on resolved target panels. The first slice is SNAP FY2025 state payment error rates: every baseline uses only FY2023 and FY2024 values, while the agent row is the immutable published forecast. Everything here is a retrospective reconstruction — strategy rows are replayed over outcomes that already resolved, carry no chronology verification, and never enter headline calibration, rewards, or leaderboards.
LLM Judge Layer
Judge records score forecast process quality from the public trace: base rates, source grounding, resolution clarity, uncertainty, mechanisms, counterarguments, and coherence. They are auxiliary diagnostics; reward still comes only from resolved CRPS.
| band | judged | scored | judge | nCRPS | 80% cover |
|---|---|---|---|---|---|
weak 0.0-2.0 | 0 | 0 | pending | pending | pending |
developing 2.0-3.0 | 422 | 30 | 2.729 | pending | 87% |
strong 3.0-4.0 | 765 | 170 | 3.515 | 1.747 | 75% |
| disagreement | target | run | judge | nCRPS | 80% cover |
|---|---|---|---|---|---|
high judge bad score | initial-claims-week-2026-07-18 | Ledger persistence baseline | 3.160 | 2.996 | outside |
high judge bad score | initial-claims-week-2026-07-18 | Headline | 3.780 | 2.845 | outside |
Agent Adjustments From Time-Series Prior
Each eligible agent run now gets a last-print persistence prior. The adjustment is the primary Brier point estimate minus that prior, scaled where useful by the prior's 80% interval.
| target | latest | prior | agent | adjustment | interval share |
|---|---|---|---|---|---|
| us.dol.initial_claims.sa.week_2026-07-11 US initial claims, week ending Jul 11 2026 | 215k 2026-07-04 | 215k | 216k | +1k | +0.06x |
| us.dol.initial_claims.sa.week_2026-07-18 US initial claims, week ending July 18 | 215k 2026-07-04 | 215k | 216k | +1k | +0.06x |
SNAP FY2025 state payment error panel
Resolved state and territory payment error rate forecasts scored against simple baselines and the recorded Brier agent run.
Strategy Scoreboard
lower nCRPS and MAE are better| strategy | mode | rows | nCRPS | MAE | vs persistence | bias | 80% cover |
|---|---|---|---|---|---|---|---|
Last-print persistence Hold the most recent official value flat and score it as a replayable benchmark. | historical replay deterministic baseline interval seeded | 53 | pending | 1.38pp | 0.00pp | +0.01pp | 87% |
Panel shrinkage trend Blend 25% of the entity's last annual change with 75% of the panel median change. | historical replay deterministic baseline interval seeded | 53 | pending | 1.79pp | +0.41pp | +0.66pp | 75% |
Brier primary agent The public Thesis forecast run actually recorded before resolution. | forward only agent forward interval seeded | 53 | pending | 1.81pp | +0.43pp | +0.72pp | 72% |
Largest Agent nCRPS Misses Against Persistence
top 0| target | actual | persistence | agent | nCRPS delta | MAE delta |
|---|