Research Note
Probability Calibration Practice
A calibrated forecaster's events assigned 70 percent should occur about 70 percent of the time across a sufficiently comparable set. One outcome cannot establish calibrat
In this article
Probability Calibration Practice
A calibrated forecaster's events assigned 70 percent should occur about 70 percent of the time across a sufficiently comparable set. One outcome cannot establish calibration.
Useful practice starts with clearly defined, time-bounded, resolvable questions. Record one probability before the outcome, preserve the definition, and score comparable forecasts later.
The Brier score for a binary forecast is the squared difference between the forecast probability and the observed result, averaged across forecasts. Lower is better in the common binary formulation. A score combines calibration and resolution, so it should not be interpreted alone.
Probability language should not create false precision. Use evidence, base rates, assumptions, and ranges where appropriate. Update a forecast when evidence changes, but preserve the earlier forecast and reason for the update.
Sources
Follow the evidence.
- WSOP report on the 2010 heads-up championshipwsop.com
- Outcome-bias replication and extensionspmc.ncbi.nlm.nih.gov
- Brier probability-forecast verification paperjournals.ametsoc.org
- WSOP official media guidewsop.com
- Thinking in Bets publisher recordpenguinrandomhouse.com
- WSOP Tournament of Champions historywsop.com
- Prospective hindsight studyonlinelibrary.wiley.com
- Outcome bias in decision evaluationpubmed.ncbi.nlm.nih.gov
- NFL account of the Malcolm Butler interceptionnfl.com
- Spotify episodeopen.spotify.com