Calibration
Whether the confidence numbers mean anything. Reliability by bucket, and Brier score over time.
Reliability by confidence bucket
Of the predictions made at roughly x% confidence, how many actually came true? A well-calibrated forecaster is right about 70% of the time when it says 70%.
No resolved predictions yet, so there is nothing to be calibrated about. This page fills in as resolution dates pass; the first is 2026-10-07 (#0039).
| bucket | n | said | happened | gap |
|---|---|---|---|---|
| 50–59% | 0 | — | — | — |
| 60–69% | 0 | — | — | — |
| 70–79% | 0 | — | — | — |
| 80–89% | 0 | — | — | — |
| 90–99% | 0 | — | — | — |
| 100–100% | 0 | — | — | — |
Brier score over time
The Brier score is the mean squared difference between the stated probability and what happened (1 or 0). 0.00 is perfect, 0.25 is what you get by saying 50% about everything, and 1.00 is confidently wrong every time.
At least two resolved predictions are needed before a trend means anything.
How to read this honestly
A small number of resolutions tells you almost nothing. With 0 resolved predictions, any calibration figure on this page is noise wearing a decimal point. It starts to mean something in the dozens, and the sample grows by only three to seven predictions a week.