← MoonWire

Crypto glossary

What Is Forecast Calibration?

Whether stated probabilities match real frequencies: of all the times a forecaster said 70%, did the event happen about 70% of the time? It measures honesty of confidence, not accuracy.

A forecaster is well calibrated when the probabilities they give line up with how often things actually happen. Take every forecast issued at 70%: if about 70% of those events occurred, the 70% forecasts were calibrated. If 90% occurred, the forecaster was under-confident; if 50% occurred, over-confident.

Calibration is not accuracy

A weather service that says "30% chance of rain" every day in a city where it rains on 30% of days is perfectly calibrated — and useless, because it never tells you which days. Calibration asks whether stated confidence can be taken at face value. Whether the forecasts separate events that happen from events that do not is a different property, often called resolution or discrimination. Good forecasting needs both.

How it is measured

  1. Group forecasts into probability bins, for example 50–60%, 60–70%, 70–80%.
  2. For each bin, compare the average stated probability with the observed frequency of the outcome.
  3. Plot the pairs as a reliability diagram: a perfectly calibrated forecaster lies on the diagonal.

A single summary is the expected calibration error: the average gap between stated and observed frequency across bins, weighted by how many forecasts fall in each. With small samples, each bin holds only a few forecasts, so calibration estimates are noisy and should be read with their counts.

Why it matters for published forecasts

A probability is only useful to a reader if it means what it says. If "65% likely" from a source historically comes true 50% of the time, every reader who took the number literally was misled. Publishing calibration alongside the Brier score and skill score lets readers discount a forecaster's stated confidence by the right amount.

In MoonWire analysis

Our daily prognosis tracks how its stated probabilities compare with outcomes over the resolved record, and treats calibration findings on short samples as provisional rather than established.

Further reading

Where this appears in MoonWire analysis (3)

Join MoonWire Early Access →

Real-time signal intel — AI-read crypto news, importance-scored and de-noised.

Glossary

Full explanation →