The Brier skill score (BSS) converts a Brier score into a statement about added value: did these forecasts beat a naive alternative that required no skill?
The formula
BSS = 1 − (Brier score of the forecasts ÷ Brier score of the reference)
- If the forecasts and the reference score the same, BSS = 0: no skill shown.
- If the forecasts score better (lower), BSS is positive. A BSS of +0.10 means the forecasts cut the reference's squared error by 10%.
- If the forecasts score worse, BSS is negative. There is no floor: a confident forecaster who is often wrong can go well below zero.
- A perfect set of forecasts would score +1.
Choosing the reference
The reference is what makes the number honest. The most common choice is the climatological or base-rate forecast — always predicting the long-run frequency of the event. If Bitcoin has closed higher over seven days 56% of the time, the reference says "56% up" every week. Beating that is harder than beating a coin flip, and it is the right bar: a forecaster who cannot beat the historical frequency is adding noise, not information.
Sample size matters
A skill score from a handful of forecasts can swing from strongly positive to negative with one or two outcomes. That is why a skill claim needs a sample size stated beside it, and why a result on a short sample is described as provisional. The same score on 200 resolved calls means far more than on 15.
In MoonWire analysis
Our daily prognosis reports skill against the base rate with the number of graded calls behind it — for example "a provisional +0.0426 on 18 graded calls" — and states when the sample is still below the threshold at which we would treat the figure as established. When the skill score is negative, the article says so. Calibration is reported separately, because a forecaster can have positive skill and still be over- or under-confident.