Search runs in your browser, across every published page on this site.

Home/Guides/Forecasting

Why Brier score beats accuracy for judging a forecaster

Accuracy throws away the confidence and rewards predicting the base rate. A proper score does neither. The arithmetic, the decomposition, and how to read a forecasting scoreboard.

Accuracy asks whether the forecast landed on the right side of 50%, which discards the only interesting thing a probability carries. Brier score is the mean squared difference between the probability and the outcome, so it rewards being confident when you should be and punishes it when you should not. It is also proper: your expected score is best when you report exactly what you believe, which no accuracy measure can claim.

Accuracy fails on two forecasters who are equally right

One forecaster says 99% on a hundred questions and every one resolves yes. Another says 51% on the same hundred, and every one resolves yes. Both score 100% accuracy. Their Brier scores are 0.0001 and 0.2401, and the difference is the whole point: the first forecaster told you something and the second told you almost nothing.

The failure runs the other way too. On a question that resolves yes 5% of the time, always answering no scores 95% accuracy while carrying zero information. Any measure that a base-rate parrot can top is not measuring forecasting.

What proper means, and why it is the whole argument

A scoring rule is proper when your expected score is optimised by reporting your true belief. Brier is proper. Log loss is proper. Accuracy is not: if you believe an event is 55% likely, accuracy pays you exactly the same for saying 55% as for saying 99%, so it actively invites overstatement. That is a design flaw, not a nuance, and it is the reason every serious forecasting programme scores probabilities rather than calls.

Brier is the mean of (p minus outcome) squared, where the outcome is 1 or 0. A permanent 50% forecaster scores 0.25, which is the number to anchor on. Log loss is the other standard choice and it differs in one important way: it goes to infinity on a confident miss, so a single 99% forecast on something that does not happen wrecks the average. Which behaviour you want depends on whether a catastrophic overconfidence should be survivable in the record. Brier says yes, log loss says no.

The decomposition tells you which part is broken

Brier score splits into three terms, following Murphy: reliability, resolution, and uncertainty. Reliability is calibration, whether things you call 70% happen about 70% of the time, and it is the fixable part. Resolution is how far your forecasts move away from the base rate while staying calibrated, and it is the hard part, the part that is actually skill. Uncertainty is a property of the questions rather than of you, which is why raw Brier scores across different question sets are not comparable.

A forecaster with poor calibration and good resolution is worth listening to after a correction. A forecaster with perfect calibration and no resolution is a base-rate table with a byline.

Reading a scoreboard without being sold one

Four things make a published score meaningful, and their absence is usually deliberate. A sample size, because a Brier score over twelve questions is a mood. A predeclared reference, since the only interpretable comparison is against a benchmark chosen before the results, which is what a Brier skill score is for. Event deduplication, because forty forecasts about one election are close to one forecast. And a frozen issue time, since a probability that can be edited after the fact is not a forecast.

Our position: a scoreboard with no n, no reference, and no timestamp is marketing. Score sets rather than calls, publish the losses in the same table as the wins, and treat a single confident hit as the least informative event in forecasting.

  • Check the sample size and how many independent events it covers.
  • Find the reference model and confirm it was chosen in advance.
  • Look for the calibration curve, not just the headline number.
  • Confirm the forecasts were timestamped before resolution.

Get the Pro launch email

This guide stays free. Join the list for one email when the forecast feed, machine-readable research files, alerts, and saved research workflows open.

One email at launch. No weekly newsletter, no spam. Address handling is covered by the privacy policy.