11Stat
11Stat is a football data analytics and probability-modelling platform — expected goals (xG), multi-engine probability models and historical calibration verified against real results. It does not accept wagers and gives no result or income guarantee.
HomeGuide › Model Calibration Explained

Model Calibration Explained

What is model calibration? Calibration measures whether the numbers a probability model speaks can be trusted: a well-calibrated model is right on about 70% of the events it calls 70%, and on about 40% of those it calls 40% — claiming neither more nor less than it earns. This page explains why calibration says more than any single accuracy percentage, how to read ECE (expected calibration error) and the Brier score, and how 11Stat's 0-10 confidence score is derived from this data. It is football data analytics, not betting advice.

🎯 Analyze a Match Now →← Home

What calibration means: is 70% really 70%?

A model saying "70%" is not information by itself; the real question is: when the model said 70% in the past, how often was it right? If the answer is about 70%, the model is calibrated — its numbers speak the same language as reality. If the answer is 55%, the model is overconfident; if it is 85%, the model is too timid.

Calibration is therefore a measure of honesty: it separates being bold from being right. In a high-variance sport like football it cannot be tested on one match; hundreds of predictions must be grouped into probability bands and each band compared with its realised frequency. 11Stat's entire probability chain is run through this test; the details live on the methodology page.

Why a single accuracy number misleads

"Our model is 68% accurate" sounds impressive and says almost nothing. A rote model that marks the favourite in every match can reach a similar figure, collecting hits without producing information. An accuracy percentage ignores both the difficulty of each prediction and the probability the model attached to it.

Calibration fills exactly those two gaps: it tests the 55% claims and the 85% claims separately and shows whether the model's confidence is deserved at every level. Two models can share one accuracy figure while one is calibrated and the other reckless — and over the long run, only the calibrated model's numbers support sound reasoning. The same distinction is critical when reading simulated ROI.

How to read a reliability diagram and ECE

A reliability diagram is calibration's X-ray: predictions are grouped into probability bands (50-60%, 60-70%, …), and for each band the model's average claim (x-axis) is plotted against the realised frequency (y-axis). Perfect calibration sits on the 45-degree diagonal; points below the diagonal expose overconfidence, points above it excessive caution.

ECE (expected calibration error) compresses that picture into one number: the average gap per band, weighted by how many predictions each band contains. An ECE near 0 says the numbers can be taken at face value; a growing ECE says the model's language is drifting from reality. Because no metric suffices alone, 11Stat reports ECE together with the Brier score and the sample size.

Brier score and log-loss: proper scoring rules

The Brier score squares the gap between each stated probability and the realised outcome (1 or 0) and averages it: closer to 0 is better. It rewards and punishes calibration and discrimination (sharpness) at the same time. Log-loss punishes overconfident misses far more brutally: a model that says 95% and is wrong pays dearly.

Both are "proper scoring rules": the model's best possible strategy is to state its true belief honestly — no points can be gained by inflating or shading the number. That property is the mathematical foundation of a calibration culture: an honest probability is, by the definition of the game, the most rewarding statement. The terms are also defined in the glossary.

How 11Stat's 0-10 confidence score is derived

The 0-10 confidence score on an analysis card is not one engine's feeling but the summary of a calibrated chain: the degree of agreement among 12 independent engines, the probability after the calibration layer, and the market's historical reliability merge into a single scale. How many engines point the same way ("N-of-M model consent") is displayed alongside it.

A high score does not mean "certain"; it means "on its historical report card, the model is more consistent on this type of call". Even scores above 7 miss with non-trivial frequency — in a calibrated system they must, otherwise the numbers would be inflated. The score's components are walked through on the how it works page.

How 11Stat recalibrates itself

Calibration is not a task you finish once; as leagues, seasons and squads change, a model's language drifts. 11Stat therefore maintains a rolling calibration store per league and per market: every settled match updates the relevant curve, and when the gap widens, the correction coefficients are tightened.

Two further safety mechanisms apply: markets whose calibration degrades are disabled automatically, and every historical window — including the bad ones — is published in the simulated metric reports. The goal is never to "look right all the time" but to protect the meaning of the numbers: where it says 70%, the long run should genuinely show something near 70%.

Frequently Asked Questions

Does a well-calibrated model guarantee correct predictions?

No, it does not. Calibration secures the honesty of the probabilities, not individual outcomes: about 30% of the events it calls 70% still go the other way — and that is the sign the system works. No statistical model can make the result of a football match certain.

Is a 10/10 confidence score a sure bet?

No — there is no such thing, and 11Stat never presents any output that way. A high confidence score only says that engine agreement and historical consistency are high; in a calibrated system even the top scores miss regularly. 11Stat offers no betting advice.

What is a good ECE value?

It depends on context; in a noisy domain like football, values below 0.05 are generally healthy, while anything above 0.10 deserves investigation. What matters is not one snapshot but the trajectory of ECE over time and the sample size it was measured on.

Can a model be accurate yet poorly calibrated?

Yes, and it is common. A model that always picks the favourite can collect a high hit-rate while inflating its probabilities to 90%-style claims — its numbers are then unusable as decision support. The reverse also exists: a modestly accurate model with trustworthy numbers is far more valuable for analysis.

What is the difference between the Brier score and ECE?

ECE measures calibration only — the match between stated probability and realised frequency. The Brier score additionally captures sharpness: how boldly, and how justifiably, the model departs from 50-50. Read together, they reveal both the model's honesty and its informational power.

How often does 11Stat recalibrate its models?

Continuously: every settled match updates the per-league, per-market calibration curves, and markets whose drift grows are disabled automatically. Periodic backtests additionally re-examine the whole chain end to end, and the results are reported including the negative windows.

See today's model analysis →