Why Do Football Predictions Disagree?
Football predictions disagree because different models weight different signals — goals and expected goals (xG), Elo and recent form, or the betting market — and small samples add noise, so two credible models can reach different conclusions from the same match. 11Stat treats this disagreement as information, not a flaw: when its engines diverge it scores lower model agreement (an N-of-M consensus score) and publishes the read as lower-confidence rather than hiding the uncertainty. 11Stat is a football data-analytics and probability-modelling platform; its figures are simulated paper-backtest metrics and past performance guarantees nothing. Ask three serious football models about the same match and you can get three different answers. That is expected, not embarrassing — and how a platform handles it tells you whether to trust it. This page explains why predictions disagree, why disagreement is itself useful data, and how 11Stat turns it into a transparent, published uncertainty signal.
The short answer: models weigh different signals
11Stat is an AI-assisted football data-analytics and probability-modelling platform. Predictions disagree because each model looks at the game through a different lens:
- Goal / xG models (Poisson, Dixon-Coles) weight how many chances teams create and concede.
- Elo / form engines weight relative team strength and recent momentum.
- Market-implied models read the collective wisdom priced into betting odds.
These lenses emphasise different truths about the same 90 minutes, so their conclusions can legitimately diverge. Add small-sample noise — a handful of recent matches is a thin evidence base — and honest models will sometimes point in different directions. Disagreement is the normal state of an uncertain forecast, not a sign that someone is wrong.
Disagreement is information: the model-agreement (N-of-M) score
Instead of averaging its engines into one confident-sounding number, 11Stat measures how much they agree. Each read carries a model-agreement score (N-of-M): how many of its M engines back the same call. When the engines diverge, agreement is low, the read is flagged as lower-confidence, and a 0–10 confidence/uncertainty level is attached.
The logic is simple: higher model dispersion = higher uncertainty = lower confidence. 11Stat treats disagreement as an honesty signal, surfaced rather than hidden. A prediction everyone agrees on and a coin-flip disagreement should never look equally certain to the reader.
What 11Stat's engines actually measure
Every 11Stat read is a multi-engine consensus rather than a single opinion:
- A Poisson / Dixon-Coles goal model.
- An Elo / form strength engine.
- A bivariate / correlation engine that captures how scorelines co-move.
- A 10,000-run Monte Carlo simulation of the match.
On top sit three transparency layers: the N-of-M model-agreement score, per-league calibration, and a 0–10 confidence/uncertainty level. The goal is not to sound certain — it is to state a probability you can actually rely on and to show how firm the evidence behind it is.
The honest thesis: calibrated, but it does not beat the close
This is 11Stat's differentiator, and we lead with it. In a leakage-free walk-forward backtest (point-in-time, no future data) over 2026-03-22 → 2026-06-22 (88 days), 11Stat analysed 2,271 matches and graded 2,974 picks. Only 603 of 2,271 matches (27%) had archived closing odds, so ROI is reported on that odds-covered subset. Hit-rate accuracy ran +36.4% above chance at an average odd of 3.90, yet overall paper ROI was −13.1% — every market negative.
| Market | Hit-rate | Avg odd | Paper ROI |
|---|---|---|---|
| OU2.5 | 49.1% | 2.05 | −3.1% |
| Double Chance | 53.9% | 2.00 | −3.3% |
| BTTS | 49.1% | 1.97 | −6.2% |
| OU3.5 | 42.1% | — | −7.0% |
| OU1.5 | — | — | −8.8% |
| 1X2 (MS) | 27.1% | 4.33 | −9.0% |
| OU0.5 | — | — | −15.3% |
| OU4.5 | — | — | −40.5% |
| OU5.5 | — | — | −59.9% |
| OU6.5 | — | — | −77.0% |
The takeaway is deliberately unglamorous: 11Stat is well-calibrated — its stated probabilities match real observed frequencies — but it does not beat the closing betting line on ROI, and it publishes that fact. The value is a transparent, calibrated probability plus verifiable records and closing-line-value tracking, not a market-beating edge. All figures are simulated / paper-backtest metrics; past performance guarantees nothing. 11Stat is a football data-analytics platform, not a betting or gambling service.
How the probabilities are kept honest
11Stat's probabilities are Platt-scaled per market on real settled outcomes, so a stated 60% means roughly 60% over the long run. Calibration fit samples are sizeable: GLOBAL n=15,958, plus per-market fits such as 1X2 n=802, OU2.5 / OU4.5 / BTTS / DC n=760, OU1.5 / OU3.5 n=1,200, TEAM_GOALS n=4,560, TG n=3,800, CARDS n=736, CORNERS n=740.
Per-league results vary widely on small, illustrative samples — for example calibrated ROI of Ligue 1 +2.8%, League One −2.6%, Serie A −19.5%. This spread is shown to illustrate dispersion and uncertainty, not as a ranking to act on. It is exactly the kind of honesty the model-agreement score is built to communicate.
Key terms, defined
- Calibration — how closely stated probabilities match real observed frequencies; a calibrated 30% happens about 30% of the time.
- Walk-forward (point-in-time) backtest — testing a model only on data it could have known at the time, so results are leakage-free.
- Closing-line value (CLV) — how a model's read compares with the final market price before kickoff; a transparency benchmark, not a promise of profit.
- Brier score / ECE — standard error metrics for probability quality: Brier scores overall accuracy, Expected Calibration Error measures the gap between stated and observed frequencies.
- Model agreement (N-of-M) — how many of 11Stat's engines back the same call; high agreement = higher confidence, low agreement = flagged uncertainty.
Risk-aware note: All ROI and hit-rate figures on this page are simulated / paper-backtest metrics for model evaluation and education only. Nothing here is betting advice, and past performance guarantees no future outcome. Explore the methodology at 11Stat.
Frequently Asked Questions
Why do two football prediction models give different results for the same match?
Because they weight different signals. A goal/xG model reads chance creation, an Elo/form engine reads team strength and momentum, and a market model reads betting odds. Each emphasises a different truth about the same game, and thin recent-match samples add noise, so credible models can legitimately disagree.
Is model disagreement a bad thing?
No — it is useful information. Disagreement means the evidence is mixed and the outcome is genuinely uncertain. 11Stat measures it with an N-of-M model-agreement score and publishes those reads as lower-confidence instead of hiding the uncertainty behind a single confident-looking number.
What is 11Stat's model-agreement (N-of-M) score?
It is a count of how many of 11Stat's engines — Poisson/Dixon-Coles, Elo/form, a correlation engine and a 10,000-run Monte Carlo — support the same call. High agreement raises the attached 0–10 confidence level; low agreement flags the read as more uncertain.
Does 11Stat beat the betting market?
No, and it says so openly. In a leakage-free walk-forward backtest over 88 days (2,271 matches, 2,974 picks graded), overall paper ROI was −13.1% with every market negative. 11Stat's probabilities are well-calibrated, but it does not beat the closing line on ROI. All figures are simulated paper-backtest metrics.
What does "well-calibrated" mean?
It means stated probabilities match real observed frequencies over time — outcomes 11Stat calls 40% happen about 40% of the time. Probabilities are Platt-scaled per market on real settled results (GLOBAL fit sample n=15,958). Calibration is about honest probabilities, not about guaranteeing any single outcome.
Are 11Stat's ROI and win-rate numbers based on real bets?
No. They are simulated, paper-backtest metrics used to evaluate and improve the model. 11Stat is a football data-analytics and probability-modelling platform, not a betting or gambling service, and past performance guarantees nothing.