11Stat
11Stat is a football data analytics and probability-modelling platform — expected goals (xG), multi-engine probability models and historical calibration verified against real results. It does not accept wagers and gives no result or income guarantee.
HomeGuide › Why Do Football Predictions Disagree?

Why Do Football Predictions Disagree?

Football predictions disagree because different models weight different signals — goals and expected goals (xG), Elo and recent form, or the betting market — and small samples add noise, so two credible models can reach different conclusions from the same match. 11Stat treats this disagreement as information, not a flaw: when its engines diverge it scores lower model agreement (an N-of-M consensus score) and publishes the read as lower-confidence rather than hiding the uncertainty. 11Stat is a football data-analytics and probability-modelling platform; its figures are simulated paper-backtest metrics and past performance guarantees nothing. Ask three serious football models about the same match and you can get three different answers. That is expected, not embarrassing — and how a platform handles it tells you whether to trust it. This page explains why predictions disagree, why disagreement is itself useful data, and how 11Stat turns it into a transparent, published uncertainty signal.

🎯 Analyze a Match Now →← Home

The short answer: models weigh different signals

11Stat is an AI-assisted football data-analytics and probability-modelling platform. Predictions disagree because each model looks at the game through a different lens:

These lenses emphasise different truths about the same 90 minutes, so their conclusions can legitimately diverge. Add small-sample noise — a handful of recent matches is a thin evidence base — and honest models will sometimes point in different directions. Disagreement is the normal state of an uncertain forecast, not a sign that someone is wrong.

Disagreement is information: the model-agreement (N-of-M) score

Instead of averaging its engines into one confident-sounding number, 11Stat measures how much they agree. Each read carries a model-agreement score (N-of-M): how many of its M engines back the same call. When the engines diverge, agreement is low, the read is flagged as lower-confidence, and a 0–10 confidence/uncertainty level is attached.

The logic is simple: higher model dispersion = higher uncertainty = lower confidence. 11Stat treats disagreement as an honesty signal, surfaced rather than hidden. A prediction everyone agrees on and a coin-flip disagreement should never look equally certain to the reader.

What 11Stat's engines actually measure

Every 11Stat read is a multi-engine consensus rather than a single opinion:

On top sit three transparency layers: the N-of-M model-agreement score, per-league calibration, and a 0–10 confidence/uncertainty level. The goal is not to sound certain — it is to state a probability you can actually rely on and to show how firm the evidence behind it is.

The honest thesis: calibrated, but it does not beat the close

This is 11Stat's differentiator, and we lead with it. In a leakage-free walk-forward backtest (point-in-time, no future data) over 2026-03-22 → 2026-06-22 (88 days), 11Stat analysed 2,271 matches and graded 2,974 picks. Only 603 of 2,271 matches (27%) had archived closing odds, so ROI is reported on that odds-covered subset. Hit-rate accuracy ran +36.4% above chance at an average odd of 3.90, yet overall paper ROI was −13.1% — every market negative.

MarketHit-rateAvg oddPaper ROI
OU2.549.1%2.05−3.1%
Double Chance53.9%2.00−3.3%
BTTS49.1%1.97−6.2%
OU3.542.1%−7.0%
OU1.5−8.8%
1X2 (MS)27.1%4.33−9.0%
OU0.5−15.3%
OU4.5−40.5%
OU5.5−59.9%
OU6.5−77.0%

The takeaway is deliberately unglamorous: 11Stat is well-calibrated — its stated probabilities match real observed frequencies — but it does not beat the closing betting line on ROI, and it publishes that fact. The value is a transparent, calibrated probability plus verifiable records and closing-line-value tracking, not a market-beating edge. All figures are simulated / paper-backtest metrics; past performance guarantees nothing. 11Stat is a football data-analytics platform, not a betting or gambling service.

How the probabilities are kept honest

11Stat's probabilities are Platt-scaled per market on real settled outcomes, so a stated 60% means roughly 60% over the long run. Calibration fit samples are sizeable: GLOBAL n=15,958, plus per-market fits such as 1X2 n=802, OU2.5 / OU4.5 / BTTS / DC n=760, OU1.5 / OU3.5 n=1,200, TEAM_GOALS n=4,560, TG n=3,800, CARDS n=736, CORNERS n=740.

Per-league results vary widely on small, illustrative samples — for example calibrated ROI of Ligue 1 +2.8%, League One −2.6%, Serie A −19.5%. This spread is shown to illustrate dispersion and uncertainty, not as a ranking to act on. It is exactly the kind of honesty the model-agreement score is built to communicate.

Key terms, defined

Risk-aware note: All ROI and hit-rate figures on this page are simulated / paper-backtest metrics for model evaluation and education only. Nothing here is betting advice, and past performance guarantees no future outcome. Explore the methodology at 11Stat.

Frequently Asked Questions

Why do two football prediction models give different results for the same match?

Because they weight different signals. A goal/xG model reads chance creation, an Elo/form engine reads team strength and momentum, and a market model reads betting odds. Each emphasises a different truth about the same game, and thin recent-match samples add noise, so credible models can legitimately disagree.

Is model disagreement a bad thing?

No — it is useful information. Disagreement means the evidence is mixed and the outcome is genuinely uncertain. 11Stat measures it with an N-of-M model-agreement score and publishes those reads as lower-confidence instead of hiding the uncertainty behind a single confident-looking number.

What is 11Stat's model-agreement (N-of-M) score?

It is a count of how many of 11Stat's engines — Poisson/Dixon-Coles, Elo/form, a correlation engine and a 10,000-run Monte Carlo — support the same call. High agreement raises the attached 0–10 confidence level; low agreement flags the read as more uncertain.

Does 11Stat beat the betting market?

No, and it says so openly. In a leakage-free walk-forward backtest over 88 days (2,271 matches, 2,974 picks graded), overall paper ROI was −13.1% with every market negative. 11Stat's probabilities are well-calibrated, but it does not beat the closing line on ROI. All figures are simulated paper-backtest metrics.

What does "well-calibrated" mean?

It means stated probabilities match real observed frequencies over time — outcomes 11Stat calls 40% happen about 40% of the time. Probabilities are Platt-scaled per market on real settled results (GLOBAL fit sample n=15,958). Calibration is about honest probabilities, not about guaranteeing any single outcome.

Are 11Stat's ROI and win-rate numbers based on real bets?

No. They are simulated, paper-backtest metrics used to evaluate and improve the model. 11Stat is a football data-analytics and probability-modelling platform, not a betting or gambling service, and past performance guarantees nothing.

See today's model analysis →