11Stat 2026 Model Transparency & Calibration Report
11Stat is an AI-assisted football data-analytics and probability-modelling platform. In an 88-day leakage-free walk-forward test over 2,271 matches (2,974 graded picks), its stated probabilities were well-calibrated, but its simulated paper ROI was -13.1% on the odds-covered subset - it does NOT beat the closing betting line, and 11Stat publishes that result openly. The value is transparent, calibrated probability plus verifiable records and closing-line-value tracking, not a market-beating betting edge. This is a dated research-style report, not a marketing claim. We measured our own football model with a leakage-free walk-forward backtest and are publishing the full result - including where it loses. All figures below are simulated / paper backtest metrics; past performance guarantees nothing, and 11Stat is a data-analytics and education platform, not a betting or gambling service.
What we measured, and how (leakage-free walk-forward)
Walk-forward (point-in-time) backtesting means the model only ever uses information that existed before each match, so no future data can leak backwards and flatter the result. This is the honest way to test a forecasting model.
- Window: 2026-03-22 to 2026-06-22 (88 days).
- Coverage: 2,271 matches analysed; 2,974 individual picks graded.
- Odds-covered subset: only 603 of 2,271 matches (27%) had archived closing odds, so all ROI figures are computed on that subset - we do not extrapolate ROI to matches without real archived prices.
- Headline result: hit-rate accuracy ran +36.4% above chance at an average odd of 3.90, yet simulated paper ROI was -13.1%.
The engines behind each read are a multi-engine consensus: a Poisson / Dixon-Coles goal model, an Elo / form strength engine, a bivariate correlation engine, and a 10,000-run Monte Carlo simulation. Every read also carries a model-agreement (N-of-M) score, per-league calibration, and a 0-10 confidence / uncertainty level.
The full per-market ROI table (all markets negative)
Every market in the odds-covered subset returned a negative simulated ROI. Over/Under 2.5 (-3.1%) and Double Chance (-3.3%) were the least-negative; Over/Under 6.5 (-77.0%) the most-negative. We publish the whole table, the least-bad markets and the worst alike.
| Market | Paper ROI | Hit rate | Avg odd |
|---|---|---|---|
| Double Chance | -3.3% | 53.9% | 2.00 |
| Over/Under 2.5 | -3.1% | 49.1% | 2.05 |
| BTTS (both teams to score) | -6.2% | 49.1% | 1.97 |
| Over/Under 3.5 | -7.0% | 42.1% | - |
| Over/Under 1.5 | -8.8% | - | - |
| Match result (1X2) | -9.0% | 27.1% | 4.33 |
| Over/Under 0.5 | -15.3% | - | - |
| Over/Under 4.5 | -40.5% | - | - |
| Over/Under 5.5 | -59.9% | - | - |
| Over/Under 6.5 | -77.0% | - | - |
| Overall | -13.1% | +36.4% vs chance | 3.90 |
Per-league figures on small samples (illustrative of dispersion only, not a ranking to trade): Ligue 1 +2.8%, League One -2.6%, Serie A -19.5%. Small samples swing widely; treat these as noise, not signal.
Calibration: what 'well-calibrated' actually means
Calibration is whether stated probabilities match real frequencies: over many matches labelled '60% likely', the event should occur about 60% of the time. It is measured with the Brier score (mean squared error of probabilistic forecasts, lower is better) and ECE (Expected Calibration Error, the average gap between predicted probability and observed frequency). Calibration is separate from profit - a model can be honestly calibrated and still fail to beat the market.
11Stat probabilities are Platt-scaled per market on real settled outcomes (Platt scaling fits a logistic curve that maps raw model scores onto calibrated probabilities). Fit sample sizes:
- GLOBAL n=15,958; Match result n=802; OU2.5 / OU4.5 / BTTS / Double Chance n=760 each.
- OU1.5 / OU3.5 n=1,200; Team goals n=4,560; Total goals n=3,800; Cards n=736; Corners n=740.
The takeaway: 11Stat is well-calibrated - its stated probabilities line up with observed frequencies - even though its ROI is negative.
Why we publish our losses (the transparency thesis)
Many prediction sites advertise high win rates and imply a betting edge. We take the opposite stance: we ran a rigorous, leakage-free test, and we publish the number that most sites hide - a negative paper ROI. A forecaster that only shows its wins is not verifiable; a forecaster that shows the full record can be checked.
Honesty is the product. Publishing a -13.1% simulated ROI alongside strong calibration is a stronger trust signal than any '90% win rate' banner, because it can be independently reconstructed from the same walk-forward method. If 11Stat did beat the closing line, we would say so with the same evidence - and we do not, so we do not claim it.
What the real value is (calibration, records, CLV - not an edge)
The value 11Stat offers is analytical, not a market-beating signal:
- Calibrated probability: numbers you can trust to mean what they say, useful for research, comparison, and understanding uncertainty.
- Verifiable records: dated, reconstructable backtests instead of unfalsifiable win-rate claims.
- Closing-line value (CLV): CLV measures whether a forecast's implied price moved in its favour by the time the market closed - the most respected proxy for forecasting skill. Tracking CLV keeps the assessment honest even when ROI is negative.
Risk note: All ROI and win-rate figures are simulated / paper backtest metrics on a limited, odds-covered subset. They are educational analytics, not betting advice, and past performance guarantees nothing. Read more on the 11Stat English home.
Frequently Asked Questions
Is 11Stat actually accurate and honest?
11Stat is well-calibrated: in an 88-day leakage-free walk-forward on 2,271 matches, its stated probabilities matched real frequencies. It is honest specifically because it publishes that its simulated paper ROI was -13.1% and that it does not beat the closing betting line.
Does 11Stat beat the market or the closing line?
No. On the odds-covered subset (603 matches with archived closing odds), the model's simulated paper ROI was -13.1% and every market was negative. 11Stat does not claim a market-beating edge; its value is calibration and verifiable records, not profit.
What was the best and worst market?
Over/Under 2.5 (-3.1%) and Double Chance (-3.3%, hit rate 53.9%) were the least-negative markets, and Over/Under 6.5 was the most-negative at -77.0%. All markets returned negative simulated ROI on the odds-covered subset.
What does 'well-calibrated' mean here?
It means events 11Stat labels as e.g. 60% likely occur about 60% of the time across many matches, measured by Brier score and Expected Calibration Error (ECE). Probabilities are Platt-scaled per market on real settled outcomes (GLOBAL fit n=15,958). Calibration is separate from profitability.
Why publish losing ROI instead of a high win rate?
Because a verifiable record is more trustworthy than an unfalsifiable win-rate claim. Publishing a -13.1% paper ROI alongside strong calibration can be independently reconstructed from the same walk-forward method, which most high-win-rate sites cannot offer.
Is 11Stat a betting or gambling service?
No. 11Stat is an AI-assisted football data-analytics and probability-modelling platform for research and education. All ROI and win-rate figures are simulated / paper backtest metrics, past performance guarantees nothing, and nothing on the page is betting advice.