Model Calibration & Transparency
11Stat openly publishes the leakage-free (walk-forward) calibration and model-quality metrics of its football probability model, because the value of a trustworthy match-analysis tool is not a profit promise — it is calibrated probability and transparency.
Honest disclosure
The figures on this page are paper/simulation accuracy metrics. Past performance does not guarantee future results. 11Stat makes no "lock" or guaranteed-match claims and promises no profit or return — no model can guarantee a match outcome.
11Stat is not a betting or gambling service; it accepts no bets, sells no coupons and offers no betting advice. All outputs are football data analytics, modelling, simulation and educational content only. 18+.
Summary metrics (99-day walk-forward backtest)
Test window: 2026-03-22 → 2026-06-28 · point-in-time (each prediction made only with data available up to that day, no leakage) · 2,271 matches priced.
- 0.19–0.27 — Brier score range (statistical model quality; lower = more accurate probability)
- 12 engines — 3 core probability models + confidence-score consensus
- Leakage-free — walk-forward, point-in-time testing
What the model is genuinely good at: calibration
A model's worth is not just "getting it right" but how well its probabilities match reality. 11Stat's outputs are calibrated against real frequency. Measurements showed the model is overconfident in some markets, and a calibration layer corrects this:
- Draw-safety profile — n=841 · model 64% · real 50% · calibration error −13.9% · Brier 0.255
- Goal volume (2.5 line) — n=825 · model 58% · real 44% · calibration error −13.8% · Brier 0.265
- Goal volume (3.5 line) — n=810 · model 52% · real 39% · calibration error −12.9% · Brier 0.222
- Match outcome profile — n=1,317 · model 36% · real 24% · calibration error −12.2% · Brier 0.190
- Both-team scoring profile — n=737 · model 59% · real 46% · calibration error −12.1% · Brier 0.256
The calibration layer corrects this overconfidence with a shrink learned from historical data (Platt sigmoid slope, capped at ±0.15 against overfitting). The result: more realistic probabilities and a more meaningful confidence score.
What is the Model Confidence Score for?
The confidence score (0–10) combines how strongly several independent models (Poisson, Dixon-Coles, Elo/form) agree on the same outcome, data quality and calibration into a probability indicator. A higher score means the model assigned a stronger, more consistent probability to that scenario — not a "lock" guarantee. A smart user bases their own decision not on chance but on the highest confidence scores and the most model agreement — while knowing the outcome is never guaranteed.
Proof and transparency
Every published model signal is timestamped at publish time with a SHA256 fingerprint (an immutable snapshot). The result is added later as a separate record. Signals captured before kickoff are labelled "Verified Publish"; results archived afterward are "Archive Record" — so past performance is shown without inflation and is independently verifiable.
Who is 11Stat right for?
- Right user: someone who supports their own pre-match decision with calibrated probability, xG and a multi-engine confidence score, and who values transparency.
- Wrong expectation: someone looking for "guaranteed locks" or "sure wins" — no honest tool can provide that, and neither does 11Stat.
See the full About & Methodology, the FAQ, or try the live analysis panel on the home page.
Frequently Asked Questions
Does 11Stat guarantee match outcomes?
No. All figures on this page are paper/simulation accuracy metrics; past performance does not guarantee future results. 11Stat makes no "lock" or guaranteed-match claims and promises no profit or return — no model can guarantee a match outcome.
What is the Brier score?
The Brier score is a statistical quality metric that measures how close probability forecasts are to actual outcomes; a lower value means more accurate probabilities. In 11Stat's 99-day walk-forward backtest the Brier score range is 0.19–0.27.
What is a walk-forward (point-in-time) backtest?
A leakage-free testing method in which each prediction is made only with data available up to that day. The reported window covers 99 days, 2026-03-22 → 2026-06-28, with 2,271 matches priced.
What is calibration error and how is it corrected?
Calibration error is the gap between model probability and real observed frequency; measurements showed the model was overconfident in some markets. A calibration layer corrects this with a shrink learned from historical data (Platt sigmoid slope, capped at ±0.15).
Does the Model Confidence Score (0–10) mean a guarantee?
No. The confidence score is a probability indicator combining model agreement, data quality and calibration. A higher score means a stronger, more consistent probability assignment — not an outcome guarantee.