⚽ football-bot

← Dashboard

How the prediction model works

Every Home/Draw/Away percentage on the dashboard — and everything built on top of it (same-game parlays, the match recommendation, suggested bets, the custom scoring parlay) — comes from one pipeline, run fresh for each match. The short version below is for anyone; the numbered walkthrough after it is the same thing again, in full technical detail, for anyone who wants to check the actual mechanics.

The short version

  1. Recent form matters most. The bot looks at how many goals each team has scored and conceded this season, weighting recent matches heavily — a result from six weeks ago barely counts anymore, but last weekend's game counts a lot.
  2. It looks past lucky scorelines. A 3-0 win built on one good chance and two deflections isn't treated as "true" 3-0 form. Where the data allows it, the bot partly corrects for scores that were flattering or unlucky, not just the final tally.
  3. New season? It leans on last season a bit. In the first few weeks of a season there isn't enough fresh data to judge a team fairly, so the bot borrows a little from how that team did last season — and lets that fade out naturally as real games pile up.
  4. Optional: transfer market value. If turned on, a small nudge comes from how much a squad is worth on the transfer market — a rough stand-in for overall talent, used only until the team has actually played enough games to speak for itself.
  5. Missing players count against a team. Injured attackers mean the bot expects fewer goals from that side; missing defenders or the goalkeeper mean it expects to concede more. A key regular missing matters more than a bench player.
  6. Head-to-head history gets a small say. If one team has historically had the upper hand in this exact matchup, that nudges things slightly — but it takes several past meetings for that to really move the needle.
  7. Home advantage is measured, not assumed. Rather than giving every home team a flat bonus, the bot uses this league's own real pattern of how much better teams tend to do at home versus away.
  8. All of that turns into "how many goals do we expect," then odds. Once it has an expected number of goals for each side, it works out the odds of every realistic final score and adds up which ones count as a home win, a draw, or an away win.
  9. It deliberately doesn't over-trust itself. When checked against a full past season, the bot's raw predictions turned out too confident — when it said "90% sure," it was only right about 84% of the time. So every prediction gets pulled back a bit toward "what usually happens in this league," more so the more confident it feels. Less dramatic, more honest.

Bottom line: the percentages are the bot's best honest guess from real recent form, missing players, head-to-head history, and home advantage — kept deliberately humble rather than overconfident. Everything else on the dashboard (parlay ideas, same-game bet suggestions, the "best pick," suggested stakes) is this same guess, just read in different ways — never a separate calculation.

See the full technical walkthrough ↓

The full technical walkthrough

For anyone who wants the actual mechanics: the same nine points above, spelled out with the real method, constants, and formulas.

1. Team ratings

A Maher (1982) attack/defense model, not Dixon-Coles yet at this stage. Each team gets one attack multiplier and one defense multiplier, relative to the league average — not split by home/away venue; home advantage is captured separately, as the league's own real average home and away goals, fit directly from results. Built fresh from every finished match of the current season:

2. Expected goals

λ_home = league_avg_home_goals × home_attack × away_defense (and the mirror for λ_away) — each side's own attack times the opponent's defense, scaled onto the league's real home/away scoring baseline.

3. Availability

A side missing key players gets its λ discounted before anything else runs. Attackers and midfielders reduce attack; defenders and goalkeepers worsen defense — each missing player weighted by their share of the squad's total season minutes, the closest available proxy for "how much do we miss them." The combined discount is capped so several simultaneous injuries can't overcorrect.

4. Head-to-head

A shrinkage-weighted average goal difference from this pair's past meetings (any venue), shifted toward whoever's historically dominated the fixture — a single past meeting counts for ~17%, five for 50%, ten for ~67%. Capped at 1.5 goals of swing, and applied as a zero-sum nudge (added to λ_home, subtracted from λ_away) so total expected goals barely moves. Unlike season form, this isn't time-decayed: an old rivalry pattern is still a meaningful signal.

5. Scoreline grid

Independent Poisson(λ_home) × Poisson(λ_away) for every scoreline up to 8-8, with the Dixon & Coles (1997) ρ correction applied to the four low-scoring cells (0-0, 1-0, 0-1, 1-1), where plain independent Poisson is known to under-predict draws and over-predict 1-0s. This same grid is what the same-game parlays and the "most likely score" pick read directly, so those always stay consistent with the Home/Draw/Away percentages by construction — they're never a separate calculation.

6. Calibration

The raw grid is blended 70% raw / 30% league baseline (a grid built the same way from just the league's average home/away goals, i.e. what two exactly-average teams would produce). A 380-match walk-forward backtest found the unshrunk model overconfident whenever it favoured a result: picks made at "90% confident" were only right 83.9% of the time, and the gap widened further out — at 65–80% confident, only 54.7% actual. This step pulls every prediction back toward the league's own base rate, proportional to how far out on a limb it is.

What's tuned, what isn't

None of the constants above — the 45-day half-life, the 0.5 xG blend, the -0.1 Dixon-Coles ρ, the head-to-head shrinkage and cap, the injury cap, the 0.7 calibration weight — are fitted to this data. They're literature-typical starting points, each flagged in the code as a candidate for tuning with scripts/backtest.py once there's more of a track record to tune against. A "strong" pick is not automatically a good bet: a heavy favourite's short price often gives negative expected value even when the model agrees with the market.

↑ Back to the short version