How the prediction model works
Every Home/Draw/Away percentage on the dashboard — and everything built on top of it (same-game parlays, the match recommendation, suggested bets, the custom scoring parlay) — comes from one pipeline, run fresh for each match. The short version below is for anyone; the numbered walkthrough after it is the same thing again, in full technical detail, for anyone who wants to check the actual mechanics.
The short version
- Recent form matters most. The bot looks at how many goals each team has scored and conceded this season, weighting recent matches heavily — a result from six weeks ago barely counts anymore, but last weekend's game counts a lot.
- It looks past lucky scorelines. A 3-0 win built on one good chance and two deflections isn't treated as "true" 3-0 form. Where the data allows it, the bot partly corrects for scores that were flattering or unlucky, not just the final tally.
- New season? It leans on last season a bit. In the first few weeks of a season there isn't enough fresh data to judge a team fairly, so the bot borrows a little from how that team did last season — and lets that fade out naturally as real games pile up.
- Optional: transfer market value. If turned on, a small nudge comes from how much a squad is worth on the transfer market — a rough stand-in for overall talent, used only until the team has actually played enough games to speak for itself.
- Missing players count against a team. Injured attackers mean the bot expects fewer goals from that side; missing defenders or the goalkeeper mean it expects to concede more. A key regular missing matters more than a bench player.
- Head-to-head history gets a small say. If one team has historically had the upper hand in this exact matchup, that nudges things slightly — but it takes several past meetings for that to really move the needle.
- Home advantage is measured, not assumed. Rather than giving every home team a flat bonus, the bot uses this league's own real pattern of how much better teams tend to do at home versus away.
- All of that turns into "how many goals do we expect," then odds. Once it has an expected number of goals for each side, it works out the odds of every realistic final score and adds up which ones count as a home win, a draw, or an away win.
- It deliberately doesn't over-trust itself. When checked against a full past season, the bot's raw predictions turned out too confident — when it said "90% sure," it was only right about 84% of the time. So every prediction gets pulled back a bit toward "what usually happens in this league," more so the more confident it feels. Less dramatic, more honest.
Bottom line: the percentages are the bot's best honest guess from real recent form, missing players, head-to-head history, and home advantage — kept deliberately humble rather than overconfident. Everything else on the dashboard (parlay ideas, same-game bet suggestions, the "best pick," suggested stakes) is this same guess, just read in different ways — never a separate calculation.
The full technical walkthrough
For anyone who wants the actual mechanics: the same nine points above, spelled out with the real method, constants, and formulas.
1. Team ratings
A Maher (1982) attack/defense model, not Dixon-Coles yet at this stage. Each team gets one attack multiplier and one defense multiplier, relative to the league average — not split by home/away venue; home advantage is captured separately, as the league's own real average home and away goals, fit directly from results. Built fresh from every finished match of the current season:
-
Recency. Each match is weighted by
exp(-ln2 × age_days / 45)— a 45-day half-life. Recent form dominates, but a match from 6–7 weeks back still counts for something, fading out smoothly rather than dropping off a hard cutoff. -
Expected goals (xG) blending. Where a provider has xG for a
match, the goals feeding the rating are
0.5 × actual + 0.5 × xG, so a scoreline built on one great chance and two deflections counts as weaker form than the result alone suggests. Falls back to the actual score when xG isn't available (real coverage is ~84% on API-Football, not 100%). - New-season seeding. The previous season's last 50 matches (roughly its last 5 matchdays) are prepended before the current season's own results, so matchday 1 isn't predicted from an all-neutral table. Recency decay fades this seed out on its own as the new season accumulates games. A team with no history at all (freshly promoted) gets a neutral rating.
- Transfermarkt prior (currently enabled, weight 0.15). Nudges attack up and defense down together (or vice versa) for a team whose squad market value sits away from the league's average — log-scaled, capped, and fading as real matches are played (to about a third of its starting size after 8 matches, and under a tenth after 20). Only applied to a team that already has at least one finished match in the ratings window; a team with zero games — the case this prior exists for — is still purely neutral until its first result arrives. See the README's "Transfermarkt squad-value prior" section for the snapshot format and freshness rules.
2. Expected goals
λ_home = league_avg_home_goals × home_attack × away_defense
(and the mirror for λ_away) — each side's own attack
times the opponent's defense, scaled onto the league's real
home/away scoring baseline.
3. Availability
A side missing key players gets its λ discounted before anything else runs. Attackers and midfielders reduce attack; defenders and goalkeepers worsen defense — each missing player weighted by their share of the squad's total season minutes, the closest available proxy for "how much do we miss them." The combined discount is capped so several simultaneous injuries can't overcorrect.
4. Head-to-head
A shrinkage-weighted average goal difference from this pair's past meetings (any venue), shifted toward whoever's historically dominated the fixture — a single past meeting counts for ~17%, five for 50%, ten for ~67%. Capped at 1.5 goals of swing, and applied as a zero-sum nudge (added to λ_home, subtracted from λ_away) so total expected goals barely moves. Unlike season form, this isn't time-decayed: an old rivalry pattern is still a meaningful signal.
5. Scoreline grid
Independent Poisson(λ_home) × Poisson(λ_away) for every scoreline up to 8-8, with the Dixon & Coles (1997) ρ correction applied to the four low-scoring cells (0-0, 1-0, 0-1, 1-1), where plain independent Poisson is known to under-predict draws and over-predict 1-0s. This same grid is what the same-game parlays and the "most likely score" pick read directly, so those always stay consistent with the Home/Draw/Away percentages by construction — they're never a separate calculation.
6. Calibration
The raw grid is blended 70% raw / 30% league baseline (a grid built the same way from just the league's average home/away goals, i.e. what two exactly-average teams would produce). A 380-match walk-forward backtest found the unshrunk model overconfident whenever it favoured a result: picks made at "90% confident" were only right 83.9% of the time, and the gap widened further out — at 65–80% confident, only 54.7% actual. This step pulls every prediction back toward the league's own base rate, proportional to how far out on a limb it is.
What's tuned, what isn't
None of the constants above — the 45-day half-life, the 0.5 xG
blend, the -0.1 Dixon-Coles ρ, the head-to-head shrinkage and cap,
the injury cap, the 0.7 calibration weight — are fitted to this
data. They're literature-typical starting points, each flagged in
the code as a candidate for tuning with scripts/backtest.py
once there's more of a track record to tune against. A "strong"
pick is not automatically a good bet: a heavy favourite's short
price often gives negative expected value even when the model
agrees with the market.