Algorithms

Most sites show you one model and ask you to trust it. We run several, put them through an identical exam, and show you the marks — including the ones we would rather not publish. Pick whichever you want to follow.

Over 1900 held-out matches the best engine scores 0.1981 against the closing market's 0.1949 — closer to the sharpest benchmark there is than any public model we know, and every week of it is on the record, including the bad ones.
Algorithm of the week
xGB Momentum

Best ranked-probability score over the last 8 weeks of matches (since 2026-03-29). This changes as form changes — which is rather the point of running more than one model. The full week-by-week record is on the weekly record.

Every model is fitted on 2016–2020 only, then run forward week by week across 2021–2025 without ever seeing those seasons during tuning. Same data, same refit schedule, same metrics — the only thing that differs is the algorithm.

1

Closing market

benchmark — not selectable

Not ours: the betting market's closing prices, the strongest public predictor in football. The bar every engine is judged against.

RPS
0.19491
95% 0.1889–0.2007
Accuracy
55.2%
1900 matches
Log loss
0.9597
Brier
0.5700
2

xGB Consensus

The engines vote. Disagreement between them is information, and the vote is settled before the season starts — never adjusted in hindsight.

RPS
0.19815
95% 0.1921–0.2040
Accuracy
54.4%
1900 matches
Log loss
0.9702
Brier
0.5767

Against the reigning champion engine over the same 1900 matches: better by 0.00009 RPS (p = 0.4331) — inside noise, so we do not claim a difference.

3

xGB Core

The house engine. Reads the quality of every chance a side creates and concedes — not just results — adjusted for opponent and venue, weighted toward now. The champion until another engine beats it where it counts.

RPS
0.19824
95% 0.1921–0.2042
Accuracy
54.5%
1900 matches
Log loss
0.9707
Brier
0.5769
4

xGB Momentum

Short memory on purpose: what have you done lately? Streaks and slumps count for more here than anywhere else in the room.

RPS
0.20047
95% 0.1941–0.2068
Accuracy
54.2%
1900 matches
Log loss
0.9790
Brier
0.5825

Against the reigning champion engine over the same 1900 matches: worse by 0.00223 RPS (p = 0.0316).

5

xGB Ladder

A long-memory strength ladder: every result moves both clubs up or down, big upsets move them further. Slow to convince, hard to fool.

RPS
0.20075
95% 0.1946–0.2067
Accuracy
53.6%
1900 matches
Log loss
0.9781
Brier
0.5823

Against the reigning champion engine over the same 1900 matches: worse by 0.00251 RPS (p = 0.0268).

Teams are not players

The model that best picks match outcomes is not automatically the one you want for “will he score?”. A match forecast depends on the difference between two sides; a player’s chance of scoring depends on the level of his own team’s expected goals. A model can get the difference right while running high or low on the level, and the match leaderboard would never show it. So we grade each model on players too, over 9,008 appearances.

ModelGoal BrierBiasTop-10 boardSeparation
xGB Core0.07503-0.98%27.5%2.23×
xGB Momentum0.07547-0.84%28.9%2.39×
xGB LadderCannot answer this question — it rates teams, not scorelines, so it has no expected-goals number to hand a player model. We leave the row empty rather than invent one.

Separation is how much more often the model’s top-ten board returns a goal or assist than everyone else — the number that actually matters when you are picking a captain. The two are close, and they disagree about which is best: the champion is better calibrated, while rolling form separates its top picks more sharply. Graded only on players who actually appeared — the projection is conditional on playing, and we do not claim to know the teamsheet.

Tuned on 2016, 2017, 2018, 2019, 2020 · evaluated on 2021, 2022, 2023, 2024, 2025 · generated 2026-08-15