What is live, what is being tested, what is paused
Every model here was scored on past games it had never seen before it went live, is graded on every new game after, and can be paused or replaced by its own record. This page reads those receipts; it does not restate them.
Models currently live
Each sport's model families and where they stand right now.
Model status · NFL
What the statuses mean →- Game forecastsExperimentalExperimental · tested on past seasons
We publish a simulated score range and win chance for every game, clearly marked experimental (1 games forecast today under nfl-regular-season-public-v1). The win chance comes from a team rating that weights margin of victory. Tested on 4,281 past games it had never seen (2006–2021), it scored 0.629 on log loss (lower is better) against 0.642 for our previous rating, 0.693 for a coin flip and 0.610 for the sportsbooks' own odds — so it makes no claim to out-predict the sportsbook market. Every forecast is frozen before kickoff and settled against the official result.
What does this mean?
Published and labelled experimental: tested on past seasons, but not validated to out-predict the sportsbook market.
- Player rangesForward test runningBlind forward test running
Receptions, yards and touchdown chances are graded week by week against preregistered bars; a family that breaches them falls back to the previous rule automatically. 0 player-games graded so far.
What does this mean?
The model is being graded on new games as they happen, against the model it replaced. If it falls significantly behind, the previous model comes back automatically.
- Touchdown chancesHoldingExpected scorers match actual
Over 168 graded, this call has held up against its own expected-scorer count.
What does this mean?
Over the games graded so far, this call has done at least as well as its floor (a coin flip, or the model it replaced).
Model status · MLB
What the statuses mean →- Winner callsWatchWatch · slightly behind
Over 605 graded, this call is slightly behind a coin flip, not yet by enough to be sure. It has landed 50% of the time while showing about 57%.
What does this mean?
Over the games graded so far, this call is slightly behind its floor, but not by enough to be sure it is worse. It keeps publishing while the record grows.
- Game totalsPausedPaused · below a coin flip
Over 576 graded, this call has done worse than a coin flip by a clear margin. It has landed 49% of the time while showing about 59%.
What does this mean?
Over the games graded so far, this call has done worse than a coin flip by a clear margin. The public call is withdrawn; it is still made and graded every day and comes back when its record recovers.
- Run line callsHoldingHolding its record
Over 605 graded, this call has held up against a coin flip. It has landed 65% of the time while showing about 60%.
What does this mean?
Over the games graded so far, this call has done at least as well as its floor (a coin flip, or the model it replaced).
Model status · Premier League
What the statuses mean →- Match modelTested on past seasonsTested blind on nine past seasons
Scored on 3,420 Premier League matches from seasons it was never fit on, where it beat the previous model and a plain rating system. Its totals follow the league's recent scoring rate, so over-2.5 is the same for every match.
What does this mean?
Before it went live, this model was scored on past seasons it had never seen and beat the simpler rules it replaced. That is a bar, not a promise.
- Forward testForward test runningForward test · 0 of 60 matches
Every graded match scores the live model beside the one it replaced. A verdict needs 60 matches; 0 are in.
What does this mean?
The model is being graded on new games as they happen, against the model it replaced. If it falls significantly behind, the previous model comes back automatically.
- Live recordToo early to judgeToo early to judge
35 graded so far — not yet enough to judge against even odds.
What does this mean?
Not enough games have been graded to judge this call either way. The number is published; the verdict on it is not.
Model status · UFC
What the statuses mean →- Fight modelTested on past seasonsBacktest passed on held-out fights
Winner, method and round heads each beat their baseline on fights held out of the fit.
What does this mean?
Before it went live, this model was scored on past seasons it had never seen and beat the simpler rules it replaced. That is a bar, not a promise.
- Live recordToo early to judgeToo early to judge
31 graded so far — not yet enough to judge against a coin flip.
What does this mean?
Not enough games have been graded to judge this call either way. The number is published; the verdict on it is not.
Experiments and shadows
Blind forward tests grade a live model week by week; a shadow is scored privately and powers nothing public.
- NFL receptions · blind forward testForward test runningAccumulating · 0 of 300 player-games
Every week's forecast is committed before its first kickoff and graded against the official box scores; a family that breaches its preregistered bars falls back to the previous rule automatically.
- NFL receiving yards · blind forward testForward test runningAccumulating · 0 of 300 player-games
Every week's forecast is committed before its first kickoff and graded against the official box scores; a family that breaches its preregistered bars falls back to the previous rule automatically.
- NFL rushing yards · blind forward testForward test runningAccumulating · 0 of 300 player-games
Every week's forecast is committed before its first kickoff and graded against the official box scores; a family that breaches its preregistered bars falls back to the previous rule automatically.
- NFL passing yards · blind forward testForward test runningAccumulating · 0 of 300 player-games
Every week's forecast is committed before its first kickoff and graded against the official box scores; a family that breaches its preregistered bars falls back to the previous rule automatically.
- Premier League match model · blind forward testForward test runningAccumulating · 0 of 60 matches
The live match model is scored beside the model it replaced on every graded match. If it is significantly worse, the previous model publishes again from the next build.
- Premier League club-specific totals · private shadowResearch onlyResearch only · registered, first fixtures pending
A club-by-club total is scored privately beside the live model's constant total on every graded fixture. It powers no public number; adoption would be a separate decision.
Recent decisions
What the receipts decided, most recent first. A rejected candidate stays rejected; it is never retried by tweaking it.
- Sep 15, 2026 · EPLRejectedClub-specific match totals: rejected on the blind test, positive on a second look
Both candidates beat the constant total on every totals measure across four other leagues and the Premier League, but missed a preregistered result-calibration ceiling on the blind set. The verdict stands; the better candidate now runs as a private shadow.
- Sep 15, 2026 · EPLResearch onlyPremier League club-specific totals registered as a private shadow
Not adopted. Scored beside the live model on future fixtures for a later decision.
- Sep 14, 2026 · EPLAdoptedPremier League match model replaced by a rating model tested blind on nine past seasons
It beat the previous model and a plain rating system on 3,420 matches it was never fit on, and now runs a blind forward test against the model it replaced.
- Sep 14, 2026 · NFLSecond lookNFL player ranges: receptions, receiving yards, rushing yards eligible on a second look; passing yards rejected
The share rule that pulled every player toward zero was replaced. The seasons had been seen once before, so the evidence is labelled second look and a blind forward test decides whether it holds.
- Sep 14, 2026 · NFLEstimateNFL passing yards published as an estimate
The candidate missed its calibration bar. It replaced a worse live rule under a stated exception and stays labelled an estimate, never a validated forecast.
- Sep 14, 2026 · NFLAdoptedNFL touchdown chances rebuilt from opportunity shares
Tested blind on eight past seasons, where it beat the live rule and a rolling rate and was calibrated overall. It runs a blind forward test this season.
- Sep 14, 2026 · NFLAdoptedNFL win chance and margin heads adopted independently
Each head cleared its own bars on sixteen past seasons it was never fit on. Neither claims to out-predict the sportsbook market; the market's own number sits beside them.
- Sep 14, 2026 · MLBPausedMLB total call paused by the live-record gate
Over 576 graded games it did worse than a coin flip by a clear margin. The call is withdrawn from every page while it keeps being made and graded; it returns when the record recovers.
- Sep 13, 2026 · NFLAdoptedNFL game-total head replaced after a centring bias was found
The previous head carried a frozen constant that lifted every total; the replacement was tested on twenty-two past seasons and the sportsbook total is shown beside it.
Why a model can be paused
Health is an alarm. Only a preregistered forward test or the live-record gate changes what publishes.
Every public call is graded against official results. When a call has done worse than a coin flip by a clear margin over its graded record, the live-record gate withdraws it from every page. The model keeps making and grading that call every day, so the record can recover and the call comes back on its own. Nothing is edited by hand, and grading never reads the paused page.
The words we use
Lower is better for a miss or a loss score; about 8 in 10 is the target for a range, not a score to beat.
- Tested on past seasons
- Before it went live, this model was scored on past seasons it had never seen and beat the simpler rules it replaced. That is a bar, not a promise.
- Forward test running
- The model is being graded on new games as they happen, against the model it replaced. If it falls significantly behind, the previous model comes back automatically.
- Holding
- Over the games graded so far, this call has done at least as well as its floor (a coin flip, or the model it replaced).
- Watch
- Over the games graded so far, this call is slightly behind its floor, but not by enough to be sure it is worse. It keeps publishing while the record grows.
- Paused
- Over the games graded so far, this call has done worse than a coin flip by a clear margin. The public call is withdrawn; it is still made and graded every day and comes back when its record recovers.
- Too early to judge
- Not enough games have been graded to judge this call either way. The number is published; the verdict on it is not.
- Experimental
- Published and labelled experimental: tested on past seasons, but not validated to out-predict the sportsbook market.
- Research only
- A research candidate scored privately beside the live model. It never powers a public number.
A range that always hits could be made by widening it; progress is a narrower range that still lands about 8 in 10. A miss or a loss score is lower when better. Calibration means a call shown at 60% lands about 6 in 10 times. Every figure on this site is from a receipt or a graded ledger; the methodology explains how each is built.
