Skip to main content
EducationalNot betting advice · research use only.
GameTime Picks
Model Lab

What is live, what is being tested, what is paused

Every model here was scored on past games it had never seen before it went live, is graded on every new game after, and can be paused or replaced by its own record. This page reads those receipts; it does not restate them.

Models currently live

Each sport's model families and where they stand right now.

Model status · NFL

What the statuses mean →
  • Game forecastsExperimental
    Experimental · tested on past seasons

    We publish a simulated score range and win chance for every game, clearly marked experimental (1 games forecast today under nfl-regular-season-public-v1). The win chance comes from a team rating that weights margin of victory. Tested on 4,281 past games it had never seen (2006–2021), it scored 0.629 on log loss (lower is better) against 0.642 for our previous rating, 0.693 for a coin flip and 0.610 for the sportsbooks' own odds — so it makes no claim to out-predict the sportsbook market. Every forecast is frozen before kickoff and settled against the official result.

    What does this mean?

    Published and labelled experimental: tested on past seasons, but not validated to out-predict the sportsbook market.

  • Player rangesForward test running
    Blind forward test running

    Receptions, yards and touchdown chances are graded week by week against preregistered bars; a family that breaches them falls back to the previous rule automatically. 0 player-games graded so far.

    What does this mean?

    The model is being graded on new games as they happen, against the model it replaced. If it falls significantly behind, the previous model comes back automatically.

  • Touchdown chancesHolding
    Expected scorers match actual

    Over 168 graded, this call has held up against its own expected-scorer count.

    What does this mean?

    Over the games graded so far, this call has done at least as well as its floor (a coin flip, or the model it replaced).

Model status · MLB

What the statuses mean →
  • Winner callsWatch
    Watch · slightly behind

    Over 605 graded, this call is slightly behind a coin flip, not yet by enough to be sure. It has landed 50% of the time while showing about 57%.

    What does this mean?

    Over the games graded so far, this call is slightly behind its floor, but not by enough to be sure it is worse. It keeps publishing while the record grows.

  • Game totalsPaused
    Paused · below a coin flip

    Over 576 graded, this call has done worse than a coin flip by a clear margin. It has landed 49% of the time while showing about 59%.

    What does this mean?

    Over the games graded so far, this call has done worse than a coin flip by a clear margin. The public call is withdrawn; it is still made and graded every day and comes back when its record recovers.

  • Run line callsHolding
    Holding its record

    Over 605 graded, this call has held up against a coin flip. It has landed 65% of the time while showing about 60%.

    What does this mean?

    Over the games graded so far, this call has done at least as well as its floor (a coin flip, or the model it replaced).

Model status · Premier League

What the statuses mean →
  • Match modelTested on past seasons
    Tested blind on nine past seasons

    Scored on 3,420 Premier League matches from seasons it was never fit on, where it beat the previous model and a plain rating system. Its totals follow the league's recent scoring rate, so over-2.5 is the same for every match.

    What does this mean?

    Before it went live, this model was scored on past seasons it had never seen and beat the simpler rules it replaced. That is a bar, not a promise.

  • Forward testForward test running
    Forward test · 0 of 60 matches

    Every graded match scores the live model beside the one it replaced. A verdict needs 60 matches; 0 are in.

    What does this mean?

    The model is being graded on new games as they happen, against the model it replaced. If it falls significantly behind, the previous model comes back automatically.

  • Live recordToo early to judge
    Too early to judge

    35 graded so far — not yet enough to judge against even odds.

    What does this mean?

    Not enough games have been graded to judge this call either way. The number is published; the verdict on it is not.

Model status · UFC

What the statuses mean →
  • Fight modelTested on past seasons
    Backtest passed on held-out fights

    Winner, method and round heads each beat their baseline on fights held out of the fit.

    What does this mean?

    Before it went live, this model was scored on past seasons it had never seen and beat the simpler rules it replaced. That is a bar, not a promise.

  • Live recordToo early to judge
    Too early to judge

    31 graded so far — not yet enough to judge against a coin flip.

    What does this mean?

    Not enough games have been graded to judge this call either way. The number is published; the verdict on it is not.

Experiments and shadows

Blind forward tests grade a live model week by week; a shadow is scored privately and powers nothing public.

  • NFL receptions · blind forward testForward test running
    Accumulating · 0 of 300 player-games

    Every week's forecast is committed before its first kickoff and graded against the official box scores; a family that breaches its preregistered bars falls back to the previous rule automatically.

  • NFL receiving yards · blind forward testForward test running
    Accumulating · 0 of 300 player-games

    Every week's forecast is committed before its first kickoff and graded against the official box scores; a family that breaches its preregistered bars falls back to the previous rule automatically.

  • NFL rushing yards · blind forward testForward test running
    Accumulating · 0 of 300 player-games

    Every week's forecast is committed before its first kickoff and graded against the official box scores; a family that breaches its preregistered bars falls back to the previous rule automatically.

  • NFL passing yards · blind forward testForward test running
    Accumulating · 0 of 300 player-games

    Every week's forecast is committed before its first kickoff and graded against the official box scores; a family that breaches its preregistered bars falls back to the previous rule automatically.

  • Premier League match model · blind forward testForward test running
    Accumulating · 0 of 60 matches

    The live match model is scored beside the model it replaced on every graded match. If it is significantly worse, the previous model publishes again from the next build.

  • Premier League club-specific totals · private shadowResearch only
    Research only · registered, first fixtures pending

    A club-by-club total is scored privately beside the live model's constant total on every graded fixture. It powers no public number; adoption would be a separate decision.

Recent decisions

What the receipts decided, most recent first. A rejected candidate stays rejected; it is never retried by tweaking it.

  1. Sep 15, 2026 · EPLRejected
    Club-specific match totals: rejected on the blind test, positive on a second look

    Both candidates beat the constant total on every totals measure across four other leagues and the Premier League, but missed a preregistered result-calibration ceiling on the blind set. The verdict stands; the better candidate now runs as a private shadow.

  2. Sep 15, 2026 · EPLResearch only
    Premier League club-specific totals registered as a private shadow

    Not adopted. Scored beside the live model on future fixtures for a later decision.

  3. Sep 14, 2026 · EPLAdopted
    Premier League match model replaced by a rating model tested blind on nine past seasons

    It beat the previous model and a plain rating system on 3,420 matches it was never fit on, and now runs a blind forward test against the model it replaced.

  4. Sep 14, 2026 · NFLSecond look
    NFL player ranges: receptions, receiving yards, rushing yards eligible on a second look; passing yards rejected

    The share rule that pulled every player toward zero was replaced. The seasons had been seen once before, so the evidence is labelled second look and a blind forward test decides whether it holds.

  5. Sep 14, 2026 · NFLEstimate
    NFL passing yards published as an estimate

    The candidate missed its calibration bar. It replaced a worse live rule under a stated exception and stays labelled an estimate, never a validated forecast.

  6. Sep 14, 2026 · NFLAdopted
    NFL touchdown chances rebuilt from opportunity shares

    Tested blind on eight past seasons, where it beat the live rule and a rolling rate and was calibrated overall. It runs a blind forward test this season.

  7. Sep 14, 2026 · NFLAdopted
    NFL win chance and margin heads adopted independently

    Each head cleared its own bars on sixteen past seasons it was never fit on. Neither claims to out-predict the sportsbook market; the market's own number sits beside them.

  8. Sep 14, 2026 · MLBPaused
    MLB total call paused by the live-record gate

    Over 576 graded games it did worse than a coin flip by a clear margin. The call is withdrawn from every page while it keeps being made and graded; it returns when the record recovers.

  9. Sep 13, 2026 · NFLAdopted
    NFL game-total head replaced after a centring bias was found

    The previous head carried a frozen constant that lifted every total; the replacement was tested on twenty-two past seasons and the sportsbook total is shown beside it.

Why a model can be paused

Health is an alarm. Only a preregistered forward test or the live-record gate changes what publishes.

Every public call is graded against official results. When a call has done worse than a coin flip by a clear margin over its graded record, the live-record gate withdraws it from every page. The model keeps making and grading that call every day, so the record can recover and the call comes back on its own. Nothing is edited by hand, and grading never reads the paused page.

The words we use

Lower is better for a miss or a loss score; about 8 in 10 is the target for a range, not a score to beat.

Tested on past seasons
Before it went live, this model was scored on past seasons it had never seen and beat the simpler rules it replaced. That is a bar, not a promise.
Forward test running
The model is being graded on new games as they happen, against the model it replaced. If it falls significantly behind, the previous model comes back automatically.
Holding
Over the games graded so far, this call has done at least as well as its floor (a coin flip, or the model it replaced).
Watch
Over the games graded so far, this call is slightly behind its floor, but not by enough to be sure it is worse. It keeps publishing while the record grows.
Paused
Over the games graded so far, this call has done worse than a coin flip by a clear margin. The public call is withdrawn; it is still made and graded every day and comes back when its record recovers.
Too early to judge
Not enough games have been graded to judge this call either way. The number is published; the verdict on it is not.
Experimental
Published and labelled experimental: tested on past seasons, but not validated to out-predict the sportsbook market.
Research only
A research candidate scored privately beside the live model. It never powers a public number.

A range that always hits could be made by widening it; progress is a narrower range that still lands about 8 in 10. A miss or a loss score is lower when better. Calibration means a call shown at 60% lands about 6 in 10 times. Every figure on this site is from a receipt or a graded ledger; the methodology explains how each is built.