Author: Vladislav Musilek   

Backtesting vs Overfitting: How to Tell a Real Model From a Curve-Fitted One

Give a competent analyst historical odds and results, and they can produce a backtest showing almost any return you like. This is not fraud. It is the natural consequence of searching a large space of possible rules until one of them fits the data you already have. The resulting model describes the past beautifully and predicts the future not at all.

Since every prediction service that has ever advertised has shown you a backtest, it is worth understanding how to interrogate one.

What overfitting actually is

A model has two kinds of pattern available to it: signal, which is a real relationship that will persist, and noise, which is coincidence specific to the sample. Overfitting is learning the second kind.

The mechanism is simple. Add enough parameters, and a model can memorise its training data exactly. Test enough candidate rules, and some will look profitable purely because of how the sample happened to fall. If you evaluate a thousand random strategies on historical baseball data, a handful will show excellent returns. They are not strategies. They are the tail of a distribution.

The tell is always the same: spectacular historical performance that degrades sharply the moment the model meets data it has not seen.

The questions that actually discriminate

Was there a genuine out-of-sample holdout?

The minimum standard is that a portion of the data — typically the most recent seasons — was set aside before any development began and never touched until final evaluation. If the model was tuned, adjusted, and then re-tested on the same holdout, the holdout is contaminated and the result is a training score wearing a disguise.

For time series data, the split must respect chronology. Randomly shuffling games across seasons leaks future information into the training set and produces flattering nonsense.

Did the backtest use the odds that were actually available?

This is where most backtests quietly fall apart. Using closing odds in a backtest while betting at opening odds in reality is a mismatch that can invent an edge from nothing. So can assuming you obtained the best price across all books on every bet, or ignoring that limits at the best price are often small.

The correct question: at the moment this bet would have been placed, was this price genuinely obtainable in the size assumed?

Is there lookahead bias in the features?

Lookahead bias means the model used information that did not exist at prediction time. It is usually accidental and often subtle: a season-long statistic applied to games from earlier in that season, an injury designation recorded after the fact, a lineup that was not confirmed until after first pitch, a park factor calculated using the games being predicted.

Every feature should be answerable to one question: was this value knowable, in this form, before the event started?

How many strategies were tested before this one was chosen?

This is the question nobody volunteers an answer to, and it is the most important one. A backtest showing a 6% return means something quite different if it was the first thing tried versus the best of four hundred variants. Multiple testing inflates apparent performance mechanically, and the correction is severe.

Are the results stable across periods and segments?

A real edge tends to show up broadly, if unevenly. An overfitted one concentrates suspiciously: profitable in two of six seasons, or driven almost entirely by one market type, or dependent on a specific league in a specific year. Ask for results broken down by season, by sport, and by market. Aggregate figures conceal exactly the pattern you want to see.

Warning signs in an advertised backtest

  • Returns that are too good. Sustained returns on turnover in the high teens or above, across thousands of bets in liquid markets, are extraordinary claims. Genuine long-run edges in major markets are usually measured in low single digits.
  • Smooth equity curves. Real betting produces jagged results with long drawdowns. A gently rising line is a signature of fitted data.
  • No stated methodology. If the training period, holdout period and odds source are not specified, the number cannot be evaluated.
  • Backtest without forward test. Historical performance is a hypothesis. Live, timestamped, publicly recorded results are the test.
  • Rules that changed after poor results. A model revised each time it underperforms is being fitted continuously, in public.

Why forward testing is the only real evidence — and why it is slow

Live results cannot be fitted, because the data did not exist when the model was built. This makes forward testing the gold standard. It also makes it painfully slow, for the reasons set out in our article on sample size: distinguishing a modest edge from break-even takes on the order of a thousand bets.

This is the genuine tension in evaluating any prediction service, and there is no clever way around it. The practical resolution is to combine three imperfect signals rather than relying on any one: a methodologically sound backtest, live closing line value tracked from the start, and full disclosure of every published pick including the losers. None of the three is conclusive. Together they are considerably better than a headline return figure.

Frequently asked questions

What return should a legitimate backtest show?

Modest single-digit returns on turnover across a large sample are far more credible than double-digit figures. In liquid markets, a sustained edge of a few percent is already a strong result.

How large should the out-of-sample period be?

Enough bets to be statistically meaningful — ideally over a thousand — and covering multiple seasons so that a single unusual year cannot drive the conclusion.

Does machine learning make overfitting more or less likely?

More likely, without discipline. Flexible models with many parameters fit noise readily. The safeguards — chronological splits, regularisation, an untouched holdout, honest accounting of how many variants were tried — matter more as model capacity increases.

Can I verify a backtest I have been shown?

Rarely in full, since you lack the underlying data. You can ask about methodology, request segmented results, and compare the backtest against subsequent live performance. A service unwilling to answer methodological questions has answered the most important one.

Related reading

Ask us the awkward questions

Methodology, holdout periods, published losers. If a service will not discuss how its model was built and tested, that is your answer. We would rather have the conversation.

  • One pick a day
  • 25+ picks every 30 days
  • Full OPTIMUS II reasoning
  • Every pick published, wins and losses
Start 7 days free 7 days free on the Foundation plan · Cancel renewal anytime

18+ · Please gamble responsibly.
69 Advisory provides informational sports analysis only. Nothing above is a guarantee of results and past performance does not indicate future outcomes. Only stake what you can afford to lose.
Free confidential support: National Gambling Helpline (UK) 0808 8020 133 · National Problem Gambling Helpline (US) 1-800-GAMBLER · begambleaware.org

Other news

Yankees F5 Draw No Bet vs Blue Jays — Free MLB Analysis

MLB · Free Analysis Yankees F5 Draw No Bet vs Toronto Toronto Blue Jays @ New York Yankees · Soriano vs Rodón Analysed by OPTIMUS II The pick MLB · First 5 · Draw No Bet Yankees F5 Draw No Bet @ 1.751 1.00Unit 1.751Odds 58.3%Decided prob. +1.7%Model EV The strongest read from the supplied […]

Read article

Twins @ Padres: How We Read the Moneyline

MLB · Match Analysis How we read it: Padres moneyline vs Minnesota Minnesota Twins @ San Diego Padres · Kremer vs Mize Analysed by OPTIMUS II The read MLB · Moneyline San Diego Padres to win @ 1.751 1.00Unit 1.751Odds 59.1%Model prob. +3.5%Model EV Note: this analysis is published for the archive after the game. […]

Read article

Cardinals @ Phillies: How We Read the F5 Under (5.0)

MLB · Match Analysis How we read it: Cardinals @ Phillies F5 under 5.0 St. Louis Cardinals @ Philadelphia Phillies · Dobbins vs Luzardo Analysed by OPTIMUS II The read MLB · First 5 innings First 5 Innings – Under 5.0 Runs @ 1.649 1.00Unit 1.649Odds 57.8%Win prob. +10.7%Model EV Note: this analysis is published […]

Read article