Backtest overfitting: how to tell an edge from a lucky curve

Why so many backtests look better than they are: overfitting, data snooping, look-ahead and the candle that touches both levels, and the checks that catch each one.

Written by
The EdgeLuma team
Published
Reading time
7 min
On this page
  1. What is backtest overfitting?
  2. Why does trying many versions inflate the result?
  3. What is look-ahead bias?
  4. The candle that touches both the stop and the target
  5. How many trades does a result need?
  6. How to check that an edge is real
  7. Frequently asked questions
  8. Sources

A backtest that looks too good usually is. Not because anyone cheated, but because building a strategy on past data has a built-in pull toward the flattering answer: every rule you add and every parameter you adjust is chosen while looking at the result. This guide covers the ways a backtest flatters an idea, and the checks that catch each one.

What is backtest overfitting?

Backtest overfitting is what happens when a strategy’s rules are tuned until they fit the historical data so closely that they describe its noise rather than a behaviour of the market that repeats. An overfit strategy can look excellent on the data it was tuned on and fail on any other.

It rarely takes bad faith. Try twenty versions of a moving-average crossover, keep the best, and the best will look good partly because it was the luckiest of twenty. Researchers have shown that high simulated performance is easy to reach after trying a fairly small number of configurations of a strategy, and they proposed a way to estimate the probability that a backtest is overfit from the results of all the versions that were tried.12

The question is not whether the best version looks good. It is how good the best of that many versions would look by chance.

Why does trying many versions inflate the result?

Each version you test is another draw. With enough draws, one of them will look good by chance, and the one you keep is, by construction, the best draw. Statisticians call this data snooping: the same data is used again and again to choose between models, and the winner’s result carries a bias that comes from the choosing itself.3

The practical consequence is a discount. The more variations were tried before a result was found, the more of that result should be assumed to be luck; researchers have proposed haircuts on a backtest’s performance that grow with the number of tests behind it.4

What you can do about it:

  • Count your trials. Write down every version you run, not only the one you keep.
  • Prefer rules with few moving parts. Every free parameter is one more way to fit the noise.
  • Distrust a sharp peak. If a stop at 1.8 times the average range works and 1.7 or 1.9 do not, the 1.8 is fitted. A real effect usually survives a small change.

What is look-ahead bias?

Look-ahead bias is the use of information that did not exist yet at the moment a decision is simulated. It is the most damaging error a backtest can make, because it can turn any idea into a winner, and the hardest to see, because nothing in the result looks wrong.

Its common forms:

  • Deciding on a candle that has not closed. A signal computed on a candle’s close cannot be traded at that same candle’s open.
  • Levels confirmed after the fact. A swing high is only known once the candles after it have formed; drawn at its own candle, it lets the backtest act on it before anyone could have.
  • Knowing the order of events inside a candle. A candle records its open, high, low and close, not which of them came first.

The last one deserves its own section, because every backtest built on candles meets it.

The candle that touches both the stop and the target

When a single candle’s range covers both the stop and the target, the candle alone cannot say which of the two price reached first. A backtest that assumes the target came first books a win that may have been a loss, and over many trades that assumption adds up to an edge that never existed.

There are two honest answers. The first is to look inside the candle: read the same stretch of time on finer candles, one minute at a time, and see which level was crossed first. The second, for the cases the finer candles still cannot settle, is to assume the worst and book the stop.

One hourly candle touches both the stop and the target. Read on one-minute candles, the same hour shows the stop was reached first, so the trade is booked as a loss.TargetEntryStop?One 1h candleboth levels inside itthe stop came firstthe target, laterThe same hour, minute by minutebooked as a stop: a loss of 1R, not a win
FigureOne candle touches both levels. Inside it, the one-minute candles show which came first; when even they cannot tell, the stop is booked.

EdgeLuma does both. Every run reads its fills, stops and targets on the finest candles it holds inside each candle, one minute wherever it has them, and when a candle still touches both levels, the engine books the stop.

How many trades does a result need?

A result over a handful of trades is mostly noise. The expectancy, the average result of one trade, needs a sample before it means anything: over ten or twenty trades it is luck rather than an edge, and around a hundred trades it starts to say something. A long backtest that produces few trades is still a small sample, however many years it covers.

The same is true of every other number. A win rate measured on fifteen trades can change completely with the next five, and a maximum drawdown measured on a short history is the worst thing that happened, not the worst thing that can.

How to check that an edge is real

No single test proves an edge. Together, these make a lucky curve much harder to mistake for one:

  1. Hold data back. Split the history: tune on the first part, then run the final rules once, unchanged, on the second. If the result collapses on data the rules never saw, it was fitted.
  2. Test across conditions. A rising market, a falling one, a flat one. An edge that only exists in one of them is a bet on that condition.
  3. Move the inputs a little. Shift each parameter a step either way. A real effect weakens gently; a fitted one falls off a cliff.
  4. Count the costs. A thin edge before costs is often no edge after them, as our guide to fees and slippage shows.
  5. Read the trades. If one or two trades make the whole result, the result is those trades.
One price history split in two: the rules are tuned on the first part, then run once, unchanged, on the part held back.Tuned herein-sample: the rules are fitted on this partChecked here, onceout-of-sample: never seen
FigureTune where you look, check where you did not: the rules are fitted on the first part of the history, then run once, unchanged, on the part held back.

Frequently asked questions

What is the difference between overfitting and data snooping?

Overfitting describes the result: rules that fit the noise of the data they were tuned on. Data snooping describes the process that usually leads there: using the same data again and again to choose between many versions, so that the winner’s result is inflated by the choice itself.

How do I know if my backtest is overfit?

Look for the signs: many versions tried before this one, many free parameters, a result that changes sharply when a parameter moves a little, and performance that collapses on data the rules were not tuned on. The more of them apply, the more of the result is likely luck.

What is out-of-sample testing?

Running a strategy, unchanged, on data it was not built or tuned on. It is the closest a backtest gets to the future: if the rules only work on the data they were fitted to, the out-of-sample run usually shows it.

Is a candle that touches both the stop and the target a problem?

Yes, because the candle does not record which level was reached first. Reading the same stretch on finer candles settles most cases; for the rest, booking the stop keeps the result from counting wins that may have been losses.

Can a strategy with a good backtest still lose money?

Yes. A backtest replays the past with the rules and the costs you chose, and markets change: an edge that was real can fade. A careful backtest lowers the odds of trading an idea that never worked. It cannot promise the future.

Sources

  1. Pseudo-Mathematics and Financial Charlatanism: The Effects of Backtest Overfitting on Out-of-Sample PerformanceDavid H. Bailey, Jonathan M. Borwein, Marcos López de Prado and Qiji Jim Zhu, Notices of the American Mathematical Society, 2014
  2. The Probability of Backtest OverfittingDavid H. Bailey, Jonathan M. Borwein, Marcos López de Prado and Qiji Jim Zhu, Journal of Computational Finance
  3. A Reality Check for Data SnoopingHalbert White, Econometrica, 2000
  4. BacktestingCampbell R. Harvey and Yan Liu, The Journal of Portfolio Management, 2015

Written by the EdgeLuma teamEdgeLuma is a backtesting platform for crypto traders. Describe a trading strategy in plain words: EdgeLuma turns it into rules, replays them on years of real candles with the exchange's real fees and slippage, and shows every trade and every number, even when the strategy loses.

Every article

Keep reading

What took years of chart-watching now takes seconds.