All articles
Backtesting

How Many Years of Data Should a Backtest Cover?

More is not simply better. What a given span of history can and cannot tell you, and why what the period contains matters more than how long it is.

Arthalab6 min read
Two years is a floor, three to five is comfortable, and what the period contains matters more than how long it is. A decade of quiet markets tests less than eighteen months containing a crash.

Why more is not simply better

The intuition is that longer means more reliable. It is partly true and has two real limits.
  1. Market structure changes. Weekly expiry mechanics, lot sizes and participation have all shifted. Data from a very different structural regime describes a market that no longer exists.
  2. More data means more room to over-fit. A longer history gives you more parameters to tune against, and the result looks more impressive while meaning less.

What matters more than length

The period should containBecause it tests
A sharp crashGap risk, and whether your stop actually protects you
A volatility spikePremium selling under stress
A sustained one-way trendRange-dependent strategies
A quiet, rangebound stretchWhether cost drag eats a thin edge
Several expiry cyclesTheta and gamma behaviour near expiry
Check your test window against this list. If a row is missing, your backtest has not tested that scenario, regardless of how many years it spans.

Minimum trade count

Years are a proxy. What actually governs reliability is how many trades the test produced.
Trades in the testWhat it supports
Under 30An impression, not a conclusion
30 to 100A rough sense of direction
100 to 300Reasonable confidence in the average
300+Useful confidence, including in the tails
This is why a daily intraday strategy needs less calendar time than a weekly one. Fifty trades is fifty trades whether they took two months or two years.

Working backwards

A strategy entering once a week needs several years to reach 150 trades. One entering every day reaches it in under a year. Plan the span around the trade count you need.

Splitting what you have

However much data you have, do not use all of it for building.
1

Hold back the most recent 20 to 30 per cent

Do not look at it while designing.
2

Build and tune on the rest

This is your in-sample period.
3

Test once on the held-back period

Once. Looking repeatedly turns it into in-sample data.
4

Compare the two

A large gap means over-fitting, and the out-of-sample number is the honest one.
5

Then paper trade forward

Unseen real-time data is the strongest test available before capital.

What a short test cannot see

Concretely, the things that only appear over a longer window. Worth knowing so you can judge what your test has and has not covered.
  • A volatility regime change. A strategy tuned to a calm market behaves differently once the baseline shifts, and that shift takes months to appear.
  • Expiry-cycle effects. Theta and gamma near expiry behave distinctly, and you want many cycles, not a few.
  • Seasonal patterns, if they exist at all for your strategy. One year gives you one observation of each.
  • The tail. The worst day in your sample is not the worst day that can happen — it is just the worst that did, within the window you chose.
That last point is the one that catches people with impressive short-window results. A maximum drawdown figure is a statement about your sample, not a limit the strategy respects.

Data quality matters as much as quantity

Five years of poor data is worse than two years of good data, because it produces the same confidence on a weaker foundation.
  • Is the data at a granularity that matches your entry rule?
  • Does it include the actual option chain, or an approximation?
  • Are lot size changes handled correctly across the period?
  • Are expiry schedule changes reflected?
  • Are costs applied, or do you need to add them yourself?
The mistakes list covers what happens when the answer to the last one is no. It is the most common reason a backtest looks far better than the live result.

Practical answers by strategy type

Strategy typeSuggested span
Daily intraday index options2 to 3 years
Expiry-day only3 years or more — fewer trades per year
Weekly entry4 to 5 years
Event-drivenAs many occurrences as you can find, span aside
Running the backtest is the easy part. Choosing the window honestly is where the judgement sits.

The short version

  • Two years is a floor, three to five is comfortable
  • What the period contains matters more than how long it is
  • Trade count governs reliability — aim for 100 or more, 300 is better
  • Hold back the most recent 20 to 30 per cent and test on it once
  • More data also means more room to over-fit, so watch parameter count

Frequently asked questions

Two years is a floor and three to five is comfortable. For a daily intraday index options strategy, two to three years is usually enough.

No. Market structure changes, so very old data describes a market that no longer exists, and a longer history also gives you more room to over-fit.

What the period contains. It should include a sharp crash, a volatility spike, a sustained trend and a quiet range. A decade of calm tests less than eighteen months with a crisis in it.

Under 30 is an impression. 100 to 300 gives reasonable confidence in the average. Above 300 you start to learn something about the tails too.

No. Hold back the most recent 20 to 30 per cent, build on the rest, then test on the held-back period once. Checking it repeatedly turns it into in-sample data.

Yes, because it produces fewer trades per year. Three years or more, driven by reaching a useful trade count rather than by the calendar.

Start with a free 3-day trial

Build a strategy, backtest it and run it on paper — no broker, no IP and no money needed to try it.

Ask us on Telegram
How Many Years of Data Should a Backtest Cover? | Arthalab — Algo Trading India