All articles
Backtesting

Seven Backtesting Mistakes That Make Good Results Meaningless

Each of these produces a backtest that looks better than reality, and each is easy to make without noticing. How to spot them in your own results.

Arthalab6 min read
Every mistake on this list produces a backtest that looks better than reality. That is what makes them dangerous — none of them announce themselves with an error.

1. Ignoring costs

Brokerage, STT, exchange charges, GST, stamp duty. On an intraday options strategy these are not a rounding error.
A strategy making a small average profit per trade can be comfortably loss-making after costs. The per-trade cost arithmetic is worth doing once by hand so you know the number for your lot size.

2. Assuming you got the price you wanted

A backtest typically fills at a candle's close or a quoted price. Reality gives you whatever is available when your order arrives.
On illiquid strikes the gap is substantial. The spread is the cost you pay to transact, and far OTM strikes on a quiet day have wide ones.

3. Over-fitting to history

You tried 9:15, 9:20, 9:25 and 9:30. One worked best. Did you find an edge, or the quirk of a specific dataset?
Curve fitting is the single most common way a backtest misleads. The tell is sensitivity: if moving the entry by five minutes destroys the result, you fitted noise.

The sensitivity test

Test every parameter for sensitivity. A real edge degrades gracefully as you nudge parameters. A fitted one falls off a cliff.

4. Testing only a friendly period

A strategy tested across 2023 and 2024 has not seen a genuine crisis. Index options behaviour in a volatility spike is not a scaled version of normal behaviour.
Period typeWhy to include it
A sharp crashTests gap risk and whether stops actually help
A sustained trendMany premium-selling strategies struggle here
A quiet rangeTests whether cost drag eats a thin edge
An expiry week spikeTests behaviour when theta and gamma both bite

5. Reading the return before the drawdown

A strategy returning a strong annual figure with a severe peak-to-trough drawdown is not a good strategy. It is a strategy you will abandon partway down.
Read the drawdown first, then ask whether you would have kept going through it with real money. If the answer is no, the return figure is irrelevant.

6. Changing the strategy and keeping the old evidence

You backtested, paper traded, saw something you did not like, adjusted a parameter. The paper test now describes a different strategy.
Each change resets your evidence back to the stage where you made it. This feels pedantic until you notice how many iterations you did without ever testing on unseen data.

7. Trusting a single backtest run

One run on one period with one parameter set is a single data point dressed up as a conclusion.
  • Run it on at least three distinct periods separately
  • Run it at a few parameter values either side of your choice
  • Run it on both NIFTY and SENSEX if the logic should generalise
  • Hold back recent months and test on them last
  • Compare the worst run, not the best, when deciding
Judging by the worst run is the discipline that separates a plan from a hope. The best run is the one you will remember and the worst is the one you will experience.

A quick self-audit

1

Are costs in the numbers?

If unsure, subtract them manually and re-read.
2

What is the maximum drawdown?

Would you have continued through it?
3

How sensitive is it to parameters?

Nudge each one and watch.
4

Does the period include a volatile stretch?

If not, the test is incomplete.
5

Was any part of this tested on unseen data?

If no, that is the next step, not live trading.

The short version

  • Costs ignored is the most common and most fatal omission
  • Over-fitting shows up as parameter sensitivity — test for it directly
  • A period without a volatility spike is an incomplete test
  • Read drawdown before return, every time
  • Judge by the worst run, because that is the one you will live through

Frequently asked questions

Leaving out costs. Brokerage, STT, exchange charges, GST and stamp duty are not a rounding error on an intraday options strategy, and a thin edge disappears once they are included.

Test parameter sensitivity. A real edge degrades gracefully as you nudge entry times or thresholds. A fitted one collapses, which tells you it was describing one dataset's quirks.

Because index options behave differently in a volatility spike, and that behaviour is not a scaled version of normal. A period without one has not tested the scenario that hurts.

Drawdown. A strong return with a severe drawdown is a strategy you will abandon partway down, which makes the return figure academic.

No. Run it on several distinct periods, at a few parameter values, and judge by the worst result rather than the best.

It resets your evidence to the stage where you made the change. If you adjusted after seeing paper results, you need a fresh forward test on unseen data.

Not on its own. A strategy winning most days while losing far more on the days it loses can still be net negative. Read the average loss against the average win, not the win rate alone.

No. It can show the logic was not obviously broken on past data, which is worth knowing and is weaker evidence than it feels. Forward testing on unseen data is what moves you beyond that.

Start with a free 3-day trial

Build a strategy, backtest it and run it on paper — no broker, no IP and no money needed to try it.

Ask us on Telegram
Seven Backtesting Mistakes That Make Good Results Meaningless | Arthalab — Algo Trading India