All articles
Backtesting

Why Backtest Results Differ From Live Trading

Slippage, assumed fills, costs, survivorship in your own testing and regime change — the structural reasons a backtest overstates live performance, and how to narrow the gap.

Arthalab10 min read
A backtest almost always looks better than live trading, and the gap is structural rather than a bug. Understanding where it comes from lets you discount a backtest by roughly the right amount instead of being surprised later.

Five sources of the gap

1. Fills are assumed

A backtest fills your order at a historical price. It does not queue behind other orders, does not cross a spread that widened at that instant, and never partially fills. Real execution does all three.
The effect is small on liquid at-the-money strikes and large everywhere else — far strikes, the first minute of the session, and any moment volatility spikes. Which is unfortunate, because those are exactly the moments many strategies are designed to act on.

2. Costs are easy to underestimate

Brokerage, exchange charges, STT, stamp duty and GST all apply per leg. A four-leg strategy pays them four times on entry and four times on exit. A strategy with a thin per-trade edge disappears into that, and the thinner the edge the more it matters.

3. The backtest never runs out of margin

Historical simulation does not check your balance. Live, an option-selling position can be rejected for margin exactly when volatility — and therefore the opportunity — is highest. Margin rejections are the most common live failure, and they cluster on the days that matter most.

4. You tested more than you remember

If you ran twenty variations and kept the best one, the winner's result includes a selection effect. The strategy was not found, it was selected out of noise — and the apparent edge partly belongs to the search, not the market.

5. Regimes change

A strategy built on a particular volatility environment, expiry structure or lot size is describing conditions that can and do change. Historical performance under conditions that no longer exist is not evidence about tomorrow, however long the sample.

Which gaps you can close and which you cannot

Source of the gapCan you reduce it?How
SlippagePartlyTrade liquid strikes, smaller size, avoid the first minutes
CostsPartlyFewer legs, fewer re-entries, a broker with lower per-order charges
Margin rejectionsYesKeep headroom above the peak requirement, not the average
Selection effectYesFix parameters before testing and write them down
Regime changeNoNothing. Accept it and re-test periodically
Your own interventionYesDefine override rules in advance, or none at all
Four of those six are within your control, which is more encouraging than the usual framing suggests. The two that are not — regime change and the irreducible part of slippage — are also the two that are easiest to recognise when they happen.

A sixth reason nobody mentions

You behave differently with real money. A backtest assumes the strategy ran untouched for the whole period. Live, you will be tempted to intervene — to skip a day that feels wrong, to exit early on a position that is uncomfortable, to size down after a loss.
Every one of those interventions makes your live results something other than the strategy you tested. The backtest is a measurement of the rules; your live P&L is a measurement of the rules plus your discretion.

Narrowing the gap

  • Include realistic costs in the comparison rather than reading gross profit
  • Prefer liquid strikes — the slippage assumption holds best there
  • Fix parameters before testing, so you are not selecting on noise
  • Run the fragility check: nudge a parameter and confirm the result degrades gently
  • Paper trade before funding, then go live at minimum size
  • Compare the live period against the backtest for the same dates
  • Write down, in advance, what would make you override the strategy
The sixth step is the only way to measure your own execution cost rather than guessing at it. The difference between paper and live over the same days is a real number specific to your strategy, size and broker — and it is the number you should use to discount future backtests.

Slippage, quantified for yourself

Rather than accepting slippage as an unknown, it can be measured directly once you are live, and the measurement takes a few minutes a week.
1

For each entry, note the quoted price at your entry time

The logs record what the engine saw when it constructed the order.
2

Note the price you actually filled at

The order book has this.
3

Record the difference per leg

In points, not rupees, so it is comparable across lot-size changes.
4

Average it over twenty or so trades

One trade tells you nothing; twenty gives you a usable figure.
5

Multiply by legs and by two

Entry and exit. That total is what each round trip costs you before the strategy has done anything.
Once you have that number, every future backtest becomes more useful: you can subtract a realistic execution cost rather than hoping the edge is large enough to absorb an unknown one.

What a normal gap looks like

A strategy whose live results land somewhat below the backtest, in the same direction and with a similar shape, is behaving normally. That is the expected outcome and not a reason to change anything.
A strategy whose live results are the opposite sign, or whose drawdown is far deeper than tested, is telling you something different: that the backtest was not measuring what you thought it was.
What you see liveLikely explanationWhat to do
Slightly worse than backtestNormal — slippage and costsNothing. This is expected
Much worse, same shapeSlippage larger than assumed, or size too bigReduce size, prefer liquid strikes
Opposite signCurve fitting, or regime changeRe-run the fragility check
Deeper drawdown than testedSample did not contain this conditionExtend the test period and re-read
Fewer trades than expectedRejections, or the bot was not started<a href="/blog/deploy-vs-start-whats-the-difference">Check the logs first</a>
That last row is more common than any of the others, and it is not a strategy problem at all. The execution logs will tell you whether the strategy decided not to act, or never got the chance to.

Measuring your own gap

Rather than guessing how much to discount a backtest, you can measure it — and the measurement is specific to your strategy, your size and your broker, which is what makes it useful.
1

Run the strategy live at minimum size for a month

One lot. The purpose is information, not profit.
2

Run the same strategy on paper over the same period

Both deployments count towards your three-strategy limit, so plan for that.
3

Backtest the same date range after the fact

Now you have three numbers for the same days.
4

Compare paper against backtest

A gap here means the backtest assumptions were off, not that execution was poor.
5

Compare live against paper

This gap is your real execution cost — slippage, rejections and timing.
The two gaps have different causes and different fixes. Backtest-to-paper points at the test setup; paper-to-live points at execution. Collapsing them into one number tells you something is wrong without telling you what.
ComparisonWhat a gap meansWhat to change
Backtest vs paperTest assumptions do not match live conditionsReview the test period and costs
Paper vs liveExecution cost — slippage, rejectionsPrefer liquid strikes, reduce size
Both largeLikely curve fittingRe-run the fragility check

Deciding whether to stop

Some gap is expected, so a bad week is not a signal. The useful discipline is to decide the stop condition before you start, in rupees, and then follow it.
If you did the work in reading the report properly, you already know the max drawdown you accepted. That number is your stop condition. Reaching it means the strategy did what it was always capable of doing — not that something went wrong.

The short version

Some gap is structural and expected. Knowing which part you can close is what makes the difference actionable.
  • Slippage and costs you can reduce, by trading liquid strikes at sensible size
  • Margin rejections you can prevent, by keeping headroom above the peak requirement
  • Selection effect you can avoid, by fixing parameters before testing
  • Regime change you cannot do anything about except re-test periodically
  • Your own intervention you can control, by deciding override rules in advance
Measure your own gap rather than guessing at it: run live at minimum size alongside paper for the same days, and the difference is a number specific to you.

Frequently asked questions

Compare the quoted price the engine saw at entry against your actual fill, per leg, over about twenty trades. Average it, multiply by the number of legs and by two for entry and exit.

Generally yes, because larger orders interact with the book more. A strategy that fills cleanly at one lot can fill noticeably worse at ten on the same strike.

Once you have measured your own figure, you can subtract it when reading future reports. That is more reliable than a generic assumption, because it reflects your strategy, size and broker.

Most often slippage and costs, then margin constraints. If the gap is very large, the likeliest cause is that the strategy was fitted to the test period.

No. Liquid at-the-money strikes slip little. Far strikes, large size and volatile moments slip much more.

Closely, sometimes, on liquid instruments at small size. Exactly, no — assumed fills and real fills are different things.

Change one parameter slightly and re-run. A robust strategy degrades gently. A fitted one collapses.

Not automatically — some gap is expected. Decide the stop condition in advance, in rupees, and follow it rather than reacting to the first bad stretch.

Yes, completely. Every manual override makes your live results a different strategy from the one you tested. If you intend to override, define the override rules and test those too.

Check the logs before assuming anything about the strategy. The usual causes are rejected orders or a bot that was deployed but never started, neither of which is a strategy problem.

Start with a free 3-day trial

Build a strategy, backtest it and run it on paper — no broker, no IP and no money needed to try it.

Ask us on Telegram
Why Backtest Results Differ From Live Trading | Arthalab — Algo Trading India