All articles
Platform Guides

Best Options Backtesting Software in India: What to Look For

How to evaluate options backtesting tools for NIFTY and SENSEX — data quality, which metrics a report must include, cost models, and the features that separate a useful backtest from a flattering one.

Arthalab11 min read
The quality of a backtesting tool is decided by what it refuses to hide from you. Any tool can produce an equity curve. The ones worth using surface the drawdown, the trade count and the costs prominently enough that you cannot avoid reading them.
This guide covers what to evaluate, which report fields are non-negotiable, how cost models differ, and the specific ways a backtesting tool can mislead you without being wrong.

What a backtest actually does

Understanding the mechanism makes the evaluation criteria obvious. For each historical day in your range, the engine walks through the session: at your entry time it reads the option chain as it stood, applies your strike selection rule, and records an entry. Through the day it tracks each leg against your stop loss and target rules. At your exit time it closes whatever is open.

The assumptions every engine makes

Every backtest makes these assumptions:
  • Every order fills. There is no queue and no rejection.
  • It fills at the recorded price. No spread crossed, no partial fill.
  • Margin is unlimited. The position is never refused for funds.
  • The system was running. No missed logins, no outages.
None of those are unreasonable for a simulation, and none are true live. That gap is structural, which is why a tool that helps you see it is more valuable than one that produces prettier curves.

Data quality, which decides everything downstream

A backtest is only as good as the historical option chain behind it. Three things matter and they are worth asking about directly.
What to askWhy it matters
What granularity is the data?Minute-level data captures intraday stop losses that end-of-day data cannot
Does it include the full option chain?Strike selection rules need every strike, not just the at-the-money one
How are expiry and lot-size changes handled?A test spanning a revision must use the figures in force on each date
That third row catches more tools than people realise. If a backtest applies today's lot size to a historical period that used a different one, the rupee figures are wrong — and those figures are exactly what you use to decide whether the drawdown is tolerable.

Report fields that are non-negotiable

A report missing any of these cannot support a decision, however good it looks.
FieldWhy it is essential
Max drawdownThe worst stretch you would have had to sit through
Drawdown durationHow long it lasted — this is what makes people quit
Number of tradesGoverns how much every other number is worth
Max consecutive lossesThe psychological load the strategy imposes
Average win and average lossWin rate alone is meaningless without these
ExpectancyThe closest thing to a single honest summary
Reading a report from the risk end covers the order to work through these in, which matters as much as having them.

Cost modelling, where tools differ most

This is the single biggest source of difference between tools, and the one most likely to flatter a strategy.

How tools handle costs

Three approaches, in increasing order of honesty:
  1. Gross only. No costs applied. Every strategy looks better than it is, and multi-leg strategies look dramatically better.
  2. A flat per-trade figure. Better, but it understates multi-leg structures, where charges apply per leg on both entry and exit.
  3. Per-leg charges including taxes. The only approach that reflects what a four-leg strategy actually pays.
A four-leg condor pays brokerage, exchange charges, STT, stamp duty and GST on eight legs per round trip. A tool modelling that as one flat fee is understating the cost by a large multiple, and a thin edge disappears entirely once it is applied properly.

Features that prevent self-deception

The most valuable features in a backtesting tool are the ones that make it harder to fool yourself.
  • Easy re-runs at varied parameters. The fragility check — nudge one value and confirm the result degrades gently — is the single best test for curve fitting, and it is only practical if re-running is cheap.
  • Trade-by-trade output. Aggregate numbers hide a strategy whose profit comes from two outlier days.
  • Portfolio or bucket testing. Testing several strategies together shows correlation that individual tests cannot.
  • Date-range flexibility. Running two non-overlapping periods and comparing them is the closest thing to out-of-sample validation most retail tools offer.

Ways a backtest can be wrong without being buggy

Tools rarely produce arithmetically incorrect results. They produce correct arithmetic on assumptions that do not hold, which is harder to spot.
IssueWhat it looks likeHow to detect it
Look-aheadSuspiciously good entriesCheck the rule only uses data available at that instant
Survivorship in your searchThe best of many variantsCount how many versions you tried before this one
Stale reference dataRupee figures that feel wrongConfirm lot size follows the historical date
Optimistic fillsA strategy that is only good on thin strikesRe-test restricted to liquid strikes
Too short a sampleSmooth equity curve, few tradesCheck the trade count before anything else
The second row is the one you cause yourself, and it is invisible in the report. A tool cannot know how many variants you discarded, so only you can account for it.

Free versus paid

Free backtesting exists and is genuinely useful for learning the mechanics. The limits usually appear in the same places.
Typically limited on free tiersWhy it matters
Historical rangeA short range cannot contain multiple market regimes
Data granularityEnd-of-day data cannot test intraday stop losses
Number of legsTwo-leg caps exclude condors and butterflies
Number of runsThe fragility check needs several runs of the same strategy
Cost modellingGross-only results flatter every strategy
The third and fifth rows are the ones that most affect whether the result means anything. A free tool that tests two legs gross over six months will tell you something, but not what you need to know before funding a strategy.

How Arthalab handles it

For transparency, the specifics on this platform: 10 backtest credits a day, resetting at midnight IST. One normal backtest uses one credit. A bucket backtest uses one credit per active strategy inside it, so a bucket of four costs four. If a run fails to start, the credit is returned. Unused credits do not carry over.
Backtests use the lot size applicable to each historical date, so a report spanning a revision stays internally consistent. Testing covers NIFTY and SENSEX index options on NSE and BSE.

The limitation worth knowing

The honest constraint is the daily cap. Ten runs is enough for genuine iteration including fragility checks, but a workflow that sweeps dozens of parameter combinations in an afternoon will hit it — which, as it happens, is usually a process worth changing rather than a cap worth raising.

Backtesting is not the same as validation

A distinction worth drawing, because tools often blur it. Backtesting measures a strategy against history. Validation asks whether that measurement means anything.
A single backtest over a single period is a measurement. It becomes evidence only once you have checked that it is not an artefact of the period, the parameters or your own search process.

What validation actually involves

Three checks turn a measurement into evidence:
  1. Fragility. Change one parameter slightly. A robust result degrades gently; a fitted one collapses.
  2. A second period. Run a non-overlapping range. Two reports that look like siblings suggest the effect is real; two that look like different strategies suggest it is not.
  3. Parameters fixed in advance. If you chose the values after seeing results, part of the apparent edge belongs to your search rather than the market.
A tool that makes all three cheap is more valuable than one with a prettier chart. A tool that makes them expensive is quietly discouraging the work that matters most.

A fair evaluation sequence

1

Build one real strategy on each tool

The same strategy, with the same rules. Differences in the report are then about the tool, not the strategy.
2

Run the same date range

A tool with a different default range will produce a different answer for reasons that have nothing to do with quality.
3

Check whether costs are applied

If the tool does not say, assume gross and compute them yourself.
4

Run the fragility check on each

Nudge one parameter. A tool that makes this painful is discouraging the most important test.
5

Compare the reports field by field

Not the profit figures — the fields. A tool missing drawdown duration or trade count is not comparable.

The short version

  • Data granularity decides whether intraday rules can be tested at all
  • A report without max drawdown, duration and trade count cannot support a decision
  • Cost modelling is where tools differ most and flatter strategies most
  • Cheap re-runs matter, because the fragility check is the best anti-fitting test
  • Lot-size handling must follow the historical date, not today's figure
  • Every backtest assumes perfect fills and unlimited margin — treat results as a ceiling
A good backtest earns a strategy the right to be paper traded, not the right to be funded. No tool changes that.

Frequently asked questions

Backtesting measures a strategy against history. Validation checks whether that measurement is real rather than an artefact of the period, the parameters or your own search. A single run is a measurement, not evidence.

Treat it as a claim to verify rather than a result to expect. Check whether drawdown and trade count are shown, what period it covers, and whether costs were applied.

More is generally better, but the composition matters more than the length. A long quiet period tells you less than a shorter one containing a trending stretch, a range-bound stretch and a volatility shock.

Usually data granularity and cost modelling. Run the same strategy over the same dates on both, then compare the report fields rather than the headline profit.

The one whose report shows you max drawdown, drawdown duration, trade count and realistic costs, and which makes re-running at varied parameters cheap. Those four things matter more than the interface.

For learning the mechanics, yes. The limits usually appear in historical range, data granularity, number of legs and cost modelling — and those are exactly the limits that decide whether a result means anything.

Usually data granularity and cost modelling. A tool using end-of-day data cannot evaluate an intraday stop loss, and a tool applying gross costs will show a materially better figure than one charging per leg.

Enough to contain more than one market regime, several expiry cycles and at least one volatility shock. A quiet period will flatter almost any option-selling strategy.

It is included in the plan, with 10 credits per day resetting at midnight IST. A bucket backtest uses one credit per active strategy, and failed runs are refunded.

Assumed fills, costs, margin constraints and sometimes curve fitting. The gap is structural rather than a fault in the tool.

Yes, and per leg rather than per trade. A four-leg strategy pays charges eight times per round trip, which is enough to turn a thin edge negative.

Cheap re-runs. The fragility check — change one parameter and confirm results degrade gently rather than collapse — catches more bad strategies than any individual metric.

Start with a free 3-day trial

Build a strategy, backtest it and run it on paper — no broker, no IP and no money needed to try it.

Ask us on Telegram
Best Options Backtesting Software in India: What to Look For | Arthalab — Algo Trading India