What a backtest actually does
The assumptions every engine makes
- Every order fills. There is no queue and no rejection.
- It fills at the recorded price. No spread crossed, no partial fill.
- Margin is unlimited. The position is never refused for funds.
- The system was running. No missed logins, no outages.
Data quality, which decides everything downstream
| What to ask | Why it matters |
|---|---|
| What granularity is the data? | Minute-level data captures intraday stop losses that end-of-day data cannot |
| Does it include the full option chain? | Strike selection rules need every strike, not just the at-the-money one |
| How are expiry and lot-size changes handled? | A test spanning a revision must use the figures in force on each date |
Report fields that are non-negotiable
| Field | Why it is essential |
|---|---|
| Max drawdown | The worst stretch you would have had to sit through |
| Drawdown duration | How long it lasted — this is what makes people quit |
| Number of trades | Governs how much every other number is worth |
| Max consecutive losses | The psychological load the strategy imposes |
| Average win and average loss | Win rate alone is meaningless without these |
| Expectancy | The closest thing to a single honest summary |
Cost modelling, where tools differ most
How tools handle costs
- Gross only. No costs applied. Every strategy looks better than it is, and multi-leg strategies look dramatically better.
- A flat per-trade figure. Better, but it understates multi-leg structures, where charges apply per leg on both entry and exit.
- Per-leg charges including taxes. The only approach that reflects what a four-leg strategy actually pays.
Features that prevent self-deception
- Easy re-runs at varied parameters. The fragility check — nudge one value and confirm the result degrades gently — is the single best test for curve fitting, and it is only practical if re-running is cheap.
- Trade-by-trade output. Aggregate numbers hide a strategy whose profit comes from two outlier days.
- Portfolio or bucket testing. Testing several strategies together shows correlation that individual tests cannot.
- Date-range flexibility. Running two non-overlapping periods and comparing them is the closest thing to out-of-sample validation most retail tools offer.
Ways a backtest can be wrong without being buggy
| Issue | What it looks like | How to detect it |
|---|---|---|
| Look-ahead | Suspiciously good entries | Check the rule only uses data available at that instant |
| Survivorship in your search | The best of many variants | Count how many versions you tried before this one |
| Stale reference data | Rupee figures that feel wrong | Confirm lot size follows the historical date |
| Optimistic fills | A strategy that is only good on thin strikes | Re-test restricted to liquid strikes |
| Too short a sample | Smooth equity curve, few trades | Check the trade count before anything else |
Free versus paid
| Typically limited on free tiers | Why it matters |
|---|---|
| Historical range | A short range cannot contain multiple market regimes |
| Data granularity | End-of-day data cannot test intraday stop losses |
| Number of legs | Two-leg caps exclude condors and butterflies |
| Number of runs | The fragility check needs several runs of the same strategy |
| Cost modelling | Gross-only results flatter every strategy |
How Arthalab handles it
The limitation worth knowing
Backtesting is not the same as validation
What validation actually involves
- Fragility. Change one parameter slightly. A robust result degrades gently; a fitted one collapses.
- A second period. Run a non-overlapping range. Two reports that look like siblings suggest the effect is real; two that look like different strategies suggest it is not.
- Parameters fixed in advance. If you chose the values after seeing results, part of the apparent edge belongs to your search rather than the market.
A fair evaluation sequence
Build one real strategy on each tool
Run the same date range
Check whether costs are applied
Run the fragility check on each
Compare the reports field by field
The short version
- Data granularity decides whether intraday rules can be tested at all
- A report without max drawdown, duration and trade count cannot support a decision
- Cost modelling is where tools differ most and flatter strategies most
- Cheap re-runs matter, because the fragility check is the best anti-fitting test
- Lot-size handling must follow the historical date, not today's figure
- Every backtest assumes perfect fills and unlimited margin — treat results as a ceiling
Frequently asked questions
Backtesting measures a strategy against history. Validation checks whether that measurement is real rather than an artefact of the period, the parameters or your own search. A single run is a measurement, not evidence.
Treat it as a claim to verify rather than a result to expect. Check whether drawdown and trade count are shown, what period it covers, and whether costs were applied.
More is generally better, but the composition matters more than the length. A long quiet period tells you less than a shorter one containing a trending stretch, a range-bound stretch and a volatility shock.
Usually data granularity and cost modelling. Run the same strategy over the same dates on both, then compare the report fields rather than the headline profit.
The one whose report shows you max drawdown, drawdown duration, trade count and realistic costs, and which makes re-running at varied parameters cheap. Those four things matter more than the interface.
For learning the mechanics, yes. The limits usually appear in historical range, data granularity, number of legs and cost modelling — and those are exactly the limits that decide whether a result means anything.
Usually data granularity and cost modelling. A tool using end-of-day data cannot evaluate an intraday stop loss, and a tool applying gross costs will show a materially better figure than one charging per leg.
Enough to contain more than one market regime, several expiry cycles and at least one volatility shock. A quiet period will flatter almost any option-selling strategy.
It is included in the plan, with 10 credits per day resetting at midnight IST. A bucket backtest uses one credit per active strategy, and failed runs are refunded.
Assumed fills, costs, margin constraints and sometimes curve fitting. The gap is structural rather than a fault in the tool.
Yes, and per leg rather than per trade. A four-leg strategy pays charges eight times per round trip, which is enough to turn a thin edge negative.
Cheap re-runs. The fragility check — change one parameter and confirm results degrade gently rather than collapse — catches more bad strategies than any individual metric.
Start with a free 3-day trial
Build a strategy, backtest it and run it on paper — no broker, no IP and no money needed to try it.

