All articles
Backtesting

How to Read a Backtest Report (Without Fooling Yourself)

Max drawdown, expectancy, reward-to-risk, loss streaks and return-over-max-drawdown explained — and the order to read them in so the headline profit number does not mislead you.

Arthalab10 min read
Read a backtest report from the risk end, not the profit end. Start with the worst thing that happened, decide whether you could live through it, and only then look at what it earned. Reading it the other way round is how people fund strategies they cannot actually hold.

The order to read it in

  1. Max drawdown — the worst peak-to-trough fall. Convert it to rupees at your intended size before anything else.
  2. Drawdown duration — how long that stretch lasted. Depth is survivable; eight months underwater often is not.
  3. Max loss streak — consecutive losers. This is the number that makes people abandon a working strategy.
  4. Number of trades — your sample size. Everything below is unreliable if this is small.
  5. Expectancy — average outcome per trade. The single most honest summary figure.
  6. Win rate and reward-to-risk — always together, never apart.
  7. Return over max drawdown — profit earned per unit of worst pain.
  8. Total profit — last, because by now you know what it cost.
This ordering is not arbitrary. Each number constrains the ones after it. A drawdown you cannot tolerate makes the profit figure irrelevant, and a sample of eighteen trades makes every ratio below it noise.

Why this order and not another

Reading from profit downwards feels natural and produces worse decisions, for a specific reason: the profit figure is the one most contaminated by luck, and reading it first anchors everything after it.
Once you have seen an attractive total, every subsequent number gets evaluated against whether it would disqualify that total. A drawdown you would normally reject starts looking acceptable. A thin trade count starts looking like enough.

The anchoring problem

Reading from risk upwards inverts that. You decide what you can tolerate before you know what you would be paid for tolerating it, which is the only order in which that decision is honest.

Max drawdown is the real constraint

Everything else is theoretical if the drawdown exceeds what you will tolerate. A strategy returning well with a 40% drawdown is not a good strategy for someone who will stop it at 20% — it is a strategy that will be abandoned at the worst possible moment, locking in the loss and missing the recovery.

Why the figure is optimistic by construction

A second, subtler point: the max drawdown in your report is the worst that happened in that sample. It is not a ceiling. A longer or different period would very likely contain a worse one. Treat it as a lower bound on future pain, not an upper bound.

A report read properly, end to end

Take a report showing: total profit Rs.1,84,000; max drawdown 22%; drawdown duration 94 days; win rate 71%; 143 trades; max loss streak 7.
  1. Drawdown first. 22% of a Rs.5,00,000 deployment is Rs.1,10,000. Would you hold through losing that? If not, stop here.
  2. Duration. 94 days is roughly three months underwater. That is a long time to keep starting a bot you are currently losing on.
  3. Loss streak. Seven consecutive losers. Picture the seventh morning and whether you would start it.
  4. Sample. 143 trades is past the rough threshold, so the ratios are worth something.
  5. Win rate with reward-to-risk. 71% wins means losses must be larger than wins on average, or the profit figure would be much bigger. That is a shape where one bad run hurts.
  6. Return over drawdown. Rs.1,84,000 against a 22% worst fall — compute it and compare against alternatives rather than admiring it alone.
  7. Total profit, last. Now it means something, because you know what it cost.
Read in that order, the decision usually makes itself at step one or step three. Read in the reverse order, the same report looks like an obvious yes.

Drawdown duration, the underrated number

Depth gets the attention; duration does the damage. A 15% drawdown that recovers in three weeks is an inconvenience. The same 15% that takes seven months to recover is seven months of watching a strategy you chose underperform, while doubting every assumption that went into it.
Most strategies are abandoned during long shallow drawdowns, not sharp deep ones. The sharp ones at least feel like events; the long ones just feel like being wrong.

Win rate without reward-to-risk is meaningless

Two strategies, both with the same expectancy, read completely differently:
Strategy AStrategy B
Win rate85%35%
Average winSmallLarge
Average lossLargeSmall
Feels likeFrequent small wins, rare brutal lossesFrequent small losses, rare big wins
Fails whenOne bad day undoes monthsYou quit during a long losing run
SuitsTraders who can stomach rare shocksTraders who can stomach being wrong often
Neither is better. They demand different temperaments, and knowing which one you can actually hold matters more than the number on the summary line. Most option-selling strategies are shaped like A, which is worth knowing before you pick one.

Sample size

A strategy with 18 trades over two years has not been tested, it has been sampled. Results from small samples swing wildly on one or two outcomes, and every ratio computed from them inherits that instability.

A rough threshold

Treat anything under roughly a hundred trades as provisional. Be especially suspicious when a strategy you expected to trade frequently produces few trades — that usually means the entry condition is rarer than you thought, which is itself useful information.

The fragility check

Before trusting any report, re-run the same strategy with one parameter nudged slightly — entry time by five minutes, stop loss by a few points, strike selection by one step.
A robust strategy degrades gently: the numbers get a little worse and the shape stays the same. A curve-fitted one falls apart, sometimes flipping from profitable to not. This single test catches more bad strategies than any individual metric on the report.

Return over max drawdown

Total return divided by the worst drawdown. It expresses profit per unit of pain, which makes two strategies with different risk profiles directly comparable in a way raw returns never are.
A strategy returning 30% with a 10% drawdown and one returning 60% with a 40% drawdown are not obviously rankable by return alone. By this measure the first one is doing more with less, and can be scaled to match the second's return at lower risk — if your capital allows it.

Converting the report into a decision

A report is only useful if it produces a decision. Here is a sequence that turns the numbers into one.
1

Convert max drawdown to rupees at your intended size

This is the first and most important number. Everything else is conditional on your being able to accept it.
2

Ask whether you would hold through it

Not whether you should — whether you would. Those are different questions and only the second one predicts behaviour.
3

Check the trade count

Under roughly a hundred, treat every ratio as provisional and do not size up on it.
4

Run the fragility check

Nudge one parameter. If results collapse, reject the strategy regardless of how good the headline was.
5

Compare expectancy against costs

A thin edge per trade can be entirely consumed by brokerage and spread, especially on multi-leg structures.
6

Write down your stop condition

In rupees, before you deploy. This is the number you will need when you are least able to decide calmly.
If a strategy fails at step one or step four, nothing further matters. Most strategies that get funded and then abandoned failed one of those two and were deployed anyway.

Two reports, same strategy

A useful habit when comparing: run the same strategy over two non-overlapping periods and read both reports side by side.
A robust strategy produces two reports that look like siblings — different numbers, similar shape, comparable drawdown character. A fitted one produces two reports that look like different strategies entirely, and that disagreement is the signal.

What the report cannot show you

  • Real slippage — fills are assumed at historical prices.
  • Margin rejections — the backtest never runs out of funds.
  • Your own behaviour during the drawdown it is describing.
  • Whether the market regime that produced these results still exists.
  • What happens when one leg of a multi-leg structure fails to fill.
Why live results differ from backtests covers each of these in detail. The short version is that a backtest measures the strategy, while live trading measures the strategy plus the execution plus you.

The short version

Read the report in an order that protects you from the number you most want to see.
  • Max drawdown first, converted to rupees at your intended size
  • Then duration and loss streak, which decide whether you would hold on
  • Then sample size, because it governs how much the rest is worth
  • Win rate and reward-to-risk together, never separately
  • Total profit last, once you know what it cost
The max drawdown in any report is the worst that happened in that sample, not a ceiling. Treat it as a lower bound on future pain.

Frequently asked questions

No. Compare them on return over max drawdown, which expresses profit per unit of worst pain. Two strategies with the same profit and very different drawdowns are not equivalent.

Usually that the strategy is fitted to one of them. A robust strategy produces reports that differ in detail but agree in shape across non-overlapping periods.

Often, in practice. Deep drawdowns feel like events and are easier to sit through. Long shallow ones are where most people quit, because they feel like being persistently wrong.

The largest peak-to-trough fall in the strategy's equity over the test period. It is the single most important number on the report because it defines the worst stretch you would have had to sit through.

Positive, after costs, with a sample large enough to trust. There is no universal target — expectancy has to be read against the drawdown and the trade count that produced it.

Not on its own. A high win rate paired with large average losses can lose money overall. Always read it with reward-to-risk.

More is better. Below roughly a hundred, treat the conclusions as provisional — a couple of outcomes can swing every metric.

Total return divided by the worst drawdown. It expresses profit per unit of pain, which makes two strategies with different risk profiles easier to compare.

No. It is the worst that happened in that sample. A longer or different period would very likely contain a worse one, so treat it as a lower bound on future pain rather than a ceiling.

Because most strategies are abandoned during long drawdowns rather than deep ones. A recovery you did not stay invested for is not a recovery you received.

Start with a free 3-day trial

Build a strategy, backtest it and run it on paper — no broker, no IP and no money needed to try it.

Ask us on Telegram
How to Read a Backtest Report (Without Fooling Yourself) | Arthalab — Algo Trading India