Curve Fitting in Backtesting: How to Spot It Before It Costs You
Curve fitting is tuning a strategy until it describes one past rather than any real market behaviour. How it happens, the tests that detect it, and the process changes that prevent it.
Arthalab7 min read
Curve fitting is adjusting a strategy until it describes one specific past rather than any durable market behaviour. It produces excellent backtests and poor live results, and it happens to almost everyone who tests strategies without a process designed to prevent it.
How it happens without anyone deciding to do it
Nobody sets out to curve fit. It emerges from a sequence that feels like diligence at every step.
You build a strategy and backtest it. The result is acceptable but unexciting.
You try a slightly tighter stop loss. The result improves.
You try tightening it further. It improves again.
You adjust the entry time by a few minutes. Better still.
You settle on the combination that produced the best number and deploy it.
Every individual step looks like sensible optimisation. The outcome is a strategy tuned to the noise of one sample, and the noise will not repeat.
Why it fails live
A historical period contains both signal — genuine, repeating market behaviour — and noise, which is everything specific to that period.
A moderately tuned strategy captures mostly signal. A heavily tuned one captures signal plus an increasing share of noise, because the fine adjustments can only be fitting to the specific sequence of outcomes in that sample.
What survives and what does not
Going forward, the signal may persist. The noise certainly will not, because it was never a property of the market — only of that stretch of it.
The plateau test
The single most useful detection method, and it takes two extra backtest runs.
Run your strategy at the chosen parameter, then at values either side of it. Plot or simply compare the three results.
Pattern
What it means
Results similar across all three
A plateau — the parameter is capturing something real
Chosen value much better than both neighbours
A peak — almost certainly fitted to noise
Results swing wildly
The strategy is unstable regardless of the value
This costs backtest credits, and it is the best use of them available. Three runs of one strategy tells you more than one run of three strategies.
The two-period test
The second detection method, and arguably the stronger one.
Run the same strategy over two non-overlapping periods and compare the reports side by side.
Two reports that look like siblings — different numbers, similar shape, comparable drawdown character — suggest the effect is real.
Two reports that look like different strategies suggest the result in one of them was specific to that period.
This is the closest thing to out-of-sample validation available without building a separate process for it, and it costs two credits.
Counting your own search
The hardest form to detect, because it leaves no trace in any report.
If you tested twenty variations and kept the best, the winner's result includes a selection effect. Some of its apparent edge belongs to the search rather than to the market — and a backtest report cannot know how many versions you discarded.
The only defence that works
The defence is procedural rather than analytical: write down the parameters you intend to test before running anything, and keep a count of how many variants you actually tried. A result arrived at after twenty attempts deserves heavier discounting than one arrived at on the first.
Signs a strategy is probably fitted
The parameters are oddly specific — a 23-point stop rather than 25 or 30
Performance collapses when any single value is nudged
The strategy has many parameters relative to the number of trades
It was arrived at after a long series of adjustments
It performs very differently across two halves of the test period
You cannot articulate a reason for each parameter beyond that it tested well
The last item is the most diagnostic. If the only answer to "why 23 points?" is "it backtested best", the value is describing a sample rather than a reason.
Building a process that resists it
1
Decide parameters before testing
Choose defensible values from reasoning about the structure, then test them. Not the other way round.
2
Write them down
So you cannot quietly revise history about what you intended.
3
Test once over a long, varied period
Rather than many times over a short one.
4
Run the plateau test
Confirm the result degrades gently either side of your chosen values.
5
Run a second, non-overlapping period
Two reports that disagree are telling you something important.
6
Paper trade before funding
Paper trading runs forward in real time, which is the one thing no historical test can do.
Simplicity as a defence
A strategy with fewer parameters has fewer opportunities to be fitted. This is worth more than it sounds.
Each additional adjustable value multiplies the combinations you can search, and every combination searched increases the chance that the best one is best by accident.
Why fewer parameters help
A simple strategy with slightly worse historical numbers frequently outperforms a complex one with better numbers, because less of its result was manufactured by the search.
What this means for published backtests
The same scepticism applies to results published by anyone else, including analysts and platforms.
You cannot see how many variants were tried before the published one. A published backtest is a claim to verify, not a result to expect, and the questions are the same: how long was the sample, how many trades, what did the drawdown look like, and does it hold up across sub-periods.
The short version
Curve fitting is tuning until the strategy describes one past rather than the market
It emerges from steps that each look like sensible optimisation
Plateau test: good parameters degrade gently either side, fitted ones collapse
Two-period test: robust strategies produce reports that look like siblings
Count your own search — a result found after twenty attempts means less
Fewer parameters means fewer chances to fit noise
Frequently asked questions
Adjusting a strategy's parameters until it performs well on a specific historical sample. The result describes that period rather than durable market behaviour, so it rarely survives going forward.
Run it at values either side of your chosen parameters. A robust strategy degrades gently; a fitted one collapses. Then run a second, non-overlapping period and check the two reports resemble each other.
No, but there is a difference between choosing a defensible value and searching for the value that tested best. The first is reasoning; the second is fitting.
A 23-point stop rather than 25 or 30 usually indicates the number was found by searching rather than chosen for a reason. The precision is a symptom of the search.
Relative to the number of trades. Each adjustable value multiplies the combinations available to search, and every search increases the chance the winner is best by accident.
Not directly, because you cannot see how many variants were tried. Check the sample length, the trade count, the drawdown, and whether it holds up across sub-periods.
Partly. It runs forward in real time on data that was not available when the strategy was built, which no historical test can replicate. A fitted strategy often disappoints there first.
Running the strategy at your chosen parameter and at values either side. Similar results across all three means the parameter captures something real; a sharp peak at your chosen value means it does not.
Start with a free 3-day trial
Build a strategy, backtest it and run it on paper — no broker, no IP and no money needed to try it.