What Is Curve Fitting in Trading, and Why Do Backtests Stop Working?
Curve fitting is tuning a strategy until it matches the particular history it was tested on, and the research on it puts numbers to how fast that happens: ten variants of an idea with no edge are expected to hand you one with an in-sample Sharpe ratio of 1.57 and nothing out of sample.

Curve fitting in trading is tuning a strategy until it matches the particular stretch of history you tested it on rather than any behaviour that repeats. The research gives the problem a size: in work published in the Notices of the American Mathematical Society in May 2014, Bailey, Borwein, López de Prado and Zhu show that trying ten different configurations of a strategy with no real edge is expected to hand you one with an in-sample Sharpe ratio of 1.57, while all ten of them, the winner included, have an expected out-of-sample Sharpe ratio of zero.
That is the difficulty in one number: a backtest reports the best of however many attempts you made, and the count of attempts is what almost nobody writes down.
What is curve fitting in trading?
Curve fitting describes the same failure the research literature calls overfitting. Bailey, Borwein, López de Prado and Zhu define overfitting as a model aimed at the specific observations in a sample instead of the structure they share, and say the idea is borrowed from machine learning, where the sample a model is designed on is the training set and the sample it is judged on is the testing set.
A backtest, in the paper's terms, simulates an algorithmic strategy over a past period and works out what it would have made and lost. The same authors set one condition for calling a backtest realistic: the in-sample result has to be consistent with the out-of-sample one, which is a sharper test than it sounds, because the trader who writes the rules usually also chooses the sample.
The fitting is rarely confined to the headline logic. The same paper lists the parts of a strategy that get quietly tuned alongside the signal: the stop loss, the profit target, the assumed cost of capital, the entry threshold, the risk sizing. A trader who has optimised two moving average lengths and also moved the stop twice has made more choices against the same data than the parameter count suggests.
Is curve fitting the same thing as data snooping?
They describe the same failure at different scales. Park and Irwin, in their 2005 AgMAS research report on US futures markets, define data snooping, after White, as using one set of data more than once for inference or model selection. Curve fitting is what one trader does to one strategy; data snooping is the general case, and it covers the choices that sit above the parameters — which cost assumption lets a result survive, which estimation window flatters it, which market cooperates, which family of systems to search in the first place.
Park and Irwin also describe collective data snooping, crediting Denton for the point: if many researchers each pick one in-sample optimisation method and all of them work on the same dataset, the group is snooping the data even though no individual did anything they would recognise as snooping. For a retail trader the equivalent is inheriting a popular parameter set from everyone who tested it on the same history first.
How much does trying more variants inflate a backtest?
A great deal, and the growth is predictable. The Notices paper's Proposition 1 approximates the expected largest value in a sample of N independent standard-normal draws — which is what the best of N Sharpe ratios amounts to when none of the strategies has an edge — and supplies an upper bound for it: the square root of twice the natural log of N. Evaluating that bound at a few trial counts, with the paper's assumption that the trials are independent, shows how little searching it takes to manufacture a figure that looks like evidence.
| Independent configurations tried | Upper bound on the expected best in-sample Sharpe ratio, no real edge |
|---|---|
| 10 | 2.15 |
| 32 | 2.63 |
| 100 | 3.03 |
| 1,000 | 3.72 |
| 8,800 | 4.26 |
The last row matches the paper's own worked example, whose grid is a parameter mesh rather than the independent trials the bound assumes. The authors generated 1,000 daily prices as a random walk — about four years — and searched a four-dimensional grid of 8,800 combinations for the best monthly trading rule: which business day to enter, how many days to hold, how large a loss triggers an exit, and whether to be long or short. The best node returned a Sharpe ratio of 1.27 and a probabilistic Sharpe ratio statistic of 2.83, implying under a 1% chance that its true Sharpe ratio was below zero. No seasonal effect existed in the data to find. The authors note the same demonstration works for trend following, momentum and mean reversion.
Trial counts also grow faster than parameter counts: one switch with two settings gives N of 2, and four more switches take N to 32. Hence the paper's warning that extra caution is needed with technical analysis, which it says mostly runs on filters: conditions that have to be met before a trade is taken, and conditions multiply.
How long does a backtest have to be before its result means anything?
Long enough that the number of configurations you tried cannot have produced the result by itself. The Notices paper turns that into a quantity it calls the Minimum Backtest Length, and the figures it works out are small. To stop a skill-less search from generating an in-sample Sharpe ratio of 1 against an expected out-of-sample Sharpe ratio of zero, five years of data supports no more than 45 independent configurations, and a two-year backtest supports seven.
Seven whole variants of the idea, including the ones discarded in the first hour.
Two qualifications matter before anyone reaches for the formula. The count assumes the trials are statistically independent, which the authors say makes the estimate a conservative one. They are also explicit that the length is a necessary condition without being a sufficient one: a backtest run on a longer sample can still be overfit.
This is why the paper presses so hard on one disclosure. A researcher who does not report how many configurations were tried makes the risk of overfitting impossible to assess, the authors write, and they add that the number is almost never reported.
Does walk-forward optimisation fix curve fitting?
It removes the most blatant form without guaranteeing a profitable result. Park and Irwin's futures test ran exactly that procedure. For each system and market they simulated the previous three years across a wide parameter range, took the best-scoring parameters, traded only those for the following year, then repeated annually — the parameters used in 1993 were the ones that scored best on mean net return over 1990 to 1992, and once 1993 closed the window rolled forward to pick the settings for 1994.
Their method was taken from earlier work rather than chosen after seeing which method performed well, which would itself be a snooping decision.
Across twelve US futures markets and twelve trading systems, the aggregate annual mean net return over 1978-1984 — with the financial markets starting in 1980 — reached 4.13%, which on a Sharpe ratio of 0.53 cleared statistical significance only at the 10% level. Running the same adaptive machinery forward over 1985-2003, no system earned a positive net return on the twelve-market portfolio and the aggregate came out at -5.82%. Across the full 1978-2003 period it was -3.14%. Every one of those figures is net of transaction costs: a $100 charge on each round turn of a single contract, which Park and Irwin say covers the bid-ask spread as well as the commission.
Park and Irwin add the caveat that stops walk-forward being read as proof: a rule that performs well in both periods is less likely to have been snooped, and could still have been profitable in both by chance. Our post on the moving average crossover covers what the same test found for that rule.
Why do US rules make a published backtest carry a hindsight warning?
Because the statement the CFTC prescribes for a commodity pool operator's or trading advisor's advertising names hindsight as one of the inherent limitations of simulated performance. Under 17 CFR 4.41(b)(1), as of the 2025 edition of the Code of Federal Regulations, nobody may show anyone the simulated or hypothetical results of a commodity interest account or set of transactions belonging to a commodity pool operator, a commodity trading advisor or a principal of either without attaching one of two statements: the one written out in the rule, or one a registered futures association has prescribed. The CFTC's version tells the reader the figures do not represent actual trading, that because the trades were never executed they may overstate or understate the effect of market factors such as a lack of liquidity, and that simulated programs in general are, in the rule's words, "designed with the benefit of hindsight". Where the presentation is not oral, the statement has to be prominent and immediately beside the figures, and the section binds a CPO or CTA exempt from registration as well.
Do the CFTC and NFA hypothetical-performance rules apply to a backtest you ran yourself?
No. Both bind the promotional material of the firms and people they name, not anyone's private research: 4.41 is the CFR's advertising rule for commodity pool operators, trading advisors and their principals, and it reaches them whether or not they are registered, while NFA Compliance Rule 2-29 governs promotional material from members and associates registered as FCMs, IBs, CPOs or CTAs. Neither section addresses a backtest you ran for yourself.
The National Futures Association prescribes its own disclaimer in Compliance Rule 2-29(c)(1), printed in capitals, which makes the same point about hindsight and adds that hypothetical trading involves no financial risk and that no hypothetical record can fully reflect what financial risk does once real money is on the line. Two of the rule's conditions go past disclosure. Under 2-29(c)(3), a member referencing hypothetical results for one of its own trading systems has to give comparable figures from the customer accounts it directs under power of attorney, covering at least five years or its whole history if that is shorter. Under 2-29(c)(4), once a member has three months of actual results for a system, it may no longer use promotional material referencing hypothetical ones for that system.
Both are worth reading anyway, because hindsight is the item on the regulator's list of inherent limitations that the mathematics above puts a number on.
What does curve fitting cost on a funded evaluation account?
It costs the account, because a result that came out of a search gives no reason to expect profit going forward, while the drawdown line is in place from the first trade. How that line is calculated is in the help center's drawdown explained article, and what it implies for size per trade is in our post on futures position sizing.
Size is where the damage compounds. The Notices authors make the point directly: believing a performance figure that overfitting inflated tends to bring too much leverage with it, so the fitted edge and the position taken on the strength of it compound each other. A size chosen for a Sharpe ratio that was never there spends room the real outcomes then need, and on an evaluation that room is the distance to the drawdown line.
One hazard is specific to simulated accounts. Our prohibited conduct rules bar strategies that only work because of how a simulated environment fills orders, including stops filling at prices a real market would have gapped through. A strategy optimised against a simulator's fills has been fitted to the platform, which is curve fitting with a rule breach attached. Where else a simulated fill and a real one part company is the subject of our post on sim trading versus live trading.
The Notices paper's conclusion is blunter than the familiar disclaimer. Advisers who leave backtest overfitting uncontrolled, the authors argue, turn strong backtested performance into an indicator of negative results ahead, which is why they call the customary line about past performance not indicating future results too optimistic in this context. Our risk disclosure says most people who attempt this do not succeed.
FAQ
How many parameters is too many in a trading strategy?
There is no threshold, but the Notices paper shows why the count matters more than it looks. A single switch with two settings gives two possible configurations; five such switches give thirty-two. Since the expected best in-sample result grows with the number of configurations tried, every parameter you add raises the score a strategy with no edge is expected to post. The paper's Minimum Backtest Length puts the trade-off in years: with five years of data, no more than 45 independent configurations.
Does out-of-sample testing prove a strategy works?
No. Park and Irwin, whose futures test re-chose parameters every year from the prior three years only, state that a rule performing well in both the in-sample and out-of-sample periods is less likely to have been snooped, and could still have been profitable in both by chance. The Notices authors make the matching point about backtest length: a long enough sample is a necessary condition for avoiding overfitting, not a sufficient one.
Can I trade an automated or backtested system on a funded account here?
Your own automation is allowed under our prohibited conduct rules; purchased bots, third-party systems and subscribed signal services are not, and we may ask you to demonstrate that the strategy is yours. Separately, any strategy whose results depend on how a simulated environment fills orders is barred, which rules out systems optimised against a simulator's fill behaviour rather than against the market.
Sources
- Bailey, Borwein, López de Prado & Zhu — Pseudo-Mathematics and Financial Charlatanism: The Effects of Backtest Overfitting on Out-of-Sample Performance, Notices of the American Mathematical Society, May 2014
- Park & Irwin, AgMAS Project Research Report 2005-04 (University of Illinois) — The Profitability of Technical Trading Rules in US Futures Markets: A Data Snooping Free Test
- US Government Publishing Office — 17 CFR 4.41, Advertising by commodity pool operators, commodity trading advisors, and the principals thereof (2025 edition)
- National Futures Association — Compliance Rule 2-29, Communications with the Public and Promotional Material
- Lumen Futures Help Center — Prohibited conduct
- Lumen Futures Help Center — Drawdown explained
Educational content about futures markets and simulated trading. Not investment advice, and not a solicitation to trade. Trading futures involves substantial risk of loss. Read the full risk disclosure.