A backtest is better at rejecting a strategy than selecting one. The moment you use historical performance to choose among many variants, the backtest becomes part of the experiment, and the apparent winner starts carrying a hidden cost: selection bias.
This sounds backward. We backtest because we want to know which strategy works. But each choice guided by the same history spends some of its evidence. Change the lookback, threshold, universe, rebalance day, or stop rule after seeing results, and you’ve asked the data another question. Eventually it has heard enough questions to give you a convincing answer by chance.
Why does testing more strategies make the best result less trustworthy?
Imagine testing 100 strategies that have no real edge. Their returns will differ anyway. If their performance estimates are roughly independent and noise is normally distributed, the best Sharpe among them will be positive even though every strategy’s true Sharpe is zero.
A rough approximation for the expected maximum of 100 independent standard normal draws is 2.5 standard deviations above the mean. Treat that as an illustration, not a correction you can paste into a report: real strategy variants are correlated, Sharpe estimates have their own sampling error, and returns rarely behave like clean normal draws. The practical point survives those caveats. The more candidates you inspect, the more likely the top score reflects favorable noise.
And a “candidate” is not only a saved backtest. It’s every choice made after looking at a result: widening the universe because the chart looks choppy, dropping a year that hurts, or changing a fee assumption until the curve recovers. Those decisions count even if nobody gave them a version number.
Keep a research ledger before you run the next variant
Record the full search, including discarded ideas and the reason each change was made. That gives you a more honest account of how much the result was selected and makes it harder to remember only the attractive trials.
| Record | Example | Why it matters |
|---|---|---|
| Hypothesis | Short-term reversal is stronger after unusually large BTC perp moves | Separates the idea from its eventual parameter settings |
| Variant and timestamp | 48-hour signal, 2026-10-09, run 014 | Counts the search and ties results to code and data |
| Selection rule | Choose using net Sharpe; minimum 100 trades | Shows what “best” meant before the comparison |
| Rejected variants | 12, 24, 72-hour windows; all retained | Stops the visible winner from erasing its competition |
A ledger won’t undo the search. It makes the search legible, which is the first step toward judging how much confidence the winner deserves.
Reserve data the research process cannot keep revisiting
Split data by time, then protect the final period. Develop the strategy on one span, make decisions using a validation span, and evaluate the frozen strategy once on a later holdout. For example, you might use 2018–2022 for development, 2023 for validation, then reserve 2024–2025. Those dates are only a sketch: the market, strategy horizon, and data coverage should determine the split.
Write down what “frozen” means: signal, universe, sizing, execution assumptions, and any rules for missing data. If the holdout disappoints and you tune the strategy against it, that period has become validation data. The next honest evaluation needs fresh evidence.
This is also where small samples bite. A strategy that trades 18 times in a holdout hasn’t earned a precise verdict. Report the trade count and uncertainty alongside the return statistics; a single average can make a thin sample look settled.
Does this mean backtests are useless for choosing strategies?
No. Critics are right that strict holdouts can be wasteful: markets change, histories are short, and an untouched period may represent only one unusual regime. A carefully designed walk-forward process can show how a strategy behaves as the available history rolls forward. It still doesn’t make repeated tuning free. If you inspect each fold and revise the strategy, those folds influenced the research.
Use backtests to expose failure modes, compare a small number of ideas grounded in a market mechanism, and test whether results survive plausible fees, funding, impact, and parameter changes. Keep a ledger. Preserve a final holdout. Then take a promising, frozen version into paper trading, where operational mismatches can surface.
A backtest’s strongest answer is often, “This idea breaks under conditions we can already see.” When it says, “This is the best strategy,” ask how many other answers it was asked to give.
← All posts


