October 9, 2026 · research

A Better Backtest Can Make Your Strategy Research Worse

A Better Backtest Can Make Your Strategy Research Worse

A backtest is better at rejecting a strategy than selecting one. The moment you use historical performance to choose among many variants, the backtest becomes part of the experiment, and the apparent winner starts carrying a hidden cost: selection bias.

This sounds backward. We backtest because we want to know which strategy works. But each choice guided by the same history spends some of its evidence. Change the lookback, threshold, universe, rebalance day, or stop rule after seeing results, and you’ve asked the data another question. Eventually it has heard enough questions to give you a convincing answer by chance.

Why does testing more strategies make the best result less trustworthy?

Imagine testing 100 strategies that have no real edge. Their returns will differ anyway. If their performance estimates are roughly independent and noise is normally distributed, the best Sharpe among them will be positive even though every strategy’s true Sharpe is zero.

A rough approximation for the expected maximum of 100 independent standard normal draws is 2.5 standard deviations above the mean. Treat that as an illustration, not a correction you can paste into a report: real strategy variants are correlated, Sharpe estimates have their own sampling error, and returns rarely behave like clean normal draws. The practical point survives those caveats. The more candidates you inspect, the more likely the top score reflects favorable noise.

And a “candidate” is not only a saved backtest. It’s every choice made after looking at a result: widening the universe because the chart looks choppy, dropping a year that hurts, or changing a fee assumption until the curve recovers. Those decisions count even if nobody gave them a version number.

Keep a research ledger before you run the next variant

Record the full search, including discarded ideas and the reason each change was made. That gives you a more honest account of how much the result was selected and makes it harder to remember only the attractive trials.

RecordExampleWhy it matters
HypothesisShort-term reversal is stronger after unusually large BTC perp movesSeparates the idea from its eventual parameter settings
Variant and timestamp48-hour signal, 2026-10-09, run 014Counts the search and ties results to code and data
Selection ruleChoose using net Sharpe; minimum 100 tradesShows what “best” meant before the comparison
Rejected variants12, 24, 72-hour windows; all retainedStops the visible winner from erasing its competition

A ledger won’t undo the search. It makes the search legible, which is the first step toward judging how much confidence the winner deserves.

Reserve data the research process cannot keep revisiting

Split data by time, then protect the final period. Develop the strategy on one span, make decisions using a validation span, and evaluate the frozen strategy once on a later holdout. For example, you might use 2018–2022 for development, 2023 for validation, then reserve 2024–2025. Those dates are only a sketch: the market, strategy horizon, and data coverage should determine the split.

Write down what “frozen” means: signal, universe, sizing, execution assumptions, and any rules for missing data. If the holdout disappoints and you tune the strategy against it, that period has become validation data. The next honest evaluation needs fresh evidence.

This is also where small samples bite. A strategy that trades 18 times in a holdout hasn’t earned a precise verdict. Report the trade count and uncertainty alongside the return statistics; a single average can make a thin sample look settled.

100illustrative null variants searched
2.5σrough expected best score, under simplifying assumptions
1holdout evaluation after the strategy is frozen

Does this mean backtests are useless for choosing strategies?

No. Critics are right that strict holdouts can be wasteful: markets change, histories are short, and an untouched period may represent only one unusual regime. A carefully designed walk-forward process can show how a strategy behaves as the available history rolls forward. It still doesn’t make repeated tuning free. If you inspect each fold and revise the strategy, those folds influenced the research.

Use backtests to expose failure modes, compare a small number of ideas grounded in a market mechanism, and test whether results survive plausible fees, funding, impact, and parameter changes. Keep a ledger. Preserve a final holdout. Then take a promising, frozen version into paper trading, where operational mismatches can surface.

A backtest’s strongest answer is often, “This idea breaks under conditions we can already see.” When it says, “This is the best strategy,” ask how many other answers it was asked to give.

strategy researchbacktestingoverfittingwalk-forwardpaper trading
← All posts