Detecting Lookahead Bias in Trading Strategy Backtests
Summary
The document explains lookahead bias: a backtest can accidentally use future candle data because the full historical dataframe is loaded before indicators and signals are calculated. This can make results appear unrealistically strong. It describes an automated analysis that first runs a baseline backtest, then repeats tests for individual entries and exits, comparing indicator values and signal locations to detect changes associated with future-data leakage. The method requires historical data and enough trades to evaluate the selected signals.
Examples of potential sources include negative shifts, fixed-row dataframe access, uncontrolled loops, whole-dataframe aggregations without rolling windows, and certain indicator settings. The analysis can miss bias in signals that never trigger, and options such as limit orders or position stacking may distort findings or create false positives. Its results are therefore a diagnostic for tested signals, not proof that every part of a strategy is unbiased. The document also cautions that removing bias may greatly reduce the apparent performance of a strategy whose results depended on it.
Key ideas
- Loading all candles at once can allow indicators or signals to use information unavailable at decision time.
- The analysis compares a baseline backtest with separate verification runs for entries and exits.
- Negative shifts, full-sample aggregates, and careless row access can introduce lookahead bias.
- Signals that do not occur during analysis remain unchecked, which can produce false negatives.
- Limit orders and other backtest settings can cause misleading flags, so findings need interpretation.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.