Skip to content
All library documents

Lookahead Bias from Unrealistic Bar-Close Execution

Article Quant Q&A · Author: user20664

Summary

The document clarifies a common backtesting problem in a model that uses a completed candle’s closing price to generate a signal for the next period. The training setup is not automatically biased merely because it predicts a future value from past observations. The key issue is whether the simulated trade can actually be executed at the price used to calculate performance.

If a backtest forms a signal from close C(t) and measures profit from that same close to C(t+1), it assumes execution at C(t). In live trading, the close is only known once the bar has ended, by which time the market may have moved. The answer proposes measuring returns from the next bar’s open instead, which assumes execution there after the prior close signal. That is still an assumption: actual fills, latency, spread, and liquidity can affect results, and the size of the bias depends on market conditions.

Key ideas

  • Using completed data to predict a later period is not by itself lookahead bias.
  • A backtest can be unrealistic if it assumes execution at the close used to create the signal.
  • Using the next bar’s open models execution after a signal formed at the prior close.
  • The impact of the close-to-open gap depends on market liquidity and volatility.
  • Bar-based execution assumptions do not capture all costs or fill uncertainty.

Tags

Full text
# Trouble understanding lookahead bias


# Trouble understanding lookahead bias












I understand lookahead bias is pretty common industry knowledge. But I cannot wrap my head around how I am introducing it and could use a nice and easy explanation. Here's my thought process.

I have $N$ data points of OHLC data. Lets say for the sake of argument I pull $t = 1$ to $N$ from a database.

I train a model on the closing price of this instrument to predict the $t + 1$th value. I understand that by doing so I have introduced lookahead bias.

Where I struggle with is - when I go to predict the next time period's close with live data I'm going to

- Pull $N - 1$ data points

- Wait for this current candle to close (for the sake of argument, lets say I get this data as fast as possible)

- Add the new data point to the list of $N -1 $ data points I have pulled from my database

- Predict the $N+1$th data point (the NEXT time periods data point, in other words the now current data point's close)

- Make my trade

I feel like this is what I am doing in the training procedure, assuming that when I train on the $N$ data points, the $N$th data point has closed already.

Could someone take this example and explain to me where I'm making a mistake in my reasoning? I would really appreciate it.

EDIT

I should be more specific. Assuming I generate a $-1$ or $1$ signal for sell and buy respectively, and I am using a log return series.

## Answer by Alex C (score 2, accepted)

https://quant.stackexchange.com/a/32023

The issue is how do you evaluate the success of your trades. If the P&L in your simulation is measured as C(t+1)-C(t) then your simulation is not completely realistic, because in real life by the time you compute the signal based on C(t), even if you do so quickly, the price C(t) will not be the current price and you will not be able to buy at that price. How big a problem this is depends on how liquid and volatile the market is.

One solution might be to measure the P&L as C(t+1)-O(t+1). This assumes that you will buy at the next bar open, based on the signal computed at the previous bar close.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.