Overlapping Rolling Returns and Leakage in Forecasting
Summary
The question considers forecasting a future multi-day return from a history of rolling returns. Because adjacent rolling windows share most of their daily observations, the inputs are strongly autocorrelated. The response says that autocorrelation alone does not make the modeling approach invalid; using past returns to forecast future returns is related to momentum strategies.
The main practical warning is timing. At prediction time, features must contain only information available then, and the target must represent the return over the intended future interval. Using a return that ends at the current close assumes the strategy can observe that close and transact immediately at the same price. The answer suggests lagging the feature window to avoid that execution assumption. It does not provide empirical results or a detailed treatment of overlapping targets, validation splits, or model-specific performance, so those issues still need separate checks.
Key ideas
- Adjacent rolling returns overlap in their observations, which creates autocorrelation.
- Autocorrelation by itself does not make a forecasting model technically invalid.
- Features used at a forecast time must not include information that becomes available later.
- A signal based on the current closing price may assume an unrealistic immediate fill at that same price.
- Past-return forecasting is connected to momentum methods, but the document offers no empirical performance evidence.
Tags
Full text
# Is there an issue with estimating future returns from autocorrelated returns?
# Is there an issue with estimating future returns from autocorrelated returns?
I have a time series $X_t$ generated from a standard GBM
$$dS_t = \mu S_t dt + \sigma S_t dW_t$$
If I take the log returns over a rolling window of length $l$
$$r^{(l)}_i = \log \left( \frac{S_i}{S_{i-l}} \right)$$
then the $r^{(l)}_i$'s will be highly autocorrelated.
For example, in python, we can calculate a 5 day rolling return ($l=5$) by
```
df # pandas dataframe
>>> date price
2006-03-01 65.72
2006-03-02 62.91
...
df["rolling_5_day_returns"] = np.log(df["price"].shift(-5)) - np.log(df["price"])
```
Given the autocorrelation, is there any technical issue in training a model $f$ to estimate the return $\hat r_{i + l}$ at time $i$? That is, at time $i$ we will have an estimate of price difference between times $i$ and $l$.
Explicitly, we have
$$ f(\vec r^{(l)}_i) = \hat r_{i+l} =\log \left( \frac{S_{i+l}}{S_i} \right) $$ $$ f(\vec r^{(l)}_{i+1}) = \hat r_{i+l+1} =\log \left( \frac{S_{i+l+1}}{S_{i+1}} \right) $$ $$ ... $$
where $\vec r^{(l)}_i$ is a vector of historic rolling window returns up to and including time $i$, and $ \hat r_{i+l}$ is the estimate of the return between the current time $i$ and future time $i+l$.
EDIT: When I say technical issue I mean, is it incorrect to train a model on such data, or will the model suffer in performance if the data is autocorrelated?
Also, I plan to train an LSTM model, but should it matter what model we train whether it's a neural net, regression or ARIMA? The data is always the same.
## Answer by wjamdanf1234 (score 2)
https://quant.stackexchange.com/a/50828
If you are predicting the return from time "i" to time "i+l" then you cannot use any information beyond time "i" to train your model. As it appears you are getting returns from "i-5" to "i" and assuming that this same relationship will hold into the future from day "i" to "i+5". In theory there is nothing glaringly bad about this approach, but I would probably advise you, to be conservative, to get the return from "i-6" to "i-1", and then use this to predict the return from "i" to "i+5" to avoid problem of knowing the return at the close of day "i", and then instantly transacting in the same stock right away. I would also suggest reading up on basic momentum strategies, where you use past returns to predict future returns. Good luck!Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.