Choosing Regression Windows for Return Prediction
Summary
The document considers whether to estimate a return prediction model with a fixed-length rolling window or with all available historical observations. A shorter window can respond more quickly when the relationship changes, while a longer history may be more stable but can retain data from an outdated regime. The question concerns predicting returns from past returns, with weekly estimation as an example.
The answer gives no rule that works universally: the better choice depends on the data, absent additional assumptions. It recommends judging alternatives empirically with properly separated training and test samples. A complex model that fits noise in training should generally perform poorly on held-out data, though trying many alternatives can still produce an apparently successful result by chance. The discussion offers no specific window-selection procedure, dataset, or empirical comparison, so it is guidance about evaluation rather than evidence favoring either approach.
Key ideas
- The best regression window length depends on the data and cannot be established without further assumptions.
- A rolling fixed-length window can adapt to changing relationships but may be sensitive to outliers.
- Using all available observations can be slower to reflect structural changes.
- Compare candidate approaches with separate training and test samples.
- Testing many alternatives creates a chance of selecting a model that succeeds by luck.
Tags
Full text
# Window length for predictive regressions # Window length for predictive regressions I am building a trading strategy that predicts the current period returns using historical returns (think e.g. using an estimated OLS model to predict next weeks return based on this weeks return). However, I am at loss in picking the window I should use for estimating the model. The way I see it, there are two ways I can do this: a) pick a fixed window length - e.g. 1 year (52 weekly observations), and re-estimate the model every week. However, depending on the asset the slope of the regression tends to change, and is especially suspect to few 'outlier' cases, which makes me question whether the model is still theoretically sound. b) use all available data, and roll to window forward every week, re-estimating the model. However, if the relationship is time variant, I think this approach will lead to extended periods of negative returns if there is a break in the model that a shorter window might capture better. How should I go about determining which method to use? I could of course back test different filtering strategies based on these two, but the more complicated, the more risk of overfitting IMO, and hence I'd prefer a more simplistic but statistically sound method for determining which way to go. Any suggestions? Would also love to read good papers that deal with this issue if any come to mind. ## Answer by Richard Hardy (score 2, accepted) https://quant.stackexchange.com/a/29542 Which strategy will work better is an empirical question that depends on the data at hand. That is, you cannot prove theoretically that one approach is better than the other without some extra assumptions. > I could of course back test different filtering strategies based on these two, but the more complicated, the more risk of overfitting IMO As long as you properly split the sample into a training and test subsamples, model complexity does not play a role. A model that overfits in the training sample will perform poorly in the test sample, and you will see it. (Of course, if you have a large number of alternative models, one or a few of them may perform well in both the training and the test subsample due to pure luck; but it does not seem you are facing this kind of setting.)
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.