Choosing Overlapping or Non-Overlapping Returns for Model Training
Summary
The note considers whether to use overlapping or non-overlapping returns as the dependent variable in a predictive regression. Overlapping samples can create concurrency: observations may share return intervals and therefore are not independent in the same way as disjoint samples. The response recommends non-overlapping returns when doing so does not sharply reduce the available training data.
Because enforcing non-overlap can leave a much shorter dataset, the answer describes sequential bootstrapping as an approach for ensemble methods such as random forests and bagging classifiers. It points to a financial machine-learning text for the concurrency concept and notes that a software package implements the described methods, but provides no empirical comparison or regression example. The recommendation is therefore a practical rule of thumb, not a universal result; the appropriate choice depends on the data loss and modeling setup.
Key ideas
- Overlapping return observations can share intervals and create concurrency in model training.
- Non-overlapping returns are preferable when they do not substantially shorten the training sample.
- Removing overlap can sharply reduce the number of observations available to fit a model.
- Sequential bootstrapping is presented as a remedy for concurrency in ensemble methods such as random forests and bagging classifiers.
Tags
Full text
# Overlapping vs Non-overlapping returns
# Overlapping vs Non-overlapping returns
Suppose I want to estimate the following regression: $R_t=\alpha + \beta X_{t-1} +\epsilon_t$. Where I use asset returns as the dependent variable. Both overlapping as well as non-overlapping returns can be used as the dependent variable. Which considerations do you have to make to choose between these two? What are the advantages and disadvantages of both approaches?
## Answer by Alexandr Proskurin (score 5)
https://quant.stackexchange.com/a/46569
Actually, overlapping samples is a big problem in financial machine learning which is called concurrency. Marcos Lopez de Prado discusses this issue in Chapter 4 of his book
> Advances in Financial Machine Learning
Ideally, non-overlapping returns should be used to train the model, however this constraint massively decreases the length of your training dataset that is why you need to solve this problem in other way. If you use ensemble methods (Random Forest, Bagging Classifier), Sequential Bootstrapping is the answer. To answer your question, prefer non-overlapping returns if it doesn't decrease the length of your dataset massively.
Note: I am one of the authors of mlfinlab package which implements concepts described in Marcos' book. Project link: https://github.com/hudson-and-thames/mlfinlabShown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.