Skip to content
All library documents

Choosing Overlapping or Non-Overlapping Returns for Model Training

Article Quant Q&A · Author: amars96

Summary

The note considers whether to use overlapping or non-overlapping returns as the dependent variable in a predictive regression. Overlapping samples can create concurrency: observations may share return intervals and therefore are not independent in the same way as disjoint samples. The response recommends non-overlapping returns when doing so does not sharply reduce the available training data.

Because enforcing non-overlap can leave a much shorter dataset, the answer describes sequential bootstrapping as an approach for ensemble methods such as random forests and bagging classifiers. It points to a financial machine-learning text for the concurrency concept and notes that a software package implements the described methods, but provides no empirical comparison or regression example. The recommendation is therefore a practical rule of thumb, not a universal result; the appropriate choice depends on the data loss and modeling setup.

Key ideas

  • Overlapping return observations can share intervals and create concurrency in model training.
  • Non-overlapping returns are preferable when they do not substantially shorten the training sample.
  • Removing overlap can sharply reduce the number of observations available to fit a model.
  • Sequential bootstrapping is presented as a remedy for concurrency in ensemble methods such as random forests and bagging classifiers.

Tags

Full text
# Overlapping vs Non-overlapping returns


# Overlapping vs Non-overlapping returns












Suppose I want to estimate the following regression: $R_t=\alpha + \beta X_{t-1} +\epsilon_t$. Where I use asset returns as the dependent variable. Both overlapping as well as non-overlapping returns can be used as the dependent variable. Which considerations do you have to make to choose between these two? What are the advantages and disadvantages of both approaches?

## Answer by Alexandr  Proskurin (score 5)

https://quant.stackexchange.com/a/46569

Actually, overlapping samples is a big problem in financial machine learning which is called concurrency. Marcos Lopez de Prado discusses this issue in Chapter 4 of his book

> Advances in Financial Machine Learning

Ideally, non-overlapping returns should be used to train the model, however this constraint massively decreases the length of your training dataset that is why you need to solve this problem in other way. If you use ensemble methods (Random Forest, Bagging Classifier), Sequential Bootstrapping is the answer. To answer your question, prefer non-overlapping returns if it doesn't decrease the length of your dataset massively.

Note: I am one of the authors of mlfinlab package which implements concepts described in Marcos' book. Project link: https://github.com/hudson-and-thames/mlfinlab

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.