Purging Overlapping Labels in Financial Cross-Validation
Summary
The document explains purging in financial machine learning through a prediction task that uses current stock returns to forecast the next day’s market return. The accepted answer illustrates the issue with a label based on a multi-day future return: labels assigned to adjacent observations share some of the same future returns. If similar observations fall into training and testing sets, this overlap can make performance look better than it is.
Purging removes training observations whose label periods overlap with those in the test set. The response also explains why combinatorial purged cross-validation can place training observations later in time than a test block: it creates additional historical performance simulations, while purging and embargoing reduce leakage across the split. The answer does not give a full implementation for the user’s one-day label example. The necessary purge interval depends on the label’s time span, and the method should be matched to the specific validation design.
Key ideas
- Labels based on future returns can overlap across adjacent observations.
- Training observations with label periods that overlap the test period should be purged to limit leakage.
- Combinatorial purged cross-validation can use training periods both before and after a test block.
- Purging and embargoing address leakage risks, but the appropriate interval depends on label construction.
Tags
Full text
# How to use 'purging' in predicting stock price tomorrow based on information today? # How to use 'purging' in predicting stock price tomorrow based on information today? #### Q1. How to create an 'overlap' when we predict a stock price tomorrow based on information today? - According to the book 'Advances in Financial Machine Learning' written by Marcos Lopez de Prado, the concept 'purging' is introduced to reduce the information leakage from the training dataset to train dataset. > One way to reduce leakage is to purge from the training set all observations whose labels overlapped in time with those labels included in the testing set. I call this process “purging.” - If my trading model is predicting stock price tomorrow using any information today, how can I perform "purging"? - To be more precise, I will give you an example. Let's say I use 3 companies' daily returns, GOOG, AAPL and MSFT to predict NASDAQ's daily return tomorrow. The input for my model is daily returns of 3 stocks, and the output is 1 daily return of NASDAQ tomorrow. - How can I create an 'overlap' between today and tomorrow? #### Q2. Can 'train dataset' appear after the 'test dataset'? - In the figure 7.2 presented below, a train dataset is from the time period after the test dataset. - It seems bizarre to me because usually in finance, we predict future based on past information. We don't predict or forecast the past based on the future data. - As such, it is against my intuition to have a train dataset after the test dataset in terms of time period. ## Answer by Quantoisseur (score 2) https://quant.stackexchange.com/a/61080 I have a video that goes into detail on this. Q1: Lets say your label is sign(5-day future return). If you look at 2 back-to-back days in a historical sample, the first day's label (t -> t+4) will contain 4 future days' returns (t+1 -> t+4) that are also included in the next days label (t+1 -> t+5). You want to eliminate this overlap when creating the training and testing sets to avoid inflated results during that period from serial correlation in the features when predicting a similar label. Q2: Combinatorial purged cross-validation (CPCV) is used to generate more historical performance simulations than traditionally available from walk-forward evaluation. The purpose of purging and embargoing is to reduce the impact/leakage from having training periods after testing periods.
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.