Artificial Time-Series Gaps in Combinatorial Purged Cross-Validation
Summary
The document raises a practical concern about combinatorial purged cross-validation (CPCV) in financial backtesting. By combining separated time blocks into a training sample, the method can make observations on either side of a gap appear adjacent. That can introduce artificial jumps into prices and derived features, indicators, or labels if calculations treat the stitched sample as continuous.
CPCV is presented as a way to evaluate a strategy across more scenarios and reduce false discoveries. The author questions whether results from scenarios with artificial discontinuities reflect conditions a live strategy would face, and wonders whether prices should be adjusted before calculations. The document does not answer that question or provide experiments comparing approaches. Its concern therefore identifies a validation issue for researchers to investigate, rather than establishing that CPCV is invalid or prescribing a preprocessing method. The effects may depend on how folds are constructed and which features or labels use observations near the gaps.
Key ideas
- CPCV forms multiple backtest scenarios by combining different time blocks.
- Stitching separated blocks can create artificial jumps at the boundaries.
- Derived features, indicators, and labels may be affected if calculations cross those boundaries.
- The document questions whether such scenarios represent live market conditions but offers no empirical resolution.
- Researchers should consider how gaps affect calculations when interpreting CPCV results.
Tags
Full text
# The discontinuity when applying the combinatorial purged cross-validation # The discontinuity when applying the combinatorial purged cross-validation In Marcos Lopez de Prado's book, Advances in financial machine learning, he recommends using the combinatorial purged cross-validation(CPCV) for backtesting. His motivation is sensible. Through the method, we can test our trading strategy with more scenarios and we can reduce the false discovery rate. Our training set is not continuous anymore with his CPCV method. As an example, our training set may combine the two time series, A(20018-01-01~2019-01-01) and B(2019-07-01~2020-01-01). We would have a big price jump at the time point where A and B are connected as there is a 6 months gap. But with his cross-validation method, we consider the time series as one continuous time series and calculate all quantities. We will have discontinuities of quantities such as features, target labels, indicators, etc at the point. These discontinuities in the quantities bother me. The discontinuities are artificial at most. I don't think we will meet the price jump in real-life and I feel the scenarios generated by CPCV are not the real one. Why do we bother to check our model in the unreal scenarios? Or do we need to take an additional preprocessing step to make the price continuously? (I doubt this though as Lopez didn't mention the step at all.)
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.