Combinatorial Purged Cross-Validation and Multivariate Time Series
Summary
The document introduces combinatorial purged cross-validation (CPCV) as a way to evaluate trading strategies on time series while accounting for leakage through purging and embargoing observations. It contrasts CPCV with ordinary k-fold and walk-forward validation: CPCV combines folds into multiple backtest paths, producing a set of Sharpe ratios that can describe variation in estimated performance rather than a single score.
The text then asks how to interpret this procedure when the data contain multiple series or features, and whether it can produce matrix-valued predictions. It does not answer those questions or specify a multivariate extension. Its useful contribution is therefore the high-level distinction between multivariate model inputs and the paths generated by validation; it leaves implementation, path construction details, and the statistical validity of the resulting Sharpe distribution unresolved. Researchers should treat the description as an introductory framing, not a complete CPCV specification.
Key ideas
- CPCV partitions time-series observations while using purging and embargoing to reduce leakage.
- Combining folds into multiple paths can yield a distribution of backtest Sharpe ratios.
- The number of input features does not by itself determine whether predictions are univariate or multivariate.
- The document poses, but does not resolve, how CPCV should be applied to multivariate series or outputs.
Tags
Full text
# Multivariate combinatorial purged cross-validation # Multivariate combinatorial purged cross-validation Combinatorial purged cross-validation (CPCV) is a technique for backtesting strategies while purging and embargoing observations in a time series. CPCV improves upon classical k-fold and walk-forward cross-validation because it has an added layer of generating paths that each possesses a unique Sharpe ratio, allowing a strategy's Sharpe ratio distribution to be derived, whereas K-fold and walk forward only provided one Sharpe ratio. My understanding is that cross-validation is a process that generates univariate dataset/time series predictions (but of course the input for fitting can be multivariate data). Given that CPCV not only groups based on folds, but also groups to produce multiple backtest paths, does this still make it a univariate predictor? If so, how can CPCV be extended to generate multivariate predictions? for example, maybe a matrix whose columns are individual prediction vectors But even without an answer to extending it to multivariate format, I just want to understand what CPCV means for multivariate time series data, input and output, given its extra layer of paths to predict.
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.