Skip to content
All library documents

Using CPCV for Hyperparameter Selection and Unbiased Backtesting

Article Quant Q&A · Author: June

Summary

The document raises two questions about using combinatorial purged cross-validation (CPCV) for a binary trading model: whether to select hyperparameters by averaging metrics across splits or across reconstructed paths, and whether evaluating those parameters on the same splits creates leakage. It describes a setup with six months of data, fifteen splits, and five paths, followed by fitting the selected model on all available observations.

The text does not answer the questions or provide performance results. Its value is in highlighting the distinction between tuning and estimating out-of-sample performance: reusing folds that influenced parameter selection can make reported results optimistic. It also points toward the need to define how fold metrics are aggregated and to preserve data separation for a final evaluation. Because the document is a question rather than a resolved methodological account, it does not establish a specific CPCV procedure, explain purging or embargo choices, or show that any strategy is profitable.

Key ideas

  • The document asks whether hyperparameters should be selected using average scores across CPCV splits or paths.
  • Using the same folds for parameter selection and performance reporting can bias the backtest upward.
  • A final model fit on all available data does not by itself provide an independent estimate of performance.
  • The text raises methodological questions but does not give a definitive CPCV workflow or empirical evidence.

Tags

Full text
# Proper Use of CPCV for Hyperparameter Tuning and Backtesting in a Trading Strategy


# Proper Use of CPCV for Hyperparameter Tuning and Backtesting in a Trading Strategy












I'm working on a binary classification model for a month-end trading strategy with 6 months of data. Initially, I split the data by using the last month for evaluation and backtesting, but this left me with only one month for evaluation, which i feel isn't sufficient.

To address this, I decided to use Combinatorial Purged Cross-Validation, which provides 15 different splits and 5 different paths for hyperparameter tuning and backtesting.

I used CPCV with 15 splits to determine the optimal hyperparameters. After tuning, I trained the final model on the entire dataset using these hyperparameters. For backtesting, I'm considering using these CPCV-tuned hyperparameters on the same 15 splits.

Concerns:

- When tuning hyperparameters, am i supposed to maximize the average evaluation metric across all 15 splits, or should I focus on the average of the 5 different paths?

- If I use the obtained best parameters within the same 15 splits for backtesting, am I introducing information leakage?

I feel like I might be misinterpreting CPCV. Could someone clarify the correct approach for both hyperparameter tuning and backtesting to avoid information leakage?

Thank you!

## Answer by Yagiz Temizel (score 0)

https://quant.stackexchange.com/a/85886

Short answer to both:

1) Split-average vs. path-average for hyperparameter selection

Use the reconstructed paths, not a flat average across the 15 raw splits. Each of the 5 paths is a full, temporally-ordered OOS equity curve stitched together from the test folds that happen to cover every period exactly once for that path - it behaves like a realistic single backtest, so you can compute Sharpe/drawdown/etc. on it the way you would on any normal equity curve.

A flat average across the 15 splits throws that structure away and can hide instability: a hyperparameter set that does great on half the splits and terrible on the other half can still look fine on average, even though no single realistic backtest sequence would ever have behaved that way. Look at the distribution of your metric across the 5 paths (mean/median and dispersion) and prefer hyperparameters that are consistently decent across paths over ones that are only good on average.

2) Yes, this is leakage - and it's a well-known failure mode

Using the best hyperparameters (chosen by optimizing performance across the same 15 CPCV splits) and then reporting "backtest" performance on those same splits is selection-induced overfitting, even though CPCV's purging/embargo already prevented the temporal leakage within each individual split. The leakage here isn't look-ahead bias in time, it's the standard "I tried N configurations and reported the best one's score as if it were unbiased" problem - the same thing k-fold CV + hyperparameter search has if you don't hold out a final test set.

This exact failure mode is what Bailey, Borwein, López de Prado & Zhu's Probability of Backtest Overfitting (PBO, via CSCV) was built to quantify as a follow-up to CPCV: given a set of trial configurations evaluated the way you're describing, PBO estimates how likely it is that the apparent "best" one is only best due to the selection process itself, not genuine skill. With only 6 months of data you don't have much room for a truly untouched final holdout, so I'd treat PBO (or at least the Deflated Sharpe Ratio, which corrects the winning config's Sharpe for the number of configurations you tried) as a mandatory companion number to whatever Sharpe you report - not a nice-to-have.

If it's useful, I put together a small open-source implementation of CPCV + PBO + Deflated Sharpe in Python (couldn't find an actively maintained one at the time - mlfinlab went closed-source a while back): https://github.com/HaasEnjoyer/quanttro - running your 15 trial results through `probability_of_backtest_overfitting` would give you a direct number for exactly the risk you're worried about in (2).

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.