Financial Hyperparameter Search with Optuna and Purged Cross-Validation
Summary
The article presents an Optuna workflow for tuning financial machine learning models while preserving a financial data contract. It describes Bayesian parameter search with a TPE sampler, fold-by-fold evaluation using PurgedKFold, separate fitting and scoring weights, and persistent study storage. It also explains how intermediate fold scores can support pruning, and how study results can be converted into a scikit-learn-style results table for later analysis.
The discussion emphasizes a boundary between tuning statistical models on labeled data and optimizing trading rules on historical equity curves: efficient search may overfit a noisy strategy objective. It details Hyperband’s fold allocation and cautions against adding custom pruning rules that interfere with its bracket state; financial thresholds should be applied in the objective before invoking the pruner. The document is primarily an implementation blueprint, not evidence of improved live performance. Its claims about compute savings and fairer comparison depend on correct purging, weighting, and validation design, and require out-of-sample assessment.
Key ideas
- Use PurgedKFold so overlapping financial labels do not leak across training and validation folds.
- Optuna’s TPE sampler uses completed trials to guide later hyperparameter choices.
- Report scores after each fold to make early pruning possible and reduce unnecessary model fits.
- Keep strategy optimization on financial returns separate from model tuning on labeled-data statistics.
- Persist studies and convert their results for compatibility with existing scikit-learn analysis.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.