Integrating Bayesian Hyperparameter Optimization into a Trading ML Pipeline
Summary
The article explains how to integrate an Optuna Bayesian hyperparameter search into an existing machine-learning training pipeline while retaining a scikit-learn search path. In the Optuna path, model parameters, sample-weight scheme, decay, and linearity are optimized together; a Hyperband pruner can stop weak trials early, and SQLite storage allows a study to resume after interruption. The pipeline also detects primary versus meta-labeling models from the event data and routes preprocessing and feature handling accordingly. It describes additional production details, including retaining fitted column-dropping steps inside the final model, optional sequential bagging, cached preprocessing, monitoring and visual reports, and a companion bid/ask pipeline. The document is an implementation blueprint rather than a performance study: it supplies architecture and code examples but no comparative trading results. It cautions that candidate features must first be vetted for leakage, redundancy, and genuine importance, and that optimizing a flawed feature set can produce misleading in-sample performance.
Key ideas
- Optuna jointly searches model hyperparameters and sample-weight choices in the integrated training path.
- Pruning and persistent SQLite studies reduce wasted computation and allow interrupted searches to resume.
- The pipeline detects primary and meta-labeling tasks from whether event data contains a side field.
- Embedding the fitted preprocessor in the saved model helps preserve consistent inference-time feature selection.
- Feature validation for leakage and redundancy should precede hyperparameter optimization.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.