Cross-Validation for Machine Learning Trading Signals
Summary
This tutorial demonstrates evaluating a decision-tree model that predicts whether the next adjusted close of Apple stock will rise or fall. It constructs features from the open-close relationship, daily range, and rolling return volatility and average return. The example first scores the model on the same data used for training, then uses a chronological holdout split and four-fold cross-validation to estimate performance across several subsets. It also aggregates confusion matrices and calculates the mean and variability of fold accuracy.
The in-sample result is perfect, while the holdout and fold scores are close to chance, illustrating how training-set accuracy can overstate predictive ability. However, the described K-fold procedure trains on observations from later periods when testing earlier folds, so it does not preserve trading chronology and can introduce look-ahead bias. The example uses one equity and one classifier, and reports classification accuracy rather than trading returns, costs, or risk. Its results therefore illustrate evaluation pitfalls, not proof that the model is useful for trading.
Key ideas
- Scoring a model on its training data can greatly overstate its ability to predict unseen observations.
- The example uses price-derived features to classify the next session’s direction.
- A chronological holdout and fold-based evaluation produce results near chance in the reported example.
- Ordinary K-fold splits can train on future observations when testing earlier periods in time-series data.
- Classification accuracy alone does not show whether a trading model is profitable after costs and risk.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.