Machine Learning Trading Research: Targets, Validation, and Common Pitfalls
Summary
This tutorial outlines a data-mining approach to machine-learning signals, contrasting it with strategies that begin from an explicit market inefficiency such as trend following or mean reversion. It recommends defining the prediction target and evaluation measure first, then assembling representative, clean historical data and features that would actually be available at prediction time. A worked example predicts a future stock–futures basis using a regression target based on subsequent basis observations, with prediction error and total P&L as evaluation criteria.
The discussion stresses that forecasts are not complete trading strategies: fees, spreads, available size, and stop rules affect realized results. In the example, transaction costs and price differences sharply reduce P&L, though the excerpt does not establish broader profitability. It warns against forward-looking leakage, overfitting, repeated model selection on the same data, and retraining too frequently. Training and test separation and out-of-sample checks are advised, while the article does not provide enough complete experiment detail to assess the model's robustness.
Key ideas
- Define the target variable and a suitable evaluation measure before training a model.
- Use only features available at the time each prediction would be made.
- Clean and representative data matters because missing values and biased samples distort results.
- Trading costs and execution constraints can erase apparent backtest profits.
- Use held-out data and guard against overfitting, data-mining bias, and frequent retraining.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.