Skip to content
All library documents

Building and Evaluating Machine Learning Signals for Trading

Article FMZ digest · Author: 善

Summary

This tutorial lays out a supervised-learning workflow for developing a market-direction signal, using a stock–futures basis as an example target. It distinguishes model-based strategy design from data mining, then walks through defining the prediction target and evaluation metric, gathering and cleaning data, and separating training, validation, and test samples. Feature examples include basis changes, RSI, moving averages, bid-ask widths, and order-book volumes; the text emphasizes that feature quality may matter more than model complexity.

The workflow compares regression and classification choices, recommends starting with simple models, and discusses rolling validation and ensembles as possible ways to assess robustness. Its reported modeling metrics and P&L illustrations are qualified by the omission of trading costs in some results; a later example shows costs and spreads eroding much of the apparent performance. The article also warns about overfitting, look-ahead bias, and data-mining bias. A predictive model alone does not define entries, exits, sizing, or execution, and out-of-sample success is not guaranteed.

Key ideas

  • Define the prediction target and evaluation metric before choosing a model.
  • Separate training, validation, and test data to reduce overfitting and test-set contamination.
  • Engineer and examine features for predictive value, while treating normalization of time series carefully.
  • Use rolling validation and model ensembles to examine robustness across changing market conditions.
  • Account for fees, spreads, volume, and execution because predictive accuracy alone does not establish profitability.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.