Skip to content
All library documents

Avoiding Data Leakage and Overfitting in AI Stock Selection

Article BigQuant

Summary

This article collects practical lessons from building supervised machine learning strategies for stock selection. It recommends keeping training and test data separate and splitting financial time series chronologically, since random splits can mix market regimes and undermine out-of-sample evaluation. It also argues for a broad training universe and comparable, economically informed features, such as normalizing prices or valuing a company relative to its industry. Labels should match the intended forecast horizon: the article links slower moving valuation features to longer horizons and faster price or flow features to shorter ones.

The author cautions that adding more factors does not guarantee better predictions, suggesting feature contribution analysis and experiments to select a smaller set. Platform notes cover supplying enough historical lookback for long-window features and understanding how a volume limit can defer fills across days. These are practitioner observations, not controlled evidence; the page supplies no systematic performance results, and its platform-specific recommendations may depend on implementation settings.

Key ideas

  • Keep training and test periods separate to avoid evaluating on data the model has already seen.
  • Split market time series chronologically so tests reflect future periods and changing market conditions.
  • Use a sufficiently broad stock sample and features normalized for cross-stock comparability.
  • Align the supervised label horizon with the expected holding period and feature behavior.
  • Select factors carefully because adding features can increase overfitting rather than improve accuracy.
  • Provide enough lookback history for rolling features and account for volume limits that can delay fills.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.