Skip to content
All library documents

Choosing Training Windows for AI Quantitative Strategies

Article SuperMind

Summary

The article examines how training-window length can affect an AI stock-selection model. It compares longer historical samples with shorter windows using manual rolling training and a held-out backtest period. In the reported experiments, changing the training span altered predicted returns, Sharpe ratios, factor weights, and NDCG. Longer samples produced more stable, longer evaluation periods but could overfit; short windows sometimes showed stronger returns or NDCG, yet could be underfit and less reliable out of sample. The author reports that adding years did not consistently improve model quality.

The main recommendation is to search for a suitable window for each model and dataset, rather than assume that more history is always better. The article uses NDCG and out-of-sample returns together to interpret fit, while cautioning that favorable returns can reflect luck, market conditions, stock-universe filters, or risk controls. It also argues that factor choices should fit the strategy and market regime. These observations come from the described experiments, not a general proof; suggested metric thresholds and date ranges should be treated as heuristics.

Key ideas

  • Training-window length can change model predictions, factor weights, and evaluation metrics.
  • Longer samples may improve backtest stability while also increasing overfitting risk.
  • Short samples can produce attractive results but may be underfit and unreliable out of sample.
  • Evaluate NDCG alongside training and prediction-period returns, while accounting for filters and market effects.
  • Choose training windows and factors in light of the model, data, and market regime.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.