Choosing Training Windows for AI Quantitative Trading Models
Summary
This article examines how training-window length affects AI stock-selection models, comparing multi-year datasets with shorter windows and proposing rolling training and evaluation on a designated validation period. It reports that, in its experiments, longer training histories did not consistently improve returns or Sharpe ratios: performance rose to a peak and then fell as years were added. Feature weights and NDCG also changed, while reported NDCG remained below 0.6 in the long-window tests. A one-month example produced positive returns but low NDCG and a return curve that weakened over time.
The article argues that neither the longest nor shortest window is automatically best. Long windows may make results more stable while increasing overfitting risk; short windows may look stronger but can underfit and fail sooner. It treats NDCG and out-of-sample performance as complementary diagnostics, while noting that universe filters, timing, market regime, and initialization can distort apparent alpha. The examples are platform-specific observations, not general proof of a universally optimal window.
Key ideas
- Training-window length can change returns, risk-adjusted performance, factor weights, and ranking metrics.
- The article's experiments show performance peaking and then declining as more historical years are added.
- Long windows may improve stability while increasing overfitting risk, whereas short windows may be less reliable.
- Training quality should be judged with both ranking metrics and out-of-sample results.
- Market regime, stock-universe filters, and later strategy adjustments can make model performance appear stronger than the model's signal alone.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.