Choosing Training Windows for AI Stock Selection Models
Summary
The article compares long and short training windows for an AI stock selection strategy and recommends evaluating each window against a fixed validation period. In its rolling experiments, extending the sample from 2005 to 2021 changed labels, factor weights, predictions, and backtest outcomes. Reported returns and Sharpe peaked around a 2005–2016 training window, while the NDCG measure stayed below 0.6 in those tests. A one-month example produced positive returns but had low NDCG and a curve that weakened over time.
The author argues that longer samples can make backtests more stable but may overfit, while short samples may produce sharper gains yet risk underfitting and fast decay. The best window depends on the model, data, and initialization. NDCG is presented as a diagnostic alongside train and prediction performance, with cautions that screening rules, timing, market regimes, and luck can explain apparent alpha. The observations come from one platform’s experiments and are not a general proof of optimal windows or reliable future performance.
Key ideas
- Compare multiple training windows against a fixed validation period.
- Longer histories may stabilize results while increasing overfitting risk.
- Short histories can show stronger gains but may underfit and lose effectiveness quickly.
- NDCG should be interpreted alongside out-of-sample returns and strategy adjustments.
- Factor effectiveness depends on market conditions and the strategy’s selection logic.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.