Choosing Training Windows for AI Stock Selection Models
Summary
This article examines how training-window length affects an AI stock-selection model. It compares long windows, expanded year by year from 2005 to 2021, with shorter windows ranging from a month to a few years. The author recommends checking each window against a fixed validation period and monitoring returns, Sharpe ratio, factor weights, prediction behavior, and NDCG. In the reported experiments, performance peaked for a window around 2005–2016 and then declined as more years were added; NDCG stayed below 0.6 in the longer-window trials. A one-month example produced positive returns but low NDCG and a weakening return curve.
The article cautions that long windows may overfit while short windows may underfit, and that favorable backtests can also reflect market conditions, stock-pool effects, or later filters. It argues for finding a model-specific training range rather than assuming more data is better. Its conclusions are based on examples from a particular strategy setup, and the suggested NDCG thresholds are rules of thumb rather than broadly validated standards. It also links strategy logic to factor choice, noting that factor performance can vary with market style and rebalancing frequency.
Key ideas
- Training-window length can change a model’s returns, factor weights, and predictions.
- Adding historical data can stop helping and may worsen performance through overfitting.
- Short windows may produce attractive returns while leaving the model underfit and less reliable.
- Compare candidate windows on a fixed validation period and examine multiple diagnostics, not returns alone.
- Factor choice and strategy logic should fit each other, while their results remain sensitive to market conditions.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.