Random Forest Stock Selection with Factor Ranking
Summary
This article explains random forests through CART decision trees, Gini impurity, bootstrap sampling, and aggregation. Trees are trained on resampled data, while their predictions are combined by voting for classification or averaging for regression. It also outlines controls such as tree count, depth, feature sampling, and minimum leaf size, and describes Gini-based feature importance and common regression and classification metrics.
The stock selection example uses 18 factors to predict each A-share stock’s return over the next five trading days. After missing-value handling, a model trained on 2010–2017 data ranks stocks in a 2017–2019 backtest; it buys the top five daily, holds positions for at least five days, and allocates more capital to higher-ranked stocks within a stated exposure cap. The article reports weaker performance in 2017–2018 and says results improved when market style was more consistent. It gives no detailed performance figures here, and notes that model fit and error still need improvement; the example is illustrative rather than conclusive.
Key ideas
- Random forests combine decision trees trained on bootstrap samples to reduce variance and aggregate their predictions.
- CART splits are selected to reduce Gini impurity, and tree depth and leaf size affect fit and generalization.
- The stock selection example predicts five-day returns from 18 factors and ranks A-share stocks for portfolio entry.
- The backtest reports uneven performance across market conditions and acknowledges room to improve model accuracy.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.