Skip to content
All library documents

Random Forest for Chinese A-Share Stock Selection

Article BigQuant

Summary

This report describes a monthly stock selection model that uses a random forest, an ensemble of decision trees trained with bagging, to estimate each stock’s probability of rising in the next period. Its workflow covers feature and label preparation, preprocessing, training, cross-validation, and out-of-sample testing. The authors use seven rolling annual train-and-test stages and assess predictions with accuracy and AUC, alongside strategy returns, information ratios, and drawdowns.

The report compares industry-neutral selections from the CSI 300, CSI 500, and the full A-share universe. It reports stronger excess returns and information ratios than linear regression in most cases, but weaker drawdown control. Market capitalization and reversal receive high feature-importance scores. The authors caution that tree models can be sensitive to noise and changing market conditions: results were strong during the small-cap-led period from 2011 to 2016, while the model struggled from 2017 onward. These historical findings do not establish that the approach will generalize to other regimes.

Key ideas

  • The model uses a bagged ensemble of decision trees to estimate next-period stock appreciation probabilities.
  • Its process includes feature preparation, training, cross-validation, and out-of-sample evaluation in rolling annual stages.
  • The reported industry-neutral strategies generally outperform linear regression on excess return and information ratio, but not on drawdown control.
  • Market capitalization and reversal are among the most influential factors in the model.
  • The authors warn that tree-based predictions may weaken when market conditions shift.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.