Skip to content
All library documents

Feature Selection for Chinese A-Share Machine-Learning Strategies

Article BigQuant

Summary

This research applies feature selection to multi-factor stock selection in the full Chinese A-share universe, with the CSI 300 and CSI 500 as benchmarks. It constructs industry- and market-cap-neutral strategies and tests feature ranking methods with logistic regression and XGBoost models at different prediction horizons. The methods include statistical filters such as F-statistics and mutual information, alongside other filter, wrapper, and embedded selection approaches.

The reported tests find that F-statistic and mutual-information selection improve some model backtests. Model AUC generally rises and then falls as more features are retained for the shorter-horizon models; with the study’s 70 inputs, performance is best at roughly 50 selected features. Frequently selected inputs in one setup are mainly price-volume factors. The authors caution that results depend on the base learner and historical patterns may not persist. Since the starting factors were already screened as effective, the gains may be modest; selection can also add overfitting risk.

Key ideas

  • Feature selection can reduce model inputs, development time, and the risk of overfitting while improving interpretability.
  • F-statistics and mutual information improve some reported A-share model backtests.
  • For logistic regression and shorter-horizon XGBoost, AUC first improves and then declines as more features are retained.
  • In the study’s 70-feature set, roughly 50 selected features work best for some models.
  • The findings depend on the base learner and may not hold if market conditions change.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.