Skip to content
All library documents

Genetic Programming and Random Forests for Price-Volume Stock Selection

Article BigQuant

Summary

This report describes a Chinese equity-selection process built from daily price and volume data. Genetic programming generates candidate factors: Rank IC is used to find factors with linear predictive relationships, while mutual information targets nonlinear relationships. To search for information beyond existing factors, the approach predicts residual returns rather than repeatedly orthogonalizing factors. A random forest then combines selected signals, with feature selection and time-ordered cross-validation used to help control overfitting. SHAP values are applied to interpret the model’s contributions.

The reported evaluation uses rolling factor discovery and a 20-trading-day rebalance horizon. After neutralization against industry, size, recent return, volatility, and turnover, the report gives a mean Rank IC of 8.87% and an IC information ratio of 1.16; its top quintile reports 9.65% annualized excess return and a 3.08 information ratio. Adding the composite factor to a CSI 500 enhancement model reportedly raises average annualized excess return by 1.38% and information ratio by 0.14. These are study results, not guarantees; the excerpt gives limited detail on costs, data-snooping controls, and out-of-sample robustness.

Key ideas

  • Genetic programming searches raw price and volume data for candidate stock-selection factors.
  • Rank IC and mutual information serve as fitness measures for linear and nonlinear factor discovery, respectively.
  • Predicting residual returns is proposed as a more efficient way to mine incremental information.
  • A random forest combines factors, with feature selection and time-ordered validation intended to limit overfitting.
  • SHAP analysis is used to interpret the model, and the report presents backtest metrics for a CSI 500 enhancement application.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.