A Genetic Programming and Random Forest Framework for Quantitative Stock Selection
Summary
This research overview presents a workflow for building stock-selection signals from daily price and volume data. Genetic programming generates candidate factors, using rank information coefficients to target linear relationships or mutual information to find nonlinear ones. To seek information beyond existing signals, the process predicts residual returns; the authors also identify high-performance computing as a way to speed factor discovery.
A random forest then combines selected factors, with time-ordered cross-validation used to tune parameters and reduce overfitting. SHAP analysis is applied to explain how the fitted model uses its inputs. The reported tests use rolling factor discovery and a 20-trading-day rebalance horizon. After neutralization against industry, size, and several recent-return and trading-activity variables, the authors report positive rank IC and benchmark-relative results, including an improvement when adding the generated factor to a traditional-factor CSI 500 enhancement model. These are study-specific backtest findings; the summary does not establish live performance or robustness across markets and periods.
Key ideas
- Genetic programming can search price and volume data for candidate linear and nonlinear stock-selection factors.
- Predicting residual returns is proposed as a way to mine information incremental to existing factors.
- Random forests combine the discovered factors, while time-series cross-validation helps manage overfitting.
- SHAP values are used to inspect the model’s factor contributions.
- The document reports favorable historical tests, which do not by themselves establish future or live-market performance.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.