A Genetic Programming and Random Forest Pipeline for Equity Selection
Summary
This research summary describes an equity selection process built from daily price and volume data. Genetic programming searches for candidate factors, using RankIC to favor linear relationships or mutual information to seek nonlinear ones. To focus on incremental information, the method predicts residual returns instead of repeatedly orthogonalizing factors. A random forest then combines selected factors, with feature selection and time-ordered cross-validation used to help control overfitting. SHAP values are applied to explain the model’s decisions.
The report says the resulting composite factor was neutralized against industry, size, and several recent-return, volatility, and turnover characteristics. It reports positive ranking and portfolio results, and says adding the factor improved a CSI 500 enhancement portfolio’s reported excess return and information ratio. These are results from the described tests, not a guarantee of future performance. The summary does not provide enough detail to assess transaction costs, robustness across periods, or live trading outcomes.
Key ideas
- Genetic programming can search raw price and volume data for candidate stock selection factors.
- RankIC and mutual information provide different fitness criteria for finding linear and nonlinear relationships.
- Predicting residual returns is presented as a way to seek incremental factor information while avoiding repeated orthogonalization.
- Random forests combine factors, while feature selection and time-ordered cross-validation address overfitting risk.
- SHAP analysis is used to interpret how the model relies on its input factors.
- Reported backtest results indicate incremental performance, but do not establish live or future returns.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.