Cross-Sectional Ensemble Models for Chinese Equity Selection
Summary
This report proposes a cross-sectional stock-selection framework that combines choices of factor subspaces, model families, and ensemble rules. Its base learners include boosted trees, ExtraTrees, random forests, Ridge regression, and Lasso regression, combined through linear weighting. The report compares selected feature spaces with the full factor set and evaluates results using excess return, Sharpe and information ratios, long-short performance, drawdown, and Calmar measures.
It reports that selected industry-related subspaces outperformed the full feature space, and that an ensemble weighting based on the return of the top in-sample selection group beat weighting based on R-squared. The ensemble also reportedly exceeded individual machine-learning and linear models in several metrics. Finally, the authors define a nonlinear-effect factor from the difference between ensemble and linear-model predictions and report a significant sorted portfolio result. These are historical research claims from the supplied summary; it gives limited detail on validation design, transaction costs, and protection against in-sample selection bias, so the reported performance should not be treated as independently verified.
Key ideas
- The framework combines factor subspaces, multiple model families, and a specified weighting rule.
- Selected feature subspaces reportedly performed better than using the entire factor set.
- The report says weighting by top-group portfolio returns outperformed R-squared-based weighting.
- An ensemble of tree and linear models reportedly improved several metrics over individual models.
- Prediction differences between the ensemble and a linear model were used to define a nonlinear-effect factor.
- Validation and transaction-cost details are limited in the supplied summary.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.