Comparing Machine Learning Models for Equity Factor Selection
Summary
This research summary compares five classifiers for predicting stock returns from historical equity factors: multinomial logistic regression, linear support vector machines, random forests, XGBoost, and deep neural networks. It distinguishes the linear models from the nonlinear methods and notes that random forests and XGBoost both use decision trees but combine them differently. The reported empirical results say all five approaches generated excess returns with broadly similar return curves, and that model scores and information coefficients were correlated, especially for the two linear classifiers.
Results varied with the training sample frequency. Daily samples favored the deep neural network, while the less frequent sample setup favored XGBoost; the daily setup performed better overall. The neural network required substantially longer training, while XGBoost trained in a time closer to the linear models. The summary also reports different style-factor exposures across models. It does not provide the underlying study’s full methodology or detailed performance figures, and warns that changing market structure or crowded trading can undermine a strategy.
Key ideas
- The study compares five linear and nonlinear classifiers for predicting stock returns from equity factors.
- All five models reportedly produced excess returns, with correlated scores and similar return curves.
- Daily training samples favored the deep neural network, while XGBoost led under less frequent sampling.
- The neural network had higher training costs, and the models differed in style-factor exposure.
- The reported results may not persist if market structure changes or similar strategies become crowded.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.