Comparing Machine Learning Models for Cross-Sectional Alpha Prediction
Summary
This research summary compares 17 machine learning approaches with conventional linear methods for predicting stock excess returns from 51 alpha factors. It describes splitting the prediction into two parts: an AR(1) forecast for return dispersion and a model-based forecast of cross-sectional standardized returns using the factors. Models are trained on monthly cross-sections. The reported historical evaluation spans 2009–2018 and uses out-of-sample R-squared and the Diebold-Mariano test to compare predictive accuracy.
The summary reports that principal component analysis and Elastic Net regularization improve linear forecasts, while nonlinear methods—especially gradient-boosted trees and random forests—perform best among the tested models. Averaging model forecasts also outperforms individual models. Predictive accuracy does not necessarily translate into portfolio returns when stock weights are constrained. The models favor technical factors and can lead to high turnover, so factor selection and transaction cost control matter. These findings come from the described historical study; model failure and extreme market conditions remain risks.
Key ideas
- The study predicts return dispersion with an AR(1) model and standardized cross-sectional returns with alpha factors.
- It evaluates 17 models using 51 factors and reports out-of-sample tests over 2009–2018.
- PCA and Elastic Net improve linear forecasts, while tree-based nonlinear models rank strongly in the reported comparison.
- Averaging forecasts from multiple models is reported to improve predictive accuracy.
- Forecast accuracy can diverge from portfolio performance because of weighting constraints and trading costs.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.