Machine Learning Workflow and Ensemble Models for Factor Stock Selection
Summary
This document outlines a machine learning workflow built around feature engineering, model training, and model combination. Feature engineering includes creating, extracting, and selecting variables. It contrasts linear models, which capture linear relationships, with tree models that can represent nonlinear patterns, and neural networks trained with improved stochastic gradient methods. It also describes combining models whose predictions perform well and are not highly correlated.
A Chinese equities case combines gradient boosting, ExtraTrees, and a deep neural network by averaging their predicted scores to form a portfolio. The document reports that the ensemble modestly outperformed individual models on excess return, Sharpe ratio, information ratio, and monthly excess-return win rate, and separated portfolio groups more clearly. These figures are reported as results of the cited example, but the supplied text does not describe the sample, benchmark, costs, validation design, or robustness checks. They should therefore be treated as a case result rather than evidence that this ensemble will generalize to other markets or periods.
Key ideas
- The workflow consists of feature engineering, model training, and model combination.
- Linear models and tree models capture different kinds of relationships and can complement one another.
- Combining accurate models with low correlation may improve the consistency of predictions.
- The stock-selection example averages scores from gradient boosting, ExtraTrees, and a neural network to form portfolios.
- The reported advantage is specific to the example, and the supplied summary lacks details needed to assess out-of-sample robustness.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.