Feature Selection for Machine Learning Stock Strategies
Summary
This research summary explains feature selection as a way to reduce a model’s input set before training. It outlines filter, wrapper, and embedded approaches, with filter criteria including variance or entropy, statistical measures such as F-scores and mutual information, and importance scores from other models. The stated aims are lower computation costs, reduced overfitting, and easier interpretation.
The study applies selection methods to Chinese A-share, multi-factor stock strategies using logistic regression and XGBoost models, with industry and market-cap neutrality and large-cap and mid-cap benchmarks. It reports that F-score and mutual-information methods improved backtest performance for some model and forecast-horizon combinations. For a set of 70 pre-screened factors, predictive AUC generally improved as more features were selected before declining or flattening; about 50 features worked best for two of the models. These are study-specific findings: the source provides only an abstract, no detailed backtest metrics, and notes that selection may help more when the starting feature set contains ineffective factors.
Key ideas
- Feature selection chooses a subset of model inputs and can reduce computation, overfitting, and complexity.
- Common approaches include filter, wrapper, and embedded methods.
- The study reports gains from F-score and mutual-information selection for some Chinese stock models.
- Adding more selected features did not always improve AUC, and the preferred count varied by model.
- Feature selection adds no new information, and results may differ when the starting factors or models change.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.