Feature Selection for Machine Learning Stock Selection
Summary
The document reviews feature selection as a preprocessing step for machine learning stock selection. It distinguishes filter, wrapper, and embedded approaches, and describes selection criteria such as variance, entropy, F statistics, mutual information, chi-square scores, regularized model coefficients, and tree-based importance. The purpose is to reduce computation, help limit overfitting, and make models easier to interpret.
The reported application uses Chinese A-share stocks with industry- and market-cap-neutral portfolios benchmarked against the CSI 300 and CSI 500. It says F-score and mutual-information methods improved backtest performance for logistic regression and XGBoost models. For some six-month prediction models, AUC first rose and then fell as more features were selected; the reported best subset was around 50 of the 70 tested factors. The document cautions that selection adds no new information and may offer only incremental benefit when input factors are already individually vetted. It gives no detailed performance figures or full validation protocol in the supplied text.
Key ideas
- Feature selection can reduce model development time, overfitting risk, and difficulty of interpretation.
- Filter, wrapper, and embedded methods differ in how they rank or choose input features.
- The study reports gains from F-score and mutual-information selection in Chinese equity stock models.
- For some models, adding features improved AUC only up to a point, after which performance declined.
- Feature selection may provide limited gains when the candidate factors have already been individually validated.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.