Interpreting XGBoost Stock-Picking Models with Explainability Methods
Summary
This report introduces six ways to interpret machine-learning models and applies them to an XGBoost stock-selection model. It discusses feature importance, individual conditional expectation (ICE), partial dependence plots (PDP), a surrogate decision tree (SDT), local interpretable model-agnostic explanations (LIME), and SHAP values. These methods examine feature influence at the model or individual-prediction level, approximate complex models with simpler ones, or estimate features’ marginal contributions.
The reported analysis trains and validates on data from 2013–2018 and tests on 2019. Feature importance and the surrogate tree indicate that price-and-volume factors rank above fundamental factors overall. PDP and SHAP suggest nonlinear use of features, especially those related to size, reversal, technical signals, and sentiment; they also reveal interactions and some factors with zero marginal contribution. The report cautions that predictive models capture associations rather than causal relationships, so explanations help assess model behavior but do not establish why returns occur or guarantee future performance.
Key ideas
- Feature importance, ICE, PDP, SDT, LIME, and SHAP offer different views of model behavior.
- The example applies these methods to an XGBoost stock-selection model.
- Price-and-volume factors rank above fundamental factors in the reported feature-importance analysis.
- The model uses some features nonlinearly, with interactions and zero-contribution features also reported.
- Model explanations describe learned associations and do not establish causal relationships.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.