Interpreting Random Forests with Importance, Tree Contributions, and PDPs
Summary
The article presents four ways to inspect a random forest: permutation-based feature importance, variation among individual tree predictions, per-observation prediction paths, and partial dependence plots (PDPs). For permutation importance, it compares a baseline score with scores after shuffling each feature; a larger error after shuffling suggests the model relies on that feature. The bulldozer price example illustrates the approach, though its reported importance values and ranking depend on the sample and scoring setup.
The article estimates prediction uncertainty from percentiles of individual tree predictions, then describes tree interpretation as a way to attribute a particular prediction to successive feature splits. PDPs vary one feature across values, average the model predictions, and plot those averages to show its modeled relationship with the target. These tools can make a model easier to inspect, but the article does not establish calibrated prediction intervals or causal effects: tree-to-tree spread is only a rough uncertainty proxy, and PDP averages can obscure feature dependencies. The examples are illustrative rather than a trading strategy or trading-system evaluation.
Key ideas
- Permutation importance measures how model score changes when a feature's values are shuffled.
- The author uses variation among individual tree predictions as a rough indicator of prediction uncertainty.
- Tree interpretation attributes an individual prediction to contributions along its decision path.
- Partial dependence plots average predictions while varying one feature to show its modeled relationship with the target.
- These interpretation tools do not by themselves establish causal effects or calibrated uncertainty.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.