Decision Trees for Stock Selection: Splits, Pruning, and Limits
Summary
The article explains how classification trees make decisions by splitting observations according to features. It describes information gain and Gini impurity as criteria for choosing splits, and notes that information gain can favor features with many distinct values. It introduces gain ratio as one response to that bias. To control overfitting, it contrasts pre-pruning during tree construction with post-pruning after a full tree has been grown, using validation data to judge whether simplification helps.
For a stock-selection example, the method uses company financial indicators to predict whether each stock’s prior-month return was positive or negative, then selects stocks predicted to rise for the next rebalance. The article reports that this decision-tree approach performed worse than a random forest, but supplies no detailed performance statistics or evaluation design. The example is therefore illustrative; its usefulness depends on sound feature selection, validation, and safeguards against look-ahead bias and overfitting.
Key ideas
- A decision tree predicts classes by repeatedly splitting data on selected features.
- Information gain and Gini impurity favor splits that make resulting classes more homogeneous.
- Gain ratio can reduce information gain’s preference for features with many distinct values.
- Pre-pruning and post-pruning use generalization performance to limit tree complexity.
- The stock-selection example predicts next-period direction from financial indicators and reports weaker results than a random forest.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.