Decision Trees: Recursive Splitting, Pruning, and Classification
Summary
The document explains decision trees for regression and classification, describing them as models that divide feature space into rectangular regions using axis-aligned splits. For regression, each region predicts the mean response of its training observations; for classification, it predicts the most common class. Recursive binary splitting greedily chooses a split that reduces residual sum of squares, repeating within resulting regions until a stopping condition is reached.
Because an unconstrained tree can overfit and small data changes can produce unstable trees, the article describes growing a tree and then pruning it with a cost-complexity penalty. Classification splits can use hit rate, Gini impurity, or cross-entropy, with the latter two described as more commonly used. The discussion gives conceptual examples and formulas but no empirical trading results. It notes that individual trees may trail other supervised methods in predictive accuracy, while bagging, random forests, and boosting can make tree models more competitive. Quantitative finance applications mentioned include forecasting prices, directions, and liquidity.
Key ideas
- Decision trees partition feature space into nonoverlapping regions using axis-aligned splits.
- Regression trees predict each region's mean response, while classification trees predict its most common class.
- Greedy recursive binary splitting selects local splits that reduce residual sum of squares for regression.
- Pruning controls complexity and overfitting, while classification trees can use Gini impurity or cross-entropy.
- Individual trees can be unstable, but ensembles such as random forests and boosting may improve predictive performance.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.