Skip to content
All library documents

Decision Trees: Recursive Splitting, Split Criteria, and Ensembles

Article Cryptohopper blog

Summary

The document introduces decision trees for classification and regression. It explains how a tree recursively partitions observations using feature tests, with branches representing outcomes and leaves providing class predictions or numeric estimates. Tree growth stops when limits such as maximum depth, minimum subset size, or sufficient node purity are reached.

It outlines information gain for ID3, gain ratio for C4.5, and Gini impurity for CART as ways to choose splits, then places random forests and boosting methods such as gradient-boosted trees in the broader family of tree models. A small synthetic binary-classification example illustrates fitting and visualizing a shallow tree. The discussion is instructional rather than an evaluation: it provides no trading results, market dataset, or out-of-sample evidence, and its claims about accuracy and generalization are not supported by comparative measurements.

Key ideas

  • A decision tree recursively splits data using feature tests and produces predictions at its leaves.
  • Information gain, gain ratio, and Gini impurity are alternative criteria for choosing splits.
  • Depth, subset size, and node purity can serve as stopping conditions.
  • Random forests combine trees in parallel, while boosting methods build tree ensembles sequentially.
  • The example demonstrates tree fitting and visualization on synthetic classification data, not trading performance.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.