Skip to content
All library documents

Decision Trees: Splits, Impurity Measures, and Practical Limitations

Article MQL5 articles

Summary

This tutorial refreshes the structure and implementation of decision trees for classification and regression. It explains internal decision nodes, leaf predictions, feature thresholds, and recursive tree construction, then contrasts CART, ID3, and C4.5. The main implementation focus is selecting splits by information gain, calculated as the parent impurity minus the weighted impurity of child datasets. Entropy and Gini impurity are presented as alternative measures for classification, with candidate thresholds drawn from feature values.

The article also lists common limitations: trees can be unstable, favor dominant classes or features with many categories, settle for locally good splits, and respond to noise. It describes these concepts through code fragments and a fruit classification example, but provides no quantitative evaluation of predictive performance. Although the discussion frames the material as preparation for random forests, it does not explain an ensemble method or a trading strategy; its value is primarily as a machine learning implementation overview.

Key ideas

  • A decision tree routes observations through feature tests until a leaf supplies a class or value prediction.
  • Information gain compares parent impurity with the weighted impurity of the resulting child datasets.
  • Entropy and Gini impurity are alternative criteria for evaluating classification splits.
  • Decision trees can be unstable and may favor dominant classes, noisy patterns, or features with many categories.
  • The tutorial presents implementation concepts but no trading-specific application or measured model results.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.