Skip to content
All library documents

Using Gini Impurity to Choose Decision Tree Splits for Market Data

Article QuantInsti blog

Summary

The article explains Gini impurity as a measure of class mixing in a decision tree node. It gives the formula one minus the sum of squared class probabilities, then describes evaluating a candidate feature by calculating each resulting branch’s impurity and taking the branch-size-weighted average. A lower weighted score indicates a cleaner split. It contrasts this measure with entropy and distinguishes decision-tree Gini impurity from the economic Gini coefficient for inequality.

A small example labels market observations by past trend, open interest, trading volume, and subsequent return. The article works through weighted impurity calculations for features and uses the lower-scoring feature to illustrate split selection. This is a teaching example, not evidence of a profitable trading model: it reports no out-of-sample results or full validation. Some explanations of the numerical range are imprecise; for two classes, the maximum impurity is 0.5, while for n classes the maximum is 1 minus 1/n. Any market use would require careful validation and risk controls.

Key ideas

  • Gini impurity is zero when a decision-tree node contains only one class.
  • A split’s score is the weighted average of the impurity in its child branches.
  • Decision trees favor features that produce lower weighted impurity.
  • The worked market example classifies returns using trend, open interest, and trading volume.
  • The example teaches split calculation but does not establish trading performance.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.