Skip to content
All library documents

Feature Importance Methods for Correlated Trading Signals

Article MQL5 articles

Summary

This article compares feature importance methods for tree based trading models, with particular attention to correlated or duplicated predictors. Mean decrease impurity divides credit among interchangeable features, while mean decrease accuracy can understate them when a substitute remains available during permutation. Clustered variants group dependent features and score the group together; single feature importance instead evaluates each predictor in isolation, at the cost of missing joint effects.

The demonstration uses a synthetic market with one latent predictive signal represented by six transformations, a separate momentum feature, and noise features. In this constructed setting, ordinary impurity rankings split credit across copies, whereas clustered importance assigns the signal cluster a shared score. The article also explains cleaning a correlation matrix with eigenvalue shrinkage and adjusting the effective sample size for autocorrelation before clustering. These results illustrate known behavior in a controlled experiment; they do not establish which method best predicts live markets, and the outcome depends on the quality of the feature clusters and validation design.

Key ideas

  • Impurity importance can dilute a signal's apparent value when several columns encode the same information.
  • Permuting one feature may understate its importance if correlated substitutes remain intact.
  • Clustered importance scores dependent features jointly and gives cluster members a shared score.
  • Single feature importance avoids substitution but cannot capture features that work only in combination.
  • The synthetic experiment illustrates method behavior rather than proving live trading performance.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.