Skip to content
All library documents

Denoising and Detoning Feature Correlation Matrices Before Clustering

Article MQL5 articles

Summary

This article prepares a feature correlation matrix for clustering, a prerequisite for the clustered feature-importance analysis planned in the next part of its series. It separates two problems: finite-sample estimation noise, which spreads eigenvalues even when variables are uncorrelated, and a shared market mode, which can make distinct feature families appear alike. The proposed pipeline fits the Marcenko–Pastur noise distribution, treats eigenvalues above its fitted upper boundary as signal, replaces noise eigenvalues with a common residual value, and removes the dominant market component before clustering.

A practical issue is the sample size used in the fit. Because sequential FX bars are autocorrelated, raw bar counts can overstate independent information; the article estimates an effective sample size from average lag-one autocorrelation. It also flags silent implementation errors in the observation-to-feature ratio and KDE bandwidth convention. The discussion is methodological and tied to an example feature panel; the excerpt does not establish that the resulting clusters improve out-of-sample trading performance.

Key ideas

  • Finite samples create apparent correlations and eigenvalue spread that can mislead feature clustering.
  • The Marcenko–Pastur distribution provides a noise ceiling for classifying eigenvalues.
  • Denoising flattens eigenvalues treated as noise while retaining eigenvalues classified as signal.
  • Detoning removes a dominant shared market component that can obscure feature-family differences.
  • Serial correlation means raw bar counts may overstate the effective sample size used in the fit.
  • The article highlights ratio and bandwidth choices that can silently bias the noise estimate.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.