Building Dependence and Distance Matrices from Financial Features
Summary
This code utility builds pairwise dependence matrices from columns in a feature DataFrame. It supports information-based measures, distance correlation, rank correlation, GPR and GNPR distances, and optimal-transport dependence. Parameters let users configure discretization, normalization, estimators, dependence targets, and selected copula settings. The resulting matrix is symmetric, with the information-variation output inverted so that larger values represent greater similarity.
A second function transforms a dependence or correlation matrix into an angular, squared-angular, or absolute-angular distance matrix, filling missing results with zero. These matrices can support analysis of relationships among assets or features, including clustering workflows. The document provides implementation rather than empirical evidence: it gives no comparison of methods, financial application, or validation results. Users must choose measures and parameters suited to their data, and should verify the matrix conventions before using the outputs in portfolio or trading decisions.
Key ideas
- The utility computes pairwise dependence across DataFrame columns using several statistical methods.
- Its parameters control discretization, normalization, estimators, and optimal-transport dependence settings.
- The dependence matrix is made symmetric, and information-variation scores are inverted to represent similarity.
- Angular transformations convert dependence values into distance matrices for downstream analysis.
- The code offers no empirical comparison or evidence that a particular method suits a given trading use.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.