Skip to content
All library documents

Estimating Mutual Information and Variation of Information from Binned Data

Code Stratmill research code

Summary

This code presents discretized mutual information (MI) and variation of information (VI) measures for comparing two data series. When the user does not supply a bin count, it estimates one from the observation count and, for the bivariate case, the correlation. MI can be calculated directly from binned observations or from ranked data transformed to approximate copula variables; a third estimator calculates MI as negative copula entropy. Optional normalization scales MI by the smaller marginal entropy, while normalized VI divides by a joint-entropy expression.

The document gives implementation details and points to research notes and an explanation of copula-based estimation, but it supplies no empirical trading results or validation. The measures can describe dependence beyond linear correlation, yet their values depend on discretization and estimator choice. The code also uses a small probability offset for the copula-entropy calculation, and its normalization choices should be checked before comparing results across datasets or implementations.

Key ideas

  • The default bin count is estimated from sample size and, in the bivariate case, the correlation coefficient.
  • Mutual information is computed from a contingency table of discretized observations.
  • Two alternative MI estimators rank-transform the inputs to approximate copula variables.
  • Variation of information combines marginal entropies and mutual information, with an optional normalization.
  • The code offers no trading tests, so estimator and binning sensitivity need independent evaluation.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.