Skip to content
All library documents

Choosing Similarity and Distance Measures for Numeric Data, Distributions, and Sets

Article TradingView scripts

Summary

This Pine library collects functions for measuring distance or similarity between arrays and other representations. Its numeric-vector measures include squared difference, Euclidean, Manhattan, Minkowski, Chebyshev, correlation, cosine, Canberra, and error-based distances. Other sections cover comparisons involving strings, probability distributions, and sets, with measures such as edit distance, Hellinger, Jaccard, and chi-square. Most functions compare corresponding elements and check that input arrays have matching sizes.

The library is a catalog of computational tools rather than a trading strategy or empirical study. It gives definitions and implementations, but does not establish which measure works best for market data or show predictive performance. Measures can respond differently to scale, outliers, zero values, and data representation; users need to match the formula to the task and handle edge cases appropriately. In particular, an implementation’s output may be undefined or specially handled for zero denominators, and similarity scores should not be treated as interchangeable without checking their meaning and direction.

Key ideas

  • The library provides multiple ways to compare numeric vectors, including norm-based distances and correlation or cosine measures.
  • Additional functions address strings, probability distributions, and sets.
  • Most vector comparisons require equally sized, nonempty inputs.
  • Different measures have different sensitivities and output interpretations, so selection depends on the data and task.
  • The library supplies implementations but no evidence of trading effectiveness.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.