Skip to content
All library documents

Why Correlation Matrices Are Difficult to Estimate with Limited Data

Article Quant Q&A · Author: user28853

Summary

The document raises the problem of estimating stock return correlation matrices when the number of assets is large relative to the available history. It asks about implementing the unbiased rotationally invariant estimator described by Bun, Bouchaud, and Potters, and whether that method improves on the ordinary sample correlation estimator.

The author reports that an initial Python implementation showed no meaningful advantage over the sample estimator, despite repeated checks for coding errors. No implementation details, experimental data, or comparison results are supplied, and the post does not resolve whether the result reflects a problem in the implementation, a limitation of the method, or a mistake in the paper. It is therefore a useful prompt about finite-sample estimation challenges, but not evidence that one estimator is superior.

Key ideas

  • Sample correlation estimates can be difficult when the asset count is large relative to the return history.
  • Rotationally invariant estimation is proposed as a way to address correlation matrix estimation issues.
  • An initial implementation did not show a clear improvement over the sample estimator.
  • The post provides no data or implementation details sufficient to evaluate that finding.

Tags

Full text
# Cleaning correlation matrix, Bun Bouchaud Potters (2016) method


# Cleaning correlation matrix, Bun Bouchaud Potters (2016) method












Stock returns correlation matrices are notoriously hard to estimate, especially when the number of assets $N$ is large with respect to the size of the readily available historical returns $T$. Many methods exist to ease the estimation (and work-around the problem's ill-posedness).

I wondered if anyone tried to implement the unbiased version of the Rotational Invariant correlation matrix Estimator (RIE) described in this paper.

I've written some python code to test it and it does not seem to offer a significant advantage over the standard sample estimator of the correlation matrix.

I've already checked my code several time so I'd be interested to know if someone managed to implement it adequately (I don't need the code, just to know whether there is a typo in the paper for instance).

Thanks a lot already.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.