Skip to content
All library documents

Bayesian Estimation as a Practical Approach to Covariance Shrinkage

Article Quant Q&A · Author: Quartz

Summary

The document discusses covariance estimation through the lens of shrinkage and Bayesian methods. Its respondent argues that shrinkage estimators can be understood as Bayesian point estimates with different implicit prior distributions, and recommends Bayesian estimation when genuine prior information is available. Bayesian updating carries the previous posterior forward as the next prior, allowing new observations to be incorporated incrementally.

The response presents coherence and the ability to update sequentially as advantages, while identifying computation time as a practical constraint. It argues that a one-day data update may add little unique information relative to the accumulated data and prior, so a lagged update can sometimes produce nearly the same result. These are the respondent’s general claims and practical perspective, not a comparison of covariance estimators or supporting empirical results. It does not specify a covariance model, prior, dataset, or conditions under which the suggested update lag is reliable.

Key ideas

  • The response interprets shrinkage estimators as Bayesian estimates with differing implicit priors.
  • It recommends Bayesian estimation when meaningful prior information is available.
  • Sequential updating can use the previous posterior as the next prior.
  • Computation time is identified as a tradeoff for Bayesian approaches.
  • The claim that lagged updates lose little information is presented without empirical validation or specific model assumptions.

Tags

Full text
# Covariance estimation: shrinkage, random matrix theory, what else?


# Covariance estimation: shrinkage, random matrix theory, what else?












Shrinkage was much en-vogue before random matrix theory (RMT) took everybody's attention in covariance matrix estimation, however the latter also showed its limits. A plethora of other estimators has been presented, but I could not yet spot a golden standard. What is nowadays used most in practice (or what are you using), and why? Also, shrinkage came in different flavours, so I'd like to know which is the favourite among them.

Note that I'm not just asking for the statistical properties of different methods (this would be on Cross Validated in that case), but also their interplay with practical considerations here in the quant world, which might include even non-technical factors.

## Answer by Dave Harris (score 3)

https://quant.stackexchange.com/a/32499

I thought I would answer the question of "what am I using." All shrinkage estimators map to a Bayesian estimator that differs only in the prior distributions. In other words, you get a point estimate that is indistinguishable from a Bayesian estimate except that the calculation rule determines the prior distribution. Stein estimators for the Gaussian are nothing more than the ordinary estimator with an empirical prior distribution based off of the grand mean.

Since I usually have real prior information, and so does everybody else, I use Bayesian estimators as I get the shrinkage for free. Bayesian estimators have the nice property that they are "coherent," which means fair gambles can be placed on them, whereas this is not true for Frequentist estimators. In fact, I once did a talk on how you could always game someone using a Frequentist estimator in the right circumstances guaranteeing a sure win no matter how the universe came out.

The downside of a Bayesian estimator is computation time. This is less serious than it appears because Bayesian updating allows you to do the calculations one data point at a time. If you have calculated the solution for every data point up to today, then you can add today's data in and ignore all prior data, reducing the calculation size tremendously. You can do this because yesterday's "posterior density" becomes today's "prior density."

Speed is your enemy here, though. Speed requires compromises if you use a Bayesian method. Although Bayesian predictions are coherent and Frequentist predictions are not, you have to do some compromises or you will not run fast enough if time is of the essence. In particular, the impact of one day's data on the posterior is negligible. Unlike a Frequentist method, you do not really have to update in real time, you can use the data up to yesterday and get nearly identical results. This is because Bayesian methods use the unique information in the data that is not in the prior.

To understand this, imagine you have two data sets and a prior and you merge one data set with the prior to get a posterior density. When you merge the second set into this density, the only changes that will happen in the second posterior density will be from information that is unique to the second set that is neither in the first set nor the prior. Bayes rule ignores redundant information. In essence, this is saying that the unique information content of a single trade is negligible when compared to the joint information content of the set of all trades that have ever happened.

So you do not need to run it in real time unless you believe today is the one day that is unique in the history of trading. You can run with a lag because there is trivial information loss.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.