Skip to content
All library documents

Why Large Vendor Risk Models Can Still Be Useful Despite Covariance Error

Article Quant Q&A · Author: quant_zero

Summary

The document questions whether vendor risk models with hundreds or thousands of correlated factors can estimate covariance reliably when they have only hundreds to about a thousand observations. Exponential weighting can reduce the effective sample further, making the usual high-dimensional covariance estimation problem especially severe. The concern is that estimation error may distort the risk picture substantially, even before accounting for the fact that risk models rely on historical data.

The document does not present an answer or empirical evidence resolving the concern. Instead, it asks what features might make these models useful in practice despite the apparent mismatch between factor count and sample size. It leaves open whether vendor methods such as cleaning, shrinkage, or other adjustments explain their adoption, so it serves as a framing of the estimation problem rather than a guide to a specific model or a conclusion about model accuracy.

Key ideas

  • Covariance estimates can become unreliable when the number of dimensions approaches or exceeds the effective observation count.
  • Highly correlated factors and exponential weighting complicate the relationship between raw sample size and estimation quality.
  • The document raises skepticism about large vendor risk models but offers no solution or supporting performance evidence.
  • Historical risk estimates also inherit the limitation of relying on past market behavior.

Tags

Full text
# Why are thousand-ish-factor vendor risk models not extremely overfit and inaccurate?


# Why are thousand-ish-factor vendor risk models not extremely overfit and inaccurate?












Many vendor risk models have many hundreds, or even thousands of factors (many of which are highly correlated with each other). Underlying all these risk models is some sort of covariance matrix in which these factors are the features/dimensions generating this matrix, at least roughly speaking (all have some form of cleaning process or idiosyncratic deviation from the general process).

For fixed income, at least, the number of observations going into this covariance matrix tends to be on the order of a few hundred to a thousand, although many have some sort of an exponential weighting scheme to weight recent observations more strongly. This means that the effective number of observations can actually be quite small.

Even assuming ideal (normal) data the situation appears hopeless, because of the well known problems with estimating covariance matrices when the dimensionality is comparable to, let alone much larger than, the number of effective observations (see here and here, for classic takes). It’s seems a safe bet that the covariance matrices underlying most vendor risk models are therefore likely to be extremely inaccurate – not merely slightly wrong, but perhaps deeply wrong. Of course any risk model is necessarily backward looking and should be viewed with caution for this reason alone, but from a pure statistical perspective the problem is that even one’s look backward is likely to be wildly distorted.

Nonetheless, they are widely used and viewed as useful. Therefore I assume such a model has virtues that I must be unaware of. What are these virtues? What is the thoughtful, quantitative answer to my skeptical take? Have I overstated the case against these mega-factor models?

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.