Covariance Estimation: Rank, Sample Size, and Eigenvalue Accuracy
Summary
The document examines how many observations are needed to estimate a covariance matrix for portfolio analysis. It distinguishes the sample count needed for a matrix to be full rank from the larger amount of data that may be needed to estimate its entries and eigenvalue distribution reliably. A matrix with N assets has N(N+1)/2 distinct covariance terms, while full rank can be achieved with fewer observations under suitable conditions.
The answers discuss random matrix theory, which predicts slow convergence of eigenvalue distributions when the number of observations is not large relative to the number of assets. They also caution that financial returns may not be independent and identically distributed, reducing the effective sample size. An example translates a stated observation rule into years of daily data for a 50 asset matrix, while another answer questions that rule under IID assumptions and points to a published covariance estimation study. These are practical rules and illustrations, not a universal guarantee; dependence and conditioning matter for inversion and optimization.
Key ideas
- A covariance matrix for N assets has N(N+1)/2 distinct entries.
- Full rank and statistically reliable covariance estimates require different sample-size considerations.
- Random matrix theory links eigenvalue accuracy to the ratio of observations to assets.
- Serial or cross-asset dependence can reduce the effective number of independent observations.
- Covariance quality matters for matrix inversion and portfolio optimization.
Tags
Full text
# Number of Observations for Non-Singular Covariance Matrix Estimation
# Number of Observations for Non-Singular Covariance Matrix Estimation
Marcos López de Prado writes the following in his book Advances in Financial Machine Learning:
> In general, we need at least `\frac{1}{2} N (N+1)` independent and identically distributed (IID) observations in order to estimate a covariance matrix of size N that is not singular. For example, estimating an invertible covariance matrix of size 50 requires, at the very least, 5 years of daily IID data.
What is the reasoning for that number of observations? Where can I find some sources related to this that I can cite?
## Answer by Michael Isichenko (score 2, accepted)
https://quant.stackexchange.com/a/67882
The covariance matrix of $N$ stocks (or whatever) consists of $N(N+1)/2$ distinct elements, so, to statistically measure these elements reasonably well, your number of independent observations $ND$ ($D$ being the number of days) should be well over $O(N^2)$, or $D\gg N$. This requirement is more stringent than the covariance $C_{ij}=\sum_dR_{di}R_{dj}$ being full rank. The latter needs just $D=N$ days to accumulate.
But the catch is elsewhere: Random matrix theory (RMT) indicates that the distribution of eigenvalues, which are relevant to the covariance condition number and inversion/solving/optimization tasks, converges to the "true" distribution very slowly, with the error decreasing only as $\sqrt{N/D}$. De Prado book actually discusses RMT aspects as well. If the observations are not exactly independent, the required number of days is further increased: one can introduce a statistic for the effective number of independent observations.
## Answer by Bob Jansen (score 4)
https://quant.stackexchange.com/a/66600
Let $f(N) = \frac{1}{2} N (N + 1)$ then $f(50) = 1275$. A year has approximately 255 trading days. So you need at least 1275 / 255 = 5 years.
I believe the rule above is used in practice but I think the text is not quite correct (which surprises me, maybe I should have a ☕). If the returns are IID, 51 observations ought to be enough, see the proof in this answer or in "Improved estimation of the covariance matrix of stock returns with an application to portfolio selection" (Ledoit and Wolf, 2003) for something citeable but short.
However, empirically, daily returns of multiple stocks are not IID and it is helpful to have more data, a more advanced estimation covariance scheme or both. To illustrate, a plot of condition numbers
```
set.seed(42)
sampleSizes <- c(50:1500)
conditionNumbers <- sapply(
sampleSizes,
function(x) {
kappa(
cov(matrix(rnorm(n = 50 * x, mean = 0, sd = 0.2), ncol = 50)),
exact = TRUE
)
}
)
plot(sampleSizes, log10(conditionNumbers))
```
kappa calculates the condition number.Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.