Why Sample Covariance Becomes Singular with Too Few Observations
Summary
The document gives a rank argument for why a sample covariance matrix cannot be positive definite when the number of observations is smaller than the number of variables. Each centered observation contributes an outer product of a vector with itself, a matrix with rank at most one. Adding these contributions limits the covariance matrix’s rank; because the observations are centered, the rank is at most the smaller of the variable count and the observation count minus one.
When there are fewer observations than variables, the resulting matrix is therefore singular. The answer also distinguishes singularity from failure of positive semidefiniteness: a sample covariance matrix remains positive semidefinite because it is a Gram matrix, but it is not positive definite when singular. This matters in portfolio and statistical calculations that require an invertible covariance matrix. The argument concerns the standard sample estimator and does not discuss remedies such as regularization or alternative estimators.
Key ideas
- Each centered observation contributes a rank-one outer product to the sample covariance matrix.
- Centering limits the covariance matrix rank to at most the number of observations minus one.
- With fewer observations than variables, the standard sample covariance matrix is singular.
- Singularity removes positive definiteness, but the matrix remains positive semidefinite.
Tags
Full text
# Proof for non-positive semi-definite covariance matrix estimator
# Proof for non-positive semi-definite covariance matrix estimator
It is well known that the standard estimator of the covariance matrix can lose the property of being positive-semidefinite if the number of variables (e.g. number of stocks) exceeds the number of observations (e.g. trading days). I think the matrix can become singular. I have a clear idea why (inspired by the geometry of the problem) but does anybody have a short but rigorous proof for this fact?
## Answer by JL344 (score 6, accepted)
https://quant.stackexchange.com/a/3861
The standard estimator of the covariance matrix is: $$\widehat{ \mathrm{cov}}(X) = \frac 1 {n-1} \sum_{i=1}^n (X_i-\bar X)(X_i-\bar X)^T,$$ where $X_i$ is the column vector containing the $i$th observation of all the observables. Each summand is an outer product of a vector with itself, i.e., a square matrix having rank at most one. Therefore $$\mathrm{rk\;}\widehat{\mathrm{cov}}(X) \le n$$ and the matrix not only can be but is always singular if $$n \lt \dim X,$$ i.e., if the number of observations is less than the number of variables.
Edit: regarding positive-semidefiniteness: $\widehat{ \mathrm{cov}}(X)$ is always positive-semidefinite because it is Gramian, even if its rank is not full. It loses the property of being positive-definite if and only if it is singular.Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.