Why OLS Errors Have Covariance Matrix σ²I
Summary
This explanation interprets the OLS assumption that the error vector is multivariate normal with covariance matrix equal to a scalar variance times the identity matrix. The scalar sets the variance of each observation’s error, while the identity matrix specifies the structure across observations: equal diagonal entries and zero off-diagonal entries. The notation therefore captures both constant variance and the absence of cross-observation covariance in one expression.
Because the errors are jointly normal, zero covariance between distinct components also implies their independence. The explanation distinguishes this implication from the more general case where uncorrelated variables need not be independent. This is a statement about the assumed error distribution in the model, not evidence that real regression residuals satisfy the assumption. Its relevance to trading research is foundational: interpreting the covariance structure matters when using regression models and their statistical inferences.
Key ideas
- The identity matrix gives each error component the same variance and sets cross-component covariances to zero.
- The variance scalar determines the common variance of each observation’s error.
- Joint normality makes uncorrelated error components independent.
- The expression summarizes assumptions about the error vector rather than demonstrating that data meet them.
Tags
Full text
# Sigma squared times identity matrix in normality of errors
# Sigma squared times identity matrix in normality of errors
In OLS regression, we have the normality of the error terms
$$\varepsilon \sim N(0,\sigma^2I_n)$$
I understand that we want to have a constant variance for homoscedastic errors, but why is $\sigma^2$ multiplied with the identity matrix ($I_n$)? Is it just in order to transform $\sigma^2$ from a scalar into a matrix?
## Answer by rubikscube09 (score 3, accepted)
https://quant.stackexchange.com/a/58290
It is a neat way of succinctly saying the following:
- The error is a multivariate normal random vector with n components (in this case each component is one observation in the regression model.) That is to say: $$ \varepsilon = (\epsilon_1, \cdots, \epsilon_n) $$ where here each $\epsilon_i$ is a random variable that takes values in $\mathbb{R}$.
- The error has constant variance $\sigma^2$ in each component/observation because the variance of components correspond to the diagonal entries at each point in the covariance matrix. That is to say: $$ \mathrm{var}(\epsilon_i) = \sigma^2 $$ Moreover, the expectation of the errors is $0$.
- The error components are uncorrelated/orthogonal because covariances/correlations between different observations correspond to off-diagonal entries in the covariance matrix - which are all 0 for multiples of the identity. That is to say, the matrix with real entries: $$ \mathbb{E}[\epsilon \epsilon^T]_{i,j} = \mathrm{Cov}(\epsilon_i,\epsilon_j) = \begin{cases} = \sigma^2 & i = j \\ 0 & i \neq j\end{cases} $$ In particular: $$ \mathrm{cov}(\epsilon_i,\epsilon_j) = \mathrm{corr}(\epsilon_i,\epsilon_j) = 0 $$ for $i \neq j$, whereas: $$ \mathrm{cov}(\epsilon_i,\epsilon_i) = \mathrm{var}(\epsilon_i) = \sigma^2 $$ Because the errors have a multivariate normal distribution (they are an affine transformation of a vector whose components are independent normals), it follows that $\epsilon_i$ is in fact independent of $\epsilon_j$ (and not just uncorrelated).
All this said - putting the errors like this is a succinct notational way of summarizing the OLS assumption - $n$ error terms, all uncorrelated with each other, and all having constant variance.Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.