Skip to content
All library documents

Deriving the OLS Estimator from the Sum of Squared Errors

Article Quant Q&A · Author: Phil

Summary

The document asks how matrix notation expresses the ordinary least squares objective and how differentiation produces the estimator. It starts from the linear model, writes residuals as the observed response minus fitted values, and expands the residual sum of squares as a quadratic form. The intended first-order condition sets its derivative with respect to the coefficient vector to zero, yielding the normal equations and, when the inverse exists, the familiar closed-form OLS solution.

The question also asks whether a squared vector’s entries sum to its transpose times itself and where a factor of two enters the derivative. The excerpt includes a proposed derivation, but the displayed algebra is ambiguous and contains a likely scaling error: differentiating the quadratic objective gives a factor of two on both terms, which cancels in the first-order condition. It does not provide a complete answer or discuss cases where the design matrix lacks full column rank, so the formula’s invertibility condition matters.

Key ideas

  • The OLS objective is the sum of squared residuals, expressible as the residual vector transposed times itself.
  • Substituting the linear model expands the objective into a quadratic function of the coefficient vector.
  • Differentiating produces a factor of two on both terms, which cancels when forming the normal equations.
  • The closed-form estimator requires the relevant cross-product matrix to be invertible.

Tags

Full text
# Minimizing the sum of squared errors in linear regression (proof/matrix notation)


# Minimizing the sum of squared errors in linear regression (proof/matrix notation)












I'd appreciate you helping me understanding the proof of minimizing the sum of squared errors in linear regression models using matrix notation. I'm trying to derive by minimizing the sum of squared errors,

Look at this proof,

```
The q.c.e. basic equation in matrix form is: 
  y = Xb + e
  where y (dependent variable) is (nx1) or  (5x1)
        X (independent vars) is (nxk) or  (5x3)
        b (betas) is (kx1) or  (3x1)
        e (errors) is (nx1) or  (5x1)
Minimizing sum or squared errors using calculus results in the OLS eqn:
  b=(X'X)-1.X'y
To minimize the sum of squared errors of
a k dimensional line that describes the relationship 
between the k independent variables and y we
find the set of slopes (betas) that minimizes
Σ_{i=1 to n} e_i^2
Re-written in linear algebra we seek to min e'e
Rearranging the regression model equation, we get e = y - Xb
So e'e = (y-Xb)'(y-Xb) = y'y - 2b'X'y + b'X'Xb   (see Judge et al (1985) p14 )
Differentiating by b we get 0 = - 2X'y + X'Xb -> 2X'Xb=2X'y
Rearranging, dividing both sides by 2 -> b = X'X-1X'y
```

it is stated that if you rewrite the summation expression above, it is

Is any summation over a squared vector/matrix the product of the transpose with the vector/matrix itself?

Further in the proof, e is substituted by Y-X*beta, and e'e is differentiated with respect to beta, yielding this expression

and it is stated that this is equal to

Does anybody know where the second 2 comes from? If I rearrange it, it yields X'X*beta = 2X'Y without the 2 in front of the lefthand side of the equation.

Thank you!

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.