Maximum Likelihood Estimation Derives Ordinary Least Squares for Linear Regression
Summary
The article frames multiple linear regression as a probabilistic supervised learning model. It assumes that the response is a linear function of the features plus normally distributed error with constant variance. It also explains that transformations of the inputs, such as polynomial and interaction terms, can represent nonlinear relationships while keeping the model linear in its coefficients.
It derives ordinary least squares by maximizing the likelihood of independent observations, equivalently minimizing the negative log likelihood. Under the Gaussian error assumption, this objective reduces to minimizing residual sum of squares. Differentiating the matrix form and setting the result to zero yields the familiar coefficient estimate, provided the feature matrix has full rank so that the solution is unique. The article’s main evidence is this mathematical derivation; it gives no empirical trading test or forecast evaluation. It notes that high dimensional or collinear data can make the required matrix noninvertible, motivating subset selection and shrinkage methods in later work.
Key ideas
- The model treats responses as conditionally Gaussian around a linear mean with fixed variance.
- Independent observations allow the log likelihood to be expressed as a sum of individual log probabilities.
- With Gaussian errors, maximizing likelihood for the coefficients is equivalent to minimizing residual sum of squares.
- The ordinary least squares solution requires the design matrix to have sufficient rank for a unique estimate.
- Basis expansions can capture nonlinear feature relationships while preserving linearity in the coefficients.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.