When OLS with an Intercept Matches Regression on Demeaned Data
Summary
The document explains why ordinary least squares estimates of slope coefficients are unchanged when a regression includes an intercept or, equivalently, when the response and regressors are demeaned and the regression is constrained to have no intercept. It frames the result through the method of moments: a fitted model with residuals having a nonzero mean could be improved by adjusting the intercept. The zero-mean residual condition implies that the fitted response and regressors satisfy a relation between their means, which is removed by demeaning.
The answer stresses that this equivalence is not universal across all fitting procedures. It applies to OLS in the stated setting, while some alternatives, including certain M-estimators, need not preserve it. The explanation gives the intuition and the key moment condition rather than a full matrix-algebra proof, so readers seeking a formal derivation may need to expand the normal equations or specify the estimator and assumptions more precisely.
Key ideas
- For OLS, including an intercept yields the same slope estimates as regressing demeaned variables with no intercept.
- The equivalence follows from the zero-mean residual condition associated with fitting an intercept.
- Demeaning removes the sample means that the intercept would otherwise explain.
- The result does not automatically extend to all regression estimators, including some M-estimators.
Tags
Full text
# Constant term in linear regresion
# Constant term in linear regresion
Can someone give a mathematical proof as to why including a constant in a linear regression equivalent is to running a regression with demeaned data and zero constant?
More specifically, consider the linear regression $$Y = b_0 + b_1 X_1 + b _2 X_2 + ... b_k X_k + e$$ where $X$'s and $Y$ are vectors, and the same regression with demeaned regressors $$\bar{Y} = b_1 \bar{X}_1 + b _2 \bar{X}_2 + ... b_k \bar{X}_k + e,$$ where the $\bar{Y}$ and the $\bar{X}_i$ are demeaned. It turns out that you'll obtain the same coefficients $\{b_i\}_{1 \leq i \leq k}$ in both regressions, but I struggle to prove this.
## Answer by Brian B (score 1, accepted)
https://quant.stackexchange.com/a/11073
Including a constant is "equivalent" to working on demeaned data only for certain cases. However among these cases is ordinary least squares (OLS) regression.
Basically, if you think of your fitting procedure as following a method of moments then you will have equivalence. Take our model as $$ Y = \alpha + \beta X + u $$
Say our model coefficients were chosen such that $E(u) \neq 0$, then the presumed symmetry of the distribution of $u$ would mean a "better" model is available by adding a further constant to $\alpha$. Therefore, we have a primary statistical condition that the first moment of $u$ be zero, and therefore
$$ E(Y)=\alpha+ \beta E(X) $$
Now, there are plenty of ways to fit regressions that are not methods of moments. For example, various M-estimators are "moment-like" but do not share this equivalence.Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.