Skip to content
All library documents

Why Regression Significance Can Rise When Controls Are Added

Article Quant Q&A · Author: adrCoder

Summary

The document explains why a regression coefficient can become more statistically significant after adding explanatory variables. It gives two mechanisms. First, newly added controls can change the estimated coefficient by reducing omitted variable bias. In the example described, a control is negatively related to the variable of interest but positively related to the outcome; leaving it out can understate the original coefficient. Once included, the coefficient may rise enough to increase its t-statistic, provided its standard error does not rise too much.

Second, added variables may have little correlation with the predictor of interest while explaining more of the outcome’s variation. That can reduce residual variance and estimation uncertainty, lowering the coefficient’s standard error and raising its t-statistic. These are interpretations of a change in conditional estimates, not automatic evidence that the relationship is causal or robust. The discussion offers conceptual explanations rather than a diagnostic test, and it does not address complications such as model selection, multiple testing, or panel-specific assumptions.

Key ideas

  • Adding a control can change a coefficient by reducing omitted variable bias.
  • A larger coefficient can raise its t-statistic if the standard error does not increase substantially.
  • Controls that explain outcome variation may reduce residual variance and coefficient uncertainty.
  • Greater statistical significance after adding variables does not by itself establish causality or robustness.

Tags

Full text
# Variable becomes more significant when more variables are included


# Variable becomes more significant when more variables are included












I do some empirical research. I typically use regression analysis and panel data econometrics (with fixed effects).

Usually, when I include more variables, the initial variables of the model become less significant.

But, in a few cases, I see that one variable becomes even more significant when I include more variables.

Why is this happening? What is the interpretation of that? I suppose it has something with the correlation between the covariates and the residuals?

## Answer by zsljulius (score 4)

https://quant.stackexchange.com/a/21338

A Change in the statistical significance of a variable when more variables are included can come from two sources.

One is the fact that the estimate of the coefficient becomes larger when new variables are included. This can happen when the newly included variables are negatively correlated with the variable of interest and contribute positively to the depend variable. So essentially, you have an omitted variable bias that bias upwards your original estimates. You are wrongly contributing the negative effect of the omitted variable on the dependent variable to the variable included. Provided that the variance of the model doesn't increase much, your t-stat, $\frac{\beta}{std err}$, will increase, and thus you see a change in your statistical significance.

Another possible reason is that the newly included variables don't correlate with the variable of interest, but it reduces the model variance (an imporvement of Rsquare). This will help reduce your estimation uncertainty and thus lower the standard error of your estimate. Again, provided that the estimate doesn't change much, you will see an improvement in your t-stat.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.