Country Dummy Variables and Multicollinearity in Cross-Sectional Regression
Summary
The document considers a cross-sectional performance-attribution regression with asset returns as the response and style factors plus country indicators as predictors. It asks whether including a dummy for every country alongside a constant makes the model misspecified because of multicollinearity.
The answer confirms the concern: for a categorical feature with twenty possible countries, use twenty minus one indicator variables when the regression includes an intercept. The omitted country is represented by all included country indicators being zero, avoiding the exact linear dependence among the constant and the full set of dummies. The note provides no further discussion of coefficient interpretation, alternative coding schemes, or empirical results, so its guidance is limited to this standard setup.
Key ideas
- With an intercept, including indicators for every category creates exact multicollinearity.
- Omit one country indicator and treat that country as the reference category.
- The remaining country coefficients are interpreted relative to the omitted country, conditional on the other regressors.
Tags
Full text
# Regression based performance attribution with dummy variables # Regression based performance attribution with dummy variables I am following some work to do with a regression based performance attribution. The regression is a cross sectional one. The $y$ vector is the risk free return for say 1,000 companies. The $X$ matrix is made up of a constant, some factors such as book to price, momentum etc, say we 6 such factors (7 including the constant). Then for a stock's country there are dummy variables (twenty countries). So our matrix is 1000 x 27 (including the constant). However I thought when you have dummy variables you would not use all of them, you would use $n-1$ because it introduces multicollinearity. Is the regression above mis-specified? ## Answer by Martin Vesely (score 2) https://quant.stackexchange.com/a/53860 You are right that if you use binary dummy variables for $n$ possible values of some feature (the country in your case) you need only $n-1$ variables because the last (or first) country is indicated by all dummy variables equal to zero.
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.