Omitted Factors in Mixed-Effects Models of Asset Returns
Summary
The document examines a proposed return model combining Fama-French factors with industry random intercepts. The modeler reports a strong relationship between residuals and fitted values and suspects that correlations among industry effects or a nearly singular random-effects covariance matrix are responsible. The answer argues that the covariance matrix is not the cause, since the issue also appears in ordinary least squares and fixed-effects specifications.
Instead, it points to omitted common influences on both the explanatory factor returns and the portfolio returns being modeled. It sketches how an omitted variable correlated with the included factors can be absorbed partly into their estimated coefficients and partly into the residual, producing a residual relationship with the outcome. This is a conceptual diagnosis, not an empirical confirmation for the particular data. The document provides no tests to identify omitted factors, and it does not establish that the proposed explanation is the only possible cause of the reported pattern.
Key ideas
- The answer attributes the reported residual pattern to omitted common influences rather than the random-effects covariance matrix.
- The same issue appearing in OLS and fixed-effects models weakens the claim that random effects alone cause it.
- An omitted variable related to included factors can bias their estimated coefficients and contribute to structured residuals.
- The proposed explanation is not empirically tested on the data described in the question.
Tags
Full text
# Correlated random effects: mixed effects as a factor model
# Correlated random effects: mixed effects as a factor model
I am trying to build a fundamental factor model in the style of Fama-French. I have the FF factors as well as the industries in which my assets belong.
The model is specified as:
$$ \mathbf{Y} = \mathbf{X}\beta + \mathbf{Z}\gamma + \epsilon $$
where, $\mathbf{Y}$ is a cross-sectional time-series matrix of $N$ assets and $T$ weekly returns; $\mathbf{X}$ are the Fama-French factors (weekly returns of the HML, SMB, WML and Market-RF portfolios); $\mathbf{Z}$ is a matrix of random effects. I have a simple case whereby there are only random intercepts and no random slopes, i.e., $\mathbf{Z}$ is a matrix of 1s and 0s indicating the industry in which the asset belongs. $\epsilon$ are the residuals which are distributed $N(0,1)$, $\gamma$ are the random intercepts $N(0,\mathbf{G})$ and $\beta$ are the usual, fixed slopes (and one fixed intercept) in this case.
We can fit this model using `statsmodels`' mixed linear model.
The dataframe looks like this:
We fit the model by simply:
```
md = smf.mixedlm("y ~ MktRF + SMB + HML + WML", X, groups=X["sector"])
res = md.fit()
```
The results are horrible, however. There is almost a perfect linear relationship between the residuals and the fitted values.
My suspicion is that something is off with the covariance matrix, $\mathbf{G}$ since this is close to singular. One of the common approaches I read about regarding this issue is that the random effects might be overparametrized and that multicollinearity might be present there. Hence, I computed (with some padding out of neccessity) the correlations between the industries - our random intercepts.
Well, yes, the groups are correlated with one another (especially some groups).
Questions:
- What can we do about this? I could regroup some of the industries that are correlated, but this doesn't seem to be a great idea. For example, 'Financials' and 'Industrials' are quite correlated, but these are quite different groups and might have different correlations with other groups.
- Is the covariance really an issue? Could be something else as well?
## Answer by deblue (score 1)
https://quant.stackexchange.com/a/81105
The covariance $\mathbf{G}$ is not what causes this issue. The issue persists even in a simple OLS specification, as well as in a fixed-effects specification.
Since both $\mathbf{X}$ and $\mathbf{Y}$ are returns, there is an omitted variable bias. $\mathbf{X}$ are weekly returns of assets in SP500, while $\mathbf{Y}$ are weekly returns of certain portfolios (based on FF formulae and quantile segmentations), again, of assets in the US market (or developed) markets. Hence, it's highly likely that both $\mathbf{X}$ and $\mathbf{Y}$ have some common factors influencing both. If such a factor is omitted, this would lead to a bias estimator and correlated residuals.
This could be formally shown as follows. Suppose the true model is given by:
$$ \mathbf{Y} = \mathbf{X}\beta + \mathbf{U}\phi + \mathbf{Z}\gamma + \epsilon $$
Suppose we omit $\mathbf{U}$ and the relationship between $\mathbf{U}$ and $\mathbf{X}$
$$ \mathbf{U} = \mathbf{X}\alpha + u $$
What we will essentially estimate is actually:
$$ \mathbf{Y} = \mathbf{X}\beta + (\mathbf{X}\alpha + u)\phi + \mathbf{Z}\gamma + \epsilon \\ \mathbf{Y} = \mathbf{X}(\beta + \alpha\phi) + \mathbf{Z}\gamma + (u\phi+\epsilon) $$
Hence, if $\phi$ (slope) is non-zero, there is a linear relationship between the residuals and the dependent variable.Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.