Why Fama-French Factors Do Not Need to Be Orthogonal
Summary
The document explains why the Fama-French three factors can be correlated without invalidating the model. In a classical factor model, researchers specify factor return series and estimate each asset’s factor loadings with time-series regressions. Those series need not be orthogonal, though low correlation can make risk decomposition easier because factor variances are then closer to additive.
It contrasts this with Barra-style models, which specify exposures and estimate factor moves through cross-sectional regressions; those estimated moves also need not be orthogonal. Principal component analysis instead estimates factors and loadings together, producing orthogonal time series by construction, but the resulting components may not correspond to intuitive economic factors and can vary with the sample window. The responses caution that raw covariance alone can mislead when factors use different units or have weak linear relationships. The discussion is conceptual and does not provide a formal test of multicollinearity or a practical threshold for acceptable factor correlation.
Key ideas
- Classical Fama-French models do not require factor return series to be orthogonal.
- Low correlation can help make factor risk contributions easier to interpret.
- Barra-style models estimate factor moves from specified exposures and also need not produce orthogonal moves.
- PCA imposes orthogonality but may yield less interpretable factors that depend on the sample window.
- Raw covariance alone is not sufficient to diagnose problematic multicollinearity.
Tags
Full text
# Why aren't the Fama-French 3 factors orthogonal to each other?
# Why aren't the Fama-French 3 factors orthogonal to each other?
I am confused whether the factors in a multi-factor model should be orthogonal or not. Google searches do not give a well documented answer and I couldn't find one in our library's limited catalog either. Intuition says they should. Moreover, the covariance of `Mkt-RF` with `SMB` and `HML` (on yearly data as obtained fron K. French's data library) is about 116 and 34 respectively, far from 0.
What am I missing?
## Answer by Kiwiakos (score 9, accepted)
https://quant.stackexchange.com/a/22604
A factor model has the form $$r_{j,t}=\sum_n \beta_{j,n} f_{n,t}+\epsilon_{j,t}$$ Where $r_{j,t}$ is the return of stock $j$ at time $t$, $\beta_{j,n}$ is the sensitivity (factor loading) of stock $j$ to factor $n$, $f_{n,t}$ is the return of factor $n$ at time $t$, and $\epsilon_{j,t}$ is the idiosyncratic non-factor return. One factor can be the constant.
There are three ways to specify and/or estimate:
- The classical Capm/ Fama-French where you explicitly specify the factor series $f_{n,t}$ and use time series regressions, one per stock, to estimate the betas $\beta_{j,n}$. There is no reason for the factor time series to be orthogonal, although it is useful as a risk decomposition if they are close to orthogonal (as factor variances become additive).
- The Barra approach where you explicitly specify the loadings $\beta_{j,n}$ and use cross-sectional regressions, one per date, to estimate the corresponding factor moves $f_{n,t}$. There is no reason, again, for these estimated moves to be orthogonal, but being close to orthogonal is, again, desirable.
- The black box PCA approach, where both factor loadings and time series are estimated simultaneously. Then we have time series that are orthogonal by construction (because we have too many degrees of freedom we put orthogonality as a constraint). However, they do not map directly to an intutive set of macro factors, although they often resemble them. Also different time windows would give rise to different factors, which might not be desirable.
## Answer by Brumder (score 1)
https://quant.stackexchange.com/a/22603
Using covariance to imply an inappropriate level of multicollinearity in a model can be very misleading, especially when the factors are measured in differing units or lack linear relationships. There will almost always be some level of collinearity in a multi-factor model (otherwise you run the risk of overfitting), especially one with a relatively small amount of explanatory variables like the FF equation. Remember, the FF model was really just an improvement on the CAPM to give a better ex-post fit to stock returns.Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.