Handling Multicollinearity with PCA, SVD, and Eigenvalue Shrinkage
Summary
The discussion explains why highly correlated asset returns create problems when estimating models or working with covariance matrices. Perfect collinearity makes the relevant matrix rank deficient, while near-collinearity can leave very small eigenvalues and make estimates numerically unstable. It frames the standard response as identifying the effective lower-dimensional space rather than immediately applying a regression penalty.
Principal component analysis diagonalizes the covariance structure, allowing analysis to focus on directions with nonzero or sufficiently large eigenvalues. Singular value decomposition is suggested for rectangular return data because it works directly with the data matrix. The discussion also describes identifying redundant stocks and replacing them with an equivalent portfolio, or using random matrix theory to shrink eigenvalues. Lasso and ridge can be useful for portfolio construction, but they answer a different question when no loss function or estimation objective has been specified. A central practical caveat is that deciding whether small eigenvalues represent noise or real structure is difficult; PCA components may also be less directly interpretable as individual stocks.
Key ideas
- Perfectly collinear returns make the covariance or Gram matrix rank deficient.
- Near-collinearity can make estimates unstable because small eigenvalues are hard to distinguish from zero.
- PCA identifies directions that span the nonredundant return space.
- SVD works directly with rectangular return matrices and avoids explicitly inverting a larger cross-product matrix.
- Econometric grouping and random matrix eigenvalue shrinkage are alternative practical approaches.
Tags
Full text
# What is the textbook answer to dealing with multicollinearity?
# What is the textbook answer to dealing with multicollinearity?
I have recently struggled in interviews, for two quantitative trading positions, by producing weak answers to effectively the same (fairly basic) question. I would like to understand, from a quant perspective, what I am missing about multicollinearity.
The question assumes you have a large portfolio of assets (say n=1000 stocks). As I recall, you prepare a covariance matrix (presumably of the price returns). The implication is that many of these returns are correlated. The question basically is, 'what is the problem with this, and how do you solve it?'.
Let $X\in \mathbb{R}^{m\times n}$ represent the matrix formed by concatenating vectors of each stock's returns observed over $m$ timesteps.
- My answer to 'What is the problem?':
If the returns are correlated, then there is some 'redundancy' in the matrix $X$ (in the extreme case, where a series of returns is identical to another, the matrix is underdetermined). I think that the implication is X is our matrix of features, and we are dealing with a linear regression model. Hence we are worried about the impact of inverting matrix $X^\top X$. If we have perfect multicollinearity, then this cannot be inverted; if we just have some multicollinearity, we will fit poorly, giving large errors/instability in our estimates for $\beta$.
- My answer to 'How do you solve it?':
Regularisation; the model has 'too many' features, and we should prioritise the more informative ones. L1 regularisation, in particular, allows for us to penalise solutions with many features and simplify the model (whilst retaining interpretability), so we could use a LASSO regression instead. L2 regularisation could also be used, but this doesn't, in general, reduce the number of features.
Unfortunately, I don't think these answers are textbook, so I would love some clarifications:
- Is this even a question about model fitting? Or is it really about variance-covariance matrices, portfolio risk, and/or CAPM-style financial management?
- An interviewer suggested using PCA instead of regularisation. I am not sure why that would be superior, since the principal components do not map to the original stocks you had in your portfolio)
- Does this apply to other models, which don't involve inverting $X^\top X$, or just linear regression?
## Answer by lehalle (score 8, accepted)
https://quant.stackexchange.com/a/75635
As one of the interviewers suggested, the expected answer starts with PCA and SVD.
Before detailing it, let's take a paragraph about the way you seem to "misunderstand" the problem: suggesting LASSO or Ridge is out of scope. Indeed these techniques are based on the penalisation of a loss function and in this question: Where is the loss function you plan to penalise? I would be the interviewer, such an answer would frighten me more than the candidate not proposing PCA.
Nevertheless, you get right the fact that this collinearity makes the inversion of $X^T X$ (in $N^2$) impossible because it is not full rank. Not being full rank means that you have to operate in the orthogonal of its Kernel, and the way to identify the kernel is to diagonalise the matrix and to work in the orthogonal of its kernel. What does it mean? You get the diagonal version of $X^T X=P\Delta P^{-1}$, and keep in mind that $P^{-1}=P^T$. Because the returns of $K$ stock are collinear, you should have $K-1$ zeros in the eigenvalues of $\Delta$, that are its diagonal values. Look at the operation of multiplying $X^T X$ by a vector $v$ (whatever it is): $$X^T X\cdot v=P\Delta P^T \cdot v = P\cdot\big(\Delta(vP)^T\big).$$ To "work in the orthogonal of the kernel" means that when the work is "rotated" by $P$, the last $k-1$ components of your vector $v$ face zeros. They correspond to the last $K-1$ components of $P$. This means that you can safely remove these coordinates: You can invert your matrix, if needed, in the space spanned by the $N-K+1$ first components of the PCA.
Numerically, since $X$ is in general rectangular with far more rows than columns, it is good to use a Singular Value Decomposition (SVD) decomposition. It prevents you from inverting a $N$ by $N$ matrix. It directly deals with the rectangular matrix.
In practice, it is not that easy because you find a lot of very small eigenvalues: are they zeros or not? is a complicated question. My advice is to get this Python code on scickit-learn, to keep only the first part and to try (last time I checked it succeeded to get returns of stocks from yahoo finance). They are different approaches to deal with that: the first is to do some econometrics to identify the stocks that are collinear and to replace them with an "equivalent portfolio" (or just keep one of them), that is equivalent to position your problem in the orthogonal of the collinear returns. The second is to rely on Random Matrix Theory that will tell you how to "shrink" the eigenvalues of the $X^T X$ Matrix.
A last remark about your LASSO proposal, it is indeed far from stupid from a portfolio construction perspective. You ca have a look at Bruder, Benjamin, Nicolas Gaussel, Jean-Charles Richard, and Thierry Roncalli. "Regularization of portfolio allocation." Available at SSRN 2767358 (2013). It very clearly explains how most of the portfolio construction penalisations make sense. Nevertheless, it is not the answer that is expected first, because it opens the door to sophisticated questions about portfolio construction. Especially in an interview, but also in practice, you should start by setting a baseline model, before trying something more complicated.Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.