How Factor Covariance Structure Can Simplify Portfolio Optimization
Summary
The document compares portfolio optimization using a full asset return covariance matrix with a factor model. In the factor formulation, each asset’s return is represented by exposures to a smaller set of factors plus an idiosyncratic component. The resulting covariance estimate combines factor covariance, asset exposures, and diagonal residual variances. The optimization objective is quadratic and may also include continuous and binary constraints.
The answer gives two potential benefits. Nonzero idiosyncratic variances can improve conditioning, reducing sensitivity to numerical round-off. An optimization can also introduce factor exposure variables and express risk through the smaller factor covariance matrix rather than forming every covariance interaction directly. Nonorthogonal factors need not be a barrier: a Cholesky factorization can rotate the loadings when the factor covariance is positive definite. The discussion does not quantify computational savings, compare solver performance, or address factor estimation error and the practical effects of constraints; the gains depend on the formulation and model assumptions.
Key ideas
- A factor model expresses asset returns through factor exposures and idiosyncratic risk.
- Diagonal residual variances can improve the conditioning of the estimated covariance matrix.
- Explicit factor exposure variables allow optimization to work with the smaller factor covariance structure.
- Cholesky factorization can rotate correlated factors when their covariance matrix is positive definite.
- The answer offers no benchmark of runtime or solution quality under particular constraints.
Tags
Full text
# Optimization: Factor model versus asset-by-asset model
# Optimization: Factor model versus asset-by-asset model
In portfolio management one often has to solve problems of the quadratic form $$ w^T \Sigma w + w^T c \rightarrow \min_{\omega} $$ with portfolio weights $w \in \mathbb{R}^N$ a constant $c \in \mathbb{R}^N$ and a covariance matrix $\Sigma \in \mathbb{R}^{N \times N}$. Furthermore we assume real world continuous and binary (e.g. cardinality) constraints.
For estimating the covariance matrix $\Sigma$ we can e.g. use the sample covariance of the returns of all assets - let's call this an asset-by-asset model. It is known that some care has to be taken here if $N$ is big and so forth, but this is not the point of this question.
On the other hand we can define factors $(F_k)_{k=1}^K$ with $K<N$. These factors have a covariance matrix $\Sigma_F \in \mathbb{R}^{K \times K}$. Denoting by $r_i$ the return of asset $i$ with $i = 1, \ldots, N$, we can write: $$ r_i = \sum_{k=1}^K e_{i,k} F_k + \epsilon_i, $$ with the meaning that the variation of the return is described by the exposures $e_{i,k}$ to the factors and some purely idiosyncratic risk $\epsilon_i$. In this case the covariance matrix of returns is given by $$ \hat{\Sigma} = e \Sigma_F e^T + \mathop{diag}(Var[\epsilon_1],\ldots,Var[\epsilon_N]), $$ where $e \in \mathbb{R}^{N \times K}$ is the matrix of all exposures and the "diag" part adds the idiosyncratic parts of variance at the main diagonal. Note that $\hat{\Sigma} \in \mathbb{R}^{N \times N}$.
I sometimes read that a problem with covariance matrix $\hat{\Sigma}$ from the factor model is easier to solve than with $\Sigma$. Is this true? If yes, then how can we see this? My personal answer so far is: no. Because both are $N \times N$ matrices and the structure does not help in general.
Especially if the factors are not orthogonal - what do we gain? If the factors come from a PCA then we might gain something but I wonder how many PCs we would need e.g. in a global equities portfolio ...
## Answer by erbian (score 3, accepted)
https://quant.stackexchange.com/a/35175
A few points.
First, In a typical factor model, the idiosyncratic piece (what you call $Var[\epsilon_k]$) is non-negligible, which results in a $\hat\Sigma$ that is going to be well-conditioned. From a numerical point of view, this is very convenient. The more well-conditioned a covariance matrix, the less susceptible the optimization routine is to round-off errors.
Second, one can introduce variables $l_1, ..., l_K$ with constraints that $\sum_{i=1}^N e_{i,j} w_i = l_j$ and then work with $\Sigma_F$ directly in the formulation of an optimization and not use a full $N\times N$ covariance matrix.
Third, from a portfolio optimization standpoint we don't care if the factors are orthogonal or not. We can rotate them to be orthogonal. As long as $\Sigma_F$ is positive definite, we can use the Cholesky decomposition to write it as $LL^T = \Sigma_F$ and then replace the loadings matrix $e$ with a rotated one $\tilde e =eL$ so that now the factors are orthogonal. That is $$e\Sigma_Fe^T = eLL^Te^T = \tilde e \tilde e^T.$$Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.