Skip to content
All library documents

Using PCA and Covariance Optimization to Build Hedge Baskets

Article Quant Q&A · Author: JoshK

Summary

The document discusses using principal component analysis to understand portfolio exposures and construct hedge baskets. It corrects an example that applies PCA to equity and ETF returns: the code uses a correlation matrix, while the stated method expects a covariance matrix supplied as such. The answer relates portfolio variance to the weight vector and covariance matrix, then explains how eigenvectors and eigenvalues describe risk directions and their associated variance. Choosing the lowest-variance component is one possible construction, but PCA alone does not specify a universally correct hedge or objective.

A second response cautions that estimated covariance matrices can be unstable, especially with many assets, and suggests shrinkage as a way to improve estimation. It proposes an alternative constrained optimization: minimize tracking error while controlling long-only weights, basket size, and sector exposure. The examples are conceptual and do not demonstrate out-of-sample hedge performance. Results depend on covariance estimates, constraints, and the selected objective, so the proposed methods require careful validation for a real portfolio.

Key ideas

  • Portfolio variance for a weighted return basket is determined by its covariance matrix and weight vector.
  • PCA eigenvectors describe directions of variation, while their eigenvalues quantify the associated variance.
  • Using covariance rather than correlation changes the PCA inputs and is necessary for the covariance-based calculation described.
  • The lowest-variance principal component is one possible portfolio construction, not a general hedge prescription.
  • Unstable covariance estimates can undermine PCA baskets; shrinkage and constraints can improve robustness.

Tags

Full text
# Using R with princomp to create hedge baskets


# Using R with princomp to create hedge baskets












I am experimenting to try to find better ways to hedge some of our equity portfolios. It's easy enough to use R to get a PCA breakdown of exposure for a portfolio but I can't figure out how to then use that to actually hedge something.

In the example below I'm just using four securities. In real life we would use more, but the small set it to keep this example manageable. Here's the R code:

```
# Get publicly available data
library(dplyr)
library(TTR)
px.spy=getYahooData("SPY", 20130101, 20160101)$Close
px.iwm=getYahooData("IWM", 20130101, 20160101)$Close
px.tlt=getYahooData("TLT", 20130101, 20160101)$Close
px.gld=getYahooData("GLD", 20130101, 20160101)$Close

names(px.spy)[1]="SPY"
names(px.iwm)[1]="IWM"
names(px.tlt)[1]="TLT"
names(px.gld)[1]="GLD"

px=merge( px.spy  , px.iwm )
px=merge( px  , px.tlt )
px=merge( px , px.gld )

# turn it into something that we can put into princomp
pxList=as.data.frame(px) %>% mutate( SPYr = (lag(SPY,1) - SPY )/SPY , IWMr = (lag(IWM,1) - IWM )/IWM  ,
 TLTr = (lag(TLT,1) - TLT )/TLT , GLDr= (lag(GLD,1) - GLD )/GLD) %>% filter( complete.cases(.)) %>% select(SPYr,IWMr,TLTr,GLDr )

pxCov=cor(pxList)
pc=princomp(pxCov)
```

Now we the loadings in `pc$loadings`. Now, say we wanted to take a portfolio of $100mm of SPY and hedge it with IWM,TLT,GLD based on the PCA exposures. What's the right way to extract that simply and to create a frame with the weightings?

Thanks, Josh

## Answer by Taylor (score 2, accepted)

https://quant.stackexchange.com/a/26299

Just a heads up, I'm not going to go through all the mathematical caveats of using this approach.

Let $\Sigma$ be your covariance matrix, and $X$ a random vector of daily returns. So

$$\text{Var}(X) = \Sigma.$$ You have a bug in your code. In your code you call it `pxCov`, but you probably meant to use `cov()` insted of `cor()`. Check out the documentation to `princomp().` It expects the covariance matrix by default. This will change things.

Also, `princomp()` is interpreting your covariance matrix as a data frame. You should run the following commands instead of your last two lines:

```
pxCov=cov(pxList)
pc=princomp(covmat=pxCov)
```

Now the math stuff. All the stuff you want to do, that you mentioned in the comments, can be done by taking expectations, variances, or covariances of linear combinations of your random return vector. The weights will be constructed from the loadings or eigenvectors, and calculation of variances will be simplified using the eigenvalues or standard deviation terms corresponding to the loadings.

Let $w$ be a weight (column) vector. Assume for now that the sum of its entires is $1$. Each element denotes the share of your portfolio sitting in that stock. Basic properties tell us that $$\text{Var}(w'X) = w'\Sigma w,$$ where the apostrophe denotes the transpose.

We can take the spectral decomposition of covariance matrices (positive definite and symmetric): $$\Sigma = \sum_{i=1}^4 \lambda_i v_i v_i'$$ with the set of loadings or eigenvectors, $\{v_i\}$ being orthonormal. In your code, these are `pc$loadings`. The lambdas that correspond with these are `pc$sdev^2`. This function orders everything in decreasing order, so $\lambda_1 > \cdots > \lambda_4$.

When you say PCA, you are probably trying to use the $v_i$s to construct a ''good'' $w$. There is no single way to do this. If you want to minimize your volatility without paying any mind to your expected returns, you would set $w = v_4$. Then your portfolio variance/risk would be

$$\text{Var}(wX) = v_4' \left[ \sum_{i=1}^4 \lambda_i v_i v_i'\right] v_4 = \lambda_4 v_4'v_4 v_4'v_4 = \lambda_4,$$ because of the orthonormality of the loading vectors.

If you want to minimize some other objective function, then you can do that too. For example, if your expected return is $w'\mu$, then you might want to minimize $w'\Sigma w - m w'\mu$. I have no idea what's commonly done in practice, though.

If you want to see how much each variance each loading explains, run `plot(pc$sdev^2/sum(pc$sdev^2))`. If you want 90\% you need some mix of the first two loadings. So we need to find some $0 < \alpha < 1$ such that $w = \alpha v_1 + (1-\alpha)v_2$. Using stuff like orthogonality that I showed you earlier, you can get

$$\text{Var}(wX) = \alpha^2\lambda_1 + (1-\alpha)^2\lambda_2 \overset{\text{set}}{=} .9$$

So you have to use the quadratic formula to find $\alpha$ in order to get your $w$. You could also do a mix of the last three loadings, too. I'm not sure what's done commonly in practice here, either.

If you want to find the "exposure" of your new portfolio, $w'X$, to the returns of the SPY, $e_1'X$, which I take to mean as covariance, then you can use the bilinearity of $\text{Cov}(\cdot,\cdot)$

$$\text{Cov}(w'X, e_1'X) = w'\Sigma e_1$$

where $e_1' = (1, 0, 0, 0)$.

## Answer by Richi Wa (score 0)

https://quant.stackexchange.com/a/26339

Several issues arise no matter which approach you choose (as a reference for my claims you can go through this:

- the covariance matrix of many assets can become instable (the more assets the more instable). Then your PCA will be based on noise. Therefore first get a good stimator of covariance. Using data of something like a year of observations worked good for me. Shrinkage is a good method to estimate the covariance matrix.

- Then you could use PCA ... but why should you? PCA with long and short loadings will be instable.

- For me the place to go would be to set-up an optimization problem (and solve it):

- minimize the tracking error (based on the stable covariance matrix)

- subject to: long only weights (this will make the thing more robust)

- and subject to:a cardinality constraint: you want to use $N$ assets for the basket

- and subject to: sector constraints: if all fails then you will have stocks with enough weight from every sector.

The above will give you a basked with only positive weights, controlled sector deviations and a control over the number of assets in the basket and finally a TE estimate.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.