Skip to content
All library documents

Interpreting PCA Eigenvectors and Weights in Long-Short Portfolios

Article Quant Q&A · Author: rwb

Summary

The document asks how to turn the leading principal component of asset returns into portfolio weights when its eigenvector contains negative entries. It notes that dividing eigenvector entries by their sum creates weights that sum to one, but this normalization can leave net and gross exposure different when some holdings are short. It also asks whether weights should instead use loadings scaled by the square root of eigenvalues.

The response emphasizes that an eigenvector's overall sign is arbitrary, and distinguishes variance explained by components from weights that map original variables onto principal-component axes. It describes loadings as the coefficients used to project the original data into component scores, and gives a variance-share expression for assessing how much variance a set of components explains. However, it does not settle a portfolio construction convention for net or gross exposure, nor resolve the scaling question for a specific PCA setup. The appropriate interpretation depends on whether PCA uses a covariance or correlation matrix and on the intended portfolio constraints.

Key ideas

  • An eigenvector and its sign-reversed version represent the same principal-component direction.
  • Normalizing signed eigenvector entries to sum to one does not determine a portfolio's gross exposure.
  • Loadings describe how original variables map onto principal-component axes and are used to form component scores.
  • The document does not prescribe a unique normalization for portfolio weights, which depends on the construction objective.

Tags

Full text
# How to extract normalised portfolio weights from PCA, when the eigenvector has negative elements?


# How to extract normalised portfolio weights from PCA, when the eigenvector has negative elements?












Most of the examples of using PCA of asset returns to construct an eigen portfolio seem to tend to focus on equities, which tend to all be positively correlated. As such I usually see normalised (such that they sum to 1) asset weights along the lines of:

```
weights = eigenvector[0, :] / sum(eigenvector[0, :])
```

Where eigenvector zero is associated with the eigenvalue of the largest variance.

Occasionally negative values will appear in that PC0 (for example if there is a negatively correlated asset introduced), which will be explained as a weight towards a short sale. But I am confused about how these weights should be correctly normalised in the presence of negative value(s).

Following the same procedure as above, the weights will of course sum to one, but the net and gross exposures now will of course be different. Is this just accepted?

Follow up question, should weights here actually be the loadings (eigenvector * sq.root(eigenvalues))?

## Answer by Carson McKee (score 3)

https://quant.stackexchange.com/a/65631

I'm not entirely sure what you are carrying out the PCA on, are you using a correlation matrix or covariance matrix? As for the negative eigenvalue issue, the sign associated with an eigenvalue is not important as eigenvectors are only unique up to a constant. If you have a square matrix, $X$, an Eigenvalue $\lambda_i$ and eigenvector $v_i$ for $X$ satisfy, \begin{equation} Xv_i = \lambda_i v_i \end{equation} If you flip the sign of $\lambda_i$ to $-\lambda_i$ then it implies that we have $-v_i$ as an eigenvector of $X$ and the above equation is still valid. These eigenvalues represent the variance explained by the corresponding Eigenvector in relation to the original axis.

For your question on weights, I'm not fully sure what you are weighting. You could allocate a 'weight' to each Eigenvector as $\frac{\sum_{j=1}^{k}\lambda_j}{\sum_{i=1}^{p}\lambda_i}$ which would represent the the proportion of variance explained by including the first k components. You may then use this to drop or retain components that contribute little to the overall variance of the portfolio.

If by weights you are instead interested in how to map your original data (an equities returns portfolio??) onto the new axis (principal components). Then for this you would use the loadings matrix. For a given column in the loadings matrix you have an eigenvector. Each element in the vector represents a 'weight' corresponding to one of your original variables which maps that variable onto the new axis (principal component) for this eigenvector. It then follows that by multiplying your original data matrix by the loadings matrix, you arrive at the score matrix (your original data mapped to the new axis of principal components). I believe the loadings matrix is the matrix of weights that you are looking for. Please add a comment if I've gone in the wrong direction here.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.