Skip to content
All library documents

Using PCA Loadings and Clustering to Group Stocks by Factor Exposure

Article Quant Q&A · Author: Simon Nicholls

Summary

The discussion distinguishes PCA scores, which describe observations such as dates in principal-component space, from factor loadings, which describe each stock’s exposure to the components. It explains why clustering loadings may emphasize a component with more dispersed exposures even when that component accounts for less total return variance. In the example, the first component represents broad market beta with relatively similar loadings across stocks, while a second component differentiates groups such as domestic businesses and exporters.

The answer argues that this can be useful if the goal is to identify meaningful groups of stocks rather than preserve the overall variance captured by PCA. The interpretation depends on whether the first component is genuinely a shared market factor and whether that common exposure matters for the grouping objective. If that assumption is unsuitable, hierarchical clustering is suggested as an alternative, since its stepwise merges provide a visible structure of similar stocks and groups. No empirical validation or specific distance metric is provided.

Key ideas

  • PCA scores represent observations in component space, while loadings represent each stock’s component exposures.
  • A component with less explained variance can still distinguish stocks if its loadings vary more widely.
  • A broad market component may explain most return variation yet contribute little to separating stocks.
  • Hierarchical clustering offers a structured, inspectable way to group stocks by similarity.
  • Choose and interpret clustering based on the grouping objective and the meaning of the components.

Tags

Full text
# PCA and K-means clustering on returns


# PCA and K-means clustering on returns












I am running a PCA on a set of returns and I would like to cluster the results of the output to group stocks that have similar factor exposures.

However when I run the PCA on the covariance of the returns, the PCA score (values mapped to new plane of PCs) gives me a matrix with dates and the principal components, to cluster on this would cluster on date therefore.

I can cluster on factor coefficients for each stock but then I have found this ignores the variance. For PC1 for instance the variance of loadings is very low compared to PC2 and therefore when clustering using the loadings it simply clusters using mainly PC2 which seem inherently wrong to me?

Or is this still correct and we can assume that because most stocks are loaded in a similar way to the PC1 then the clustering can’t determine much from that PC anyway.

I’m worried that I am missing some of the information from the variance here as PC1 explains 55% of the variance compared with PC2 at 18%!

## Answer by demully (score 3, accepted)

https://quant.stackexchange.com/a/60471

A classic problem, been there, done that, didn't buy the T-shirt ;-)

PCA and clustering (K-means, or hierarchical) are similar but different. They're both "unsupervised learning" methods; but one is essentially descriptive, while the other is essentially pragmatic and expedient. People want both, but they need to prioritise one first!

Your PC1/PC2 phenomenon actually makes a lot of sense for stocks. PC1 is beta; and the factor loadings here will indeed be tight - most stocks in the long-run have betas between ~0.8 and ~1.25. Imagine, for simplicity sake's, that your benchmark was dominated by domestic banks and exporting oilcos... Your foreign/domestic/FX/dollar PC2 would indeed generate a much wider set of loadings, even if this effect was much less significant in scale than basic beta (you 18% vs 55% of variance point). And clustering your stocks between your domestics/banks versus your exporters/commodities would, to me, make a lot of intuitive sense.

The caveat - you just have to be comfortable that PC1 is indeed just basic beta, and comfortable that beta is just a factor that all stocks share in common, rather than something that really differentiates them.

Fail that test, and you may need to get your hands dirty with hierarchical clustering. The method bottom-up merges the most similar stocks (or groups of stocks). So it starts off with the easy merging of your oilies, your miners, your tech, banks, industrials etc. Until it then has to start merging industries. But you will get a structured grouping of lookalikes, with transparency how it got there.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.