Skip to content
All library documents

Choosing Between PCA and Clustering for Portfolio Risk Models

Article Quant Q&A · Author: chrisaycock

Summary

The discussion compares principal component analysis (PCA) and clustering as ways to build risk models for portfolio optimization. PCA is presented as a method for reducing a set of risk variables to a smaller number of uncorrelated factors. Clustering instead groups securities or other observations into segments with related characteristics; the methods can also be combined.

The responses offer practitioner experience but no controlled comparison for portfolio optimization. One contributor reports that PCA experiments evolved into more complex approaches without impressive results, while clustering led to classification methods that performed similarly to, or possibly better than, a nearest-neighbor approach. The contributors caution that computational methods alone may not uncover excess returns or lower risk. They may still reveal overlooked portfolio issues, and the distinction between factor variables and nominal asset-class segments can guide method selection.

Key ideas

  • PCA can reduce correlated risk variables to a smaller set of uncorrelated factors.
  • Clustering is suited to grouping securities or nominal categories into risk segments.
  • PCA and clustering can be used together because they serve different purposes.
  • Practitioner comparisons in the discussion were not conducted for portfolio optimization.
  • Data-driven methods may reveal portfolio problems without reliably finding excess returns or lower risk.

Tags

Full text
# Cluster analysis vs PCA for risk models?


# Cluster analysis vs PCA for risk models?












I built risk models using cluster analysis in a previous life. Years ago I learned about principal component analysis and I've often wondered whether that would have been more appropriate. What are the pros and cons of using PCA as opposed to clustering to derive risk factors?

If it makes a difference, I'm using the risk models for portfolio optimization, not performance attribution.

## Answer by bill_080 (score 5, accepted)

https://quant.stackexchange.com/a/745

I've played around with both schemes, but not for portfolio optimization.

I used PCA on some interest rate models. That turned into a Partial Least Squares scheme, then into some non-linear thing. I wasn't impressed with the results.

My Cluster Analysis scheme morphed into a classification scheme, and it turned out that the K-Nearest-Neighbor method worked just as well, and possibly better. Again, this wasn't for portfolio optimization, so it may not apply to your situation.

From what I've seen, if you're depending on the computational method to find excess returns (or lower risk), you'll probably be disappointed. On the other hand, it is common for various methods to highlight some problems that weren't originally obvious. For instance, bootstrapping your portfolio(s) to determine just how good they are compared to luck. I've dumped a lot of ideas because of that issue.

## Answer by Ralph Winters (score 7)

https://quant.stackexchange.com/a/756

They are not mutually exclusive. PCA and clustering are similar but used for different purposes. You could use PCA to whittle down 10 risk factors to say 4 uncorrelated factors, and you could combine securities with different FACTORS into different clusters with offsetting returns and variance characteristics. However, when you say you want to derive risk factors, that implies that you are dealing more with variables, and PCA (or factor analysis) is more appropriate. If you are really interested in risk segments across nominal variables, say asset classes, you would be interest more in clustering.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.