Skip to content
All library documents

How Hierarchical Risk Parity Uses Clusters in Portfolio Allocation

Article Quant Q&A · Author: Doggie52

Summary

The document explains a point of confusion in Hierarchical Risk Parity (HRP): the clustering tree is used to order assets through quasi-diagonalisation, but the standard allocation procedure then bisects that ordered list to form groups for recursive minimum-variance weighting. As a result, assets that were close in the original clustering tree can be separated during allocation. The weights are not necessarily built by combining every branch of the clustering tree from the leaves upward.

The cited discussion describes this as a limitation of naive bisection because it can discard information in the empirical hierarchy and split closely related assets. A proposed alternative recurses through the original clusters instead. The document reports that this cluster-aware approach does not consistently outperform naive HRP across return, variance, turnover, and concentration measures. It offers an algorithmic clarification and a caveat, but no detailed derivation or independent empirical results; conclusions about comparative performance depend on the cited study and its evaluation setup.

Key ideas

  • In standard HRP, hierarchical clustering helps order assets before allocation.
  • Naive HRP bisects the ordered asset list to build groups for minimum-variance weighting.
  • This bisection can separate assets that the clustering step identified as similar.
  • A cluster-aware alternative preserves the original hierarchy during allocation.
  • The cited comparison does not find the alternative consistently superior across reported metrics.

Tags

Full text
# Why does Hierarchical Risk Parity ignore the clusters generated?


# Why does Hierarchical Risk Parity ignore the clusters generated?












I am currently working through the Hierarchical Risk Parity algorithm (Lopez de Prado (2016) link ) and trying to understand each of the steps.

I have completed the step of creating the clusters, and visualised them using a dendrogram. Here's an example dendrogram for sake of illustration:

At this point, the way I thought HRP worked was by calculating the min-variance weights in each cluster from the bottom to the top. In the example dendrogram above, this would thus be:

- Calculate minvar weights between JPM and BoA

- Calculate minvar weights between JPM+BoA cluster (var calculated using the weights from #1) and BRK

- Calculate minvar weights between JPM+BoA+BRK cluster and Exxon

- Etc.

The final weights would then be the product of all the weights calculated in these steps.

I don't seem to be the only one with this mental heuristic. In the documentation for the PyPortfolioOpt library, the following "rough overview" is presented (my highlights):

> From a universe of assets, form a distance matrix based on the correlation of the assets. Using this distance matrix, cluster the assets into a tree via hierarchical clustering. Within each branch of the tree, form the minimum variance portfolio (normally between just two assets). Iterate over each level, optimally combining the mini-portfolios at each node.

Seemingly this is not how the HRP algorithm works. Instead, it seems to purely use the clusters to sort the assets in the quasi-diagonalisation step. It then takes this sorted list and creates its own new clusters by bisection, and then calculates minvar working from top to bottom.

In the example dendrogram, this would mean some assets we clustered in the first step would never end up in a cluster together after being sorted. E.g. Facebook and Alphabet would end up being separated in the first bisection step.

Have I understood this correctly? If so, why does this make sense? What is an intuitive way to understand why this works?

## Answer by Doggie52 (score 5, accepted)

https://quant.stackexchange.com/a/68727

This turns out to be a general drawback of the HRP algorithm, as pointed out by Pfitzinger, J., & Katzke, N. (2019) (my highlights):

> As shown in Figure 2.3, the naive bisection rule can violate the intuitive character of the result, by placing similar assets into separate clusters for allocation purposes. While centered bisection yields a symmetric allocation tree, which results in well-diversified portfolio weights, the method does not respect empirical cluster boundaries and discards information about the hierarchical structure inferred in the cluster algorithm. Figure 2.3 (top-left) demonstrates the concern, with the bisection separating two closely related assets.

The authors go on to propose an alternative method that does take into account the information contained in the clusters (by recursing into the actual clusters rather than naively bisecting), however their results are not consistently better than naive HRP across returns, variance, turnover and concentration metrics.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.