Calculating Average Asset Correlation from Portfolio Weights
Summary
The document examines two formulas for measuring average correlation among assets in a weighted portfolio. It presents an R example using simulated returns to build a covariance matrix, correlation matrix, and portfolio variance, then compares the two measures. The initial nested loops sum across incorrect index pairs, which includes self-pairs and duplicates terms, producing inconsistent results.
The accepted correction restricts the loops to unique pairs where the second asset follows the first. It also gives vectorized matrix calculations: zeroing the correlation matrix diagonal for the first measure and using the upper triangle for the volatility-weighted pair sum in the second. The example offers a practical implementation lesson, but it does not provide empirical market evidence or discuss estimation uncertainty, missing data, or behavior with degenerate weights or zero volatility.
Key ideas
- Average correlation can be computed using portfolio weights and pairwise asset correlations.
- The double sums should include each distinct asset pair once, excluding self-pairs.
- Portfolio variance and individual variances provide the numerator for a volatility-weighted correlation measure.
- Matrix operations can replace nested loops when calculating the weighted sums.
Tags
Full text
# Implementation of total correlation of assets in R
# Implementation of total correlation of assets in R
I am trying to implement the (average) total correlation of assets, which is discussed here and here in R. Specifically, I am looking at $\rho_{av(1)}$ and $\rho_{av(2)}$:
$$ \rho_{av(1)} = \frac{2 \sum_{i=1}^N \sum_{j>i}^N w_i w_j \rho_{i, j}}{1 - \sum_{i=1}^N w_i^2} $$
$$ \rho_{av(2)} = \frac{\sigma^2 - \sum_{i=1}^N w_i^2\sigma_i^2}{2 \sum_{i=1}^N \sum_{j>i}^N w_i w_j \sigma_i \sigma_j} $$
So far, I have the following code:
```
set.seed(123)
a = rnorm(100, 0, 0.02)
b = rnorm(100, 0, 0.03)
c = rnorm(100, 0, 0.04)
dt = cbind(a, b, c)
wg = runif(ncol(dt))
wg = wg/sum(wg)
wg = round(wg, digits = 3)
options(scipen = 999)
cv = cov(dt)
VARs = diag(cv)
cr = cor(dt)
length(wg)
sum_1 = 0
for (i in 1 : length(wg) ) {
for (j in 2 : length(wg) ) {
sum_1 = sum_1 + 2 * wg[i] * wg[j] * cr[i, j]
}}
rho_1 = sum_1/(1 - sum(wg^2))
sum_2 = 0
for (i in 1 : length(wg) ) {
for (j in 2 : length(wg) ) {
sum_2 = sum_2 + 2 * wg[i] * wg[j] * sqrt(VARs[i]) * sqrt(VARs[j])
}}
PV = t(wg) %*% cv %*% wg
ss = sum(wg^2 * VARs)
rho_2 = (PV - ss)/sum_2
rhos = c(rho_1, rho_2)
rhos
```
And the results are:
```
[1] 1.388473356 -0.004104948
```
I expected these two to be close to each other. I think there might be an error in my code. I would appreciate if somebody could verify the code.
Thanks.
## Answer by AK88 (score 1, accepted)
https://quant.stackexchange.com/a/34777
Here is the corrected original code (thanks @will):
```
sum_1 = 0
for (i in 1 : (length(wg) - 1) ) {
for (j in min(i+1, length(wg)) : length(wg) ) {
sum_1 = sum_1 + 2 * wg[i] * wg[j] * cr[i, j]
}}
rho_11 = sum_1/(1 - sum(wg^2))
sum_2 = 0
for (i in 1 : (length(wg) - 1) ) {
for (j in min(i+1, length(wg)) : length(wg) ) {
sum_2 = sum_2 + 2 * wg[i] * wg[j] * sqrt(VARs[i]) * sqrt(VARs[j])
}}
PV = t(wg) %*% cv %*% wg
ss = sum(wg^2 * VARs)
rho_22 = (PV - ss)/sum_2
```
A vectorized version:
```
# double summation in numerator (including multiplier 2)
diag(cr) <- 0
double_sum.1 <- c(crossprod(wg, cr %*% wg))
# single summation in denominator
single_sum.1 <- c(crossprod(wg))
rho_1 = double_sum.1/(1 - single_sum.1)
# single summation in numerator
single_sum.2 <- c(crossprod(wg * sqrt(diag(cv))))
# double summation in denominator
double_sum.2 <- sum(tcrossprod(wg * sqrt(diag(cv)))[upper.tri(diag(length(wg)))])
PV = t(wg) %*% cv %*% wg
rho_2 = (PV - single_sum.2)/ (2 * double_sum.2)
```
And the comparison of the two methods (initially they were way off):Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.