Why Simulated Sample Correlation Can Differ from Its Target
Summary
The document describes a Monte Carlo setup for a correlation swap: pairs of normally distributed asset returns are generated with a specified population correlation, and sample correlation is calculated across observation periods for each simulated path. The question is why the average of those realized correlations often appears below the input target.
The response identifies sample correlation as a biased estimator, which means its expected value need not equal the population correlation in finite samples. The observed difference therefore does not, by itself, show that the correlated random variables were generated incorrectly. The answer is brief and does not derive the direction or magnitude of the bias, quantify its dependence on observation count or target correlation, or offer a correction. Those limits matter when using simulated realized correlations to price a correlation swap; the discussion supplies a statistical explanation but not a complete pricing procedure.
Key ideas
- Sample correlation is a biased estimator of population correlation.
- Finite-sample realized correlations need not average to the target correlation.
- A discrepancy between target and realized correlation does not alone establish a simulation error.
- The response does not quantify the bias or give a correction for swap pricing.
Tags
Full text
# Simulating Correlation (but sample correlation is always too low)
# Simulating Correlation (but sample correlation is always too low)
I am trying to simulate correlation in order to price a correlation swap (via Monte-Carlo). For simplicity, let's assume we have 2 assets, and everything is correlated with $\rho$, and there is no drift or anything. Say, the swap has $M$ observation periods, and we simulate $N$ paths per asset.
So we start by using a multivariate normal distribution with $\mu=\begin{bmatrix}0\\0\end{bmatrix}$ and $\Sigma=\begin{bmatrix}1&\rho\\\rho&1\end{bmatrix}$. Now we generate $M$ correlated pairs of normal variables $R = \begin{bmatrix}r_{11}&\ldots & r_{1M}\\r_{21}&\ldots&r_{2M}\end{bmatrix} \in \mathbb{R}^{2\times M}$. We do this $N$ times, such that we obtain $R_1,\ldots,R_N$. For each sample $R_i$ we now compute the realised sample correlation $\bar{\rho}_i$ between its 2 assets.
Finally, we can compute the average of all realised sample correlations: $$\bar{\rho} = \frac1N \cdot\sum_{i=1}^N \bar{\rho}_i$$
What I have noticed is, the average realised sample correlation is disproportionally often lower than the target correlation, i.e. $$\bar{\rho} < \rho$$ I would have expected $\bar{\rho} \approx \rho$.
Below are a few sample plots, which highlight this effect (with 3 out of 4 paths lower than 50%, and 1 path around 50%):
Source Code to generate graph: zerobin
Question: Does anybody know why this is the case? Or am I generating random variables incorrectly?
## Answer by Dave (score 1)
https://quant.stackexchange.com/a/70727
Sample correlation is known to be a biased estimator.
Biased estimators are perfectly acceptable, and many common estimators are biased, such as:
- Sample standard deviation $s$
- Logistic regression maximum likelihood estimatorsShown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.