Generating Correlated Monte Carlo Variables with Cholesky Decomposition
Summary
The document considers how to generate a value for one random variable that is paired with a particular simulated value of another. It rejects reversing a regression of one variable on the other as a satisfactory way to recover the paired sample, and instead describes constructing jointly distributed samples from their covariance matrix.
The proposed method factors the covariance matrix as a Cholesky product, then transforms independent standard normal draws with the resulting lower triangular matrix and adds the means. Across many draws, the generated variables are intended to have the specified covariance structure. The response also mentions principal component analysis as an alternative and notes that the covariance matrix of samples is associated with a Wishart distribution. This is a general multivariate normal sampling method; it does not explain how to condition on one already selected geometric Brownian motion path, and its displayed matrix notation may require checking when implementing the method.
Key ideas
- A regression rearranged to estimate one variable from another is not presented as a sound way to construct correlated simulation pairs.
- Cholesky decomposition can factor a target covariance matrix into a transformation for independent normal draws.
- Applying the factor to independent standard normals and adding the means generates paired samples with the intended covariance structure in large samples.
- Principal component analysis is mentioned as another decomposition approach.
- The response does not give a method for conditioning on a specific pre-existing simulated path.
Tags
Full text
# Given a particular Monte-Carlo simulation, how will a different correlated value change
# Given a particular Monte-Carlo simulation, how will a different correlated value change
I am currently working on a project at an investment bank regarding new accounting regulations on financial instruments. The task at hand it to understand the connection between a large array of variables. In its simplest form, my question can be presented as follows:
Let $\tilde{X}$ and $\tilde{Y}$ be two random variables with sample means $\hat{\mu_{\tilde{X}}}$, $\hat{\mu_{\tilde{Y}}}$, sample variances $\hat{\sigma_{\tilde{X}}}^2$, $\hat{\sigma_{\tilde{X}}}^2$ and sample covariance $\hat{\sigma_{\tilde{X}\tilde{Y}}}^2$.
I have ran 1000 Monte-Carlo simulations for $\tilde{X}$ under a geometric Brownian motion given by
$$d{\tilde{X_t}}=\hat{\mu_{\tilde{X}}}\tilde{X_t}dt+\tilde{\sigma_{\tilde{X}}}\tilde{X_t}dW_{t},$$
where $W_{t}$ is the Wiener process; normally distributed - $\mathcal{N}(0,t)$
One particular simulation, say $i$ of 1000 is of peculiar interest. I am now interested in the possible values for $\tilde{Y}$.
I have tried to take an econometric approach. That is, run a regression on $\tilde{X}$ with $\tilde{Y}$ and using the regression equation
$$\tilde{X_i} = \tilde{\beta_0} + \tilde{\beta_1}{\tilde{Y_{i}}},$$ $$\implies \tilde{Y_i} = \frac{\tilde{X} - \tilde{\beta_0}}{\tilde{\beta_1}},$$
try and back out the values of $\tilde{Y_i}$.
I am not quite pleased with the results. I am therefore hoping for an alternative method. I have other things in mind - that is, creating a new distribution for the sample but was hoping to gather the thoughts of the experts on this forum.
Thank you in advanced to everyone who takes the time to read and/or provide their thoughts on the matter.
Gus
## Answer by Attack68 (score 1, accepted)
https://quant.stackexchange.com/a/36759
Quantuple's comment is completely relevant here. You have defined the relationship between X and Y through the covariance matrix. The traditional way of generating a set of paired variables (or multiple variables) using the monte carlo method is to use Cholesky decompisition on the matrix. You have: $$ \Sigma = \begin{bmatrix} \sigma_x \sigma_x & \rho_{xy} \sigma_x \sigma_y \\ \rho_{xy} \sigma_x \sigma_y & \sigma_y \sigma_y \end{bmatrix} $$ Define the Cholesky decomposition: $\Sigma = \mathbf{LL^T}$ where $\mathbf{L}$ is a lower diagonal matrix. For the 2x2 case we have: $$\mathbf{L}= \begin{bmatrix} \sigma_x & 0 \\ \rho_{xy} \sigma_y & \sigma_y\sqrt{(1-\rho_{xy}^2)} \end{bmatrix} $$ Now if, for each iteration you define $\mathbf{Z}=[Z_1, Z_2]$ two independent standard normal variables, the combination $\mathbf{LZ}+\mu$ will define paired samples whose covariance matrix tends to what you want for large samples. I believe (although have not looked for a long time) that the distribution of the covariance matrix for n samples is defined by the Wishart distribution. You can also use PCA decomposition for this but is more computationally expensive I believe.Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.