Generating Correlated Quantities Along an Ordered Price Series
Summary
The document asks how to construct quantity values associated with an already ordered price series, while controlling their dependence, to study how correlation changes a total-revenue curve. The responses offer two approaches. For a short fixed vector, one can enumerate permutations, calculate each permutation’s rank correlation with the original sequence, and select those near the desired level. This method directly preserves the set of quantity values, but the number of permutations limits its use to small samples.
For larger vectors, one answer suggests sampling with replacement through an inversion-style method. Another proposes a covariance construction: model quantities as prices plus stochastic disturbances, specify the covariance of prices and disturbances, and use a Cholesky decomposition to generate the joint sample. The latter approach does not itself enforce an ordering; sorting prices together with their corresponding quantities afterward can produce an ordered display. These are alternative constructions with different assumptions, and the document does not establish a single method for achieving every target correlation or preserving a particular marginal distribution.
Key ideas
- For small vectors, permutation enumeration can identify rank correlations attainable by reordering values.
- Permutation enumeration becomes impractical as vector size grows.
- An inversion-style sampling method is suggested for larger vectors when sampling with replacement is acceptable.
- A joint covariance model can generate correlated prices and quantities through Cholesky decomposition.
- Sorting prices together with matched quantities can restore ordered presentation, but does not force quantities themselves to increase.
Tags
Full text
# Ordered correlated random numbers
# Ordered correlated random numbers
I'm not sure how feasible this is. I'm aware of how to generate correlated random numbers using a cholesky decomposition. However, say I have a fixed data set in increasing order (e.g. Price series: $10, $ 20, $ 30....$50). Now, I want to generated another series - of correlated quantities that follow the same order of increasing prices (so in the extreme case of 100% correl, Quantities have the exact same values - 10, 20, 30...50. But what would this look like if the correl were less than 100%). The idea is to see the impact on Total Revenue (P*Q) curve under different correlation assumptions. Cholesky did not work - as the Q series does not follow a specific order. I also tried using a regression say Q = a + bP; so if I changed b by changing the underlying correl..but that doesn't quite work (b = correl * stdev Q/stdev P; so I'll need to generate a new series that has the specified correl and stdev; it feels circular). In any case - this question itself might be daft...if so, please feel free to throw the eggs. But, if there's something possible - I'd love to know. thanks!
## Answer by Enrico Schumann (score 2)
https://quant.stackexchange.com/a/36078
For a small vector, you could compute all the permutations of the elements and calculate their correlation with the original vector. You could then pick those permutations that lead to correlations within your desired range.
In R, for instance:
```
library("e1071")
P <- seq(10, 50, by = 10)
perm <- permutations(length(P))
rho <- apply(perm, 1, function(r) cor(P, P[r], method = "spearman"))
```
The vector `rho` gives the rank correlation for every row in the matrix `perm`.
```
head(data.frame(perm, rho))
## X1 X2 X3 X4 X5 rho
## 1 1 2 3 4 5 1.0
## 2 2 1 3 4 5 0.9
## 3 2 3 1 4 5 0.7
## 4 1 3 2 4 5 0.9
## 5 3 1 2 4 5 0.7
## 6 3 2 1 4 5 0.6
```
For larger vectors and if sampling with replacement is allowed, I would probably go for some variant of the inversion method. In R, an implementation is in function `resampleC` in the NMOF package.
## Answer by Attack68 (score 1)
https://quant.stackexchange.com/a/36058
Instead of five elements lets reduce to two for simplicity, and call prices, $P$ and quantities, $Q$.
Suggest $P_1$ and $P_2$ have covariance matrix $\Sigma$. You might say $Q_1 = P_1 + X_1$ and $Q_2 = P_2 + X_2$, where $X_1$ and $X_2$ are your stochastic elements for quantities, and we let them be independent with non-identical variances. (note: if $\mathbf{P}=[10, 20]$ and $\mathbf{X}=[0, 0]$ then you return your simple scenario $\mathbf{Q}=[10, 20]$).
Now we have four random quantities to sample, $P_1, P_2, Q_1, Q_2$.
Written in matrix vector notation (with $\mathbf{Y,Z}$ added in for shorthand):
$$(\mathbf{Y}=)\begin{bmatrix}\mathbf{P} \\ \mathbf{Q}\end{bmatrix} = \begin{bmatrix}1 & 0 & 0 & 0 \\ 0 & 1 & 0 & 0 \\ 1 & 0 & 1 & 0 \\ 0 & 1 & 0 & 1 \end{bmatrix} \begin{bmatrix}\mathbf{P} \\ \mathbf{X}\end{bmatrix} (=\mathbf{AZ})$$
Then $$Cov(\mathbf{Y,Y}) = \mathbf{A} Cov(\mathbf{Z,Z}) \mathbf{A^T}$$
where, $$Cov(\mathbf{Z,Z})=\begin{bmatrix} \Sigma_{1,1} &\Sigma_{1,2} & 0 & 0 \\ \Sigma_{1,2} & \Sigma_{2,2} & 0 & 0 \\ 0 & 0 & \sigma_{X_1}^2 & 0 \\ 0 & 0 & 0 & \sigma_{X_2}^2 \end{bmatrix}$$ You can now reduce the covariance of $\mathbf{Y}$ to lower triangular by Cholesky decomposition and sample in the usual way. If you want to define the Q's differently say letting $Q_2=(P_2-P_1)+conts.$ then you just need to rework the equations, similarly if you want the X's to be correlated with each other.
This doesn't address the issue of ordering, but one could choose after the process to sort the P's and their corresponding Q's to give an ordered structure. One might also choose to define P as a constant vector plus random disturbance for which the equations could then be slightly re-worked.Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.