Skip to content
All library documents

Extending Pairs Trading to Multiple Assets with Cointegration and Error Correction

Article Quant Q&A · Author: Xerium

Summary

The document explores how to size a trade in one asset when several other assets have lagged correlations with it. It raises the problem of double-counting signals when the predictor assets are correlated with each other, and asks how to estimate the conditional probability of the target asset’s next move. It does not provide a direct solution for that lagged-prediction setup.

The response recasts a two-asset spread as a normalized cointegrating regression, noting that the regression depends on which asset is chosen as the dependent variable. It outlines estimating the log-price relationship, testing whether its residuals are stationary, and fitting an error-correction model to describe short-term changes and adjustment toward the long-run relationship. The model may help assess whether a spread is likely to revert or widen. The discussion does not establish that the proposed delayed-correlation strategy is profitable, and it leaves the multi-asset signal construction question unresolved.

Key ideas

  • A multi-asset prediction signal must account for dependence among the predictor assets to avoid counting shared movements repeatedly.
  • A two-asset spread can be represented as a cointegrating regression, but its coefficients depend on which asset is treated as the dependent variable.
  • The suggested workflow estimates a log-price regression and checks whether its residuals are stationary.
  • An error-correction model links short-term price changes to both other asset changes and the lagged spread.
  • The response does not specify trade sizing for the lagged-correlation example.

Tags

Full text
# Pairs trading with 3+ assets


# Pairs trading with 3+ assets












I am trying to understand how you would construct a pairs trading strategy on 3+ assets.

In the 2 asset case, assuming zero drift, we trade based on:

$$dX_t = \beta_AdS_t^A - \beta_B dS_t^B,$$

where $\beta_A=\frac{cov\left(S^A,S^B\right)}{\sigma^2_A}$, $\beta_B=\frac{cov\left(S^B,S^A\right)}{\sigma^2_B}$is the beta. But these betas' just become the ratio of the variances since their covariances are the same. Therefore, we effectively choose to short/long the ratio of their variances. So if $$\sigma^2_A = 0.20, \quad \sigma^2_B = 0.10$$

If stock A goes up $\\\$10$, then if we shorted $100$ shares of stock $A$, we would long $100\cdot\frac{0.20}{0.10}= 200$ shares of stock $B$.

If we decided to instead have 3 stocks $A, B, C$ and we have the correlation matrix: $$ \Sigma=\left[\begin{array}{lll} 1 & \rho_{A,B} & \rho_{A,C} \\ \rho_{B,A} & 1 & \rho_{B,C} \\ \rho_{C,A} & \rho_{C,B} & 1 \\ \end{array}\right] $$

How could this be applied if we only traded 1 stock (let's choose $A$). (From here onwards, we assume the returns to have mean 0, $\mu=0$). In my contrived scenario, asset $A$ has the above correlation matrix, but the correlated returns movement is always delayed. In other words:

$$\text{corr}(dS^A_{t+1}, dS^B_{t}) = \rho_{A,B}$$ $$\text{corr}(dS^A_{t+1}, dS^C_{t}) = \rho_{A,C}$$

So we don’t need to long and short the pair since we have some certainty what A will be at time $t+1$ given the returns at $t$. We can just long/short $A$ depending on their returns and correlation of $B$ and $C$

In this case, what would the trade size amount be?

My initial intuition is that the size is equal to $$trade = \sum_{i=1}^{2}\text{movement} \times \text{correlation}\times \text{variance ratio}$$

So if $\sigma_B = \sigma_C$ and $\rho_{A,B} > \rho_{A,C}$, but then $dS^B_t=1$ and $dS^C_t=-1$ then we would buy $A$ in anticipation that it will rise since $A$ has a stronger correlation with $B$ (similarly if their correlations are equal, but 1 has higher volatility and or 1 has a larger movement).

This seems like a reasonable strategy given that we have this delayed correlation. The issue I run into with my logic is that if we have 4 assets instead. We can use the same logic as before:

$$trade = \sum_{i=1}^{3}\text{movement} \times \text{correlation}\times \text{variance ratio},$$ but the issue becomes what about the correlations between stocks $B,C,D$?

If $$\rho_{A,B} = 0.3$$ $$\rho_{A,C} = 0.3$$ $$ \rho_{A,D}=0.6$$ $$\sigma_A = \sigma_B = \sigma_C = \sigma_D$$ but $C$ and $D$ are highly correlated, like $\rho_{C,D} = 0.9$. Then in the scenario:

$$dS^B_t = -10$$ $$dS^C_t = 3$$ $$dS^D_t = 4$$

$D$ went up because it's highly correlated with $C$ (and vice-versa). Then we would have:

$$trade = -10 * 0.3 + 3 * 0.3 + 4 * 0.6$$ (ignoring the variance ratio since their vols are equal). To me, this doesn't seem reasonable because the movement from asset $D$ is kind of "baked into" asset $C$. So summing them would be kind of like double-counting?

How would I get around this?

I’m effectively asking: $$\mathbb{P}(dS^A_{t+1}> 0 | dS^B_t= -10, dS^C_t = 3,dS^D_t=4) $$

(If it’s bigger than $0.5$ we buy)

But I don’t think that’s possible since the events $\{dS^B_t= -10, dS^C_t = 3,dS^D_t=4\}$ aren’t measurable so conditioning on them isn’t possible.

## Answer by mark leeds (score 0)

https://quant.stackexchange.com/a/79801

Hi Xerium: This is not an answer but more like a long comment. Let me show how your 2 asset model is really just the engle-granger cointegrating regression model un-normalized. Once you follow that, then, for the 3 asset case, you can either go the way you originally described ( which I have not had a chance to go through ) or you can use Johansen's generalization which I was talking about earlier.

Take the case where you consider just 2 assets.

Your model is:

$$dX_t = \beta_AdS_t^A - \beta_B dS_t^B ~~~(1) $$

But, suppose I change some of the notation. So, I let $dX_t$ be equal to $\epsilon_t$. Also, I re-define things so that $$dS_t^{A} = log(S_t^{A})$$ and $$dS_t^{B} = log(S_t^{B})$$.

I then normalize the relation in (1) by setting $\beta_A = 1$.

These changes above result in the following model:

$$log(S_{t}^{A}) = \beta_{B} \times log(S_{t}^{B}) + \epsilon_t ~~~(2) $$

But note that (2) is just the engler-granger cointegrating regression model where the cointegrating vector is $(1, \beta_B)$ rather than $(\beta_A, \beta_B)$.

Also, we no longer have the nice symmetry in the covariance so now $\beta_{B} = \frac{cov\left(log(S_t^{A}),log(S_t^{B})\right)}{\sigma^2_{B}}$

So, the formulae are kind of the same as you had for your case, except that $\sigma^2_{A} = 1$ now because stock A and stock B are no longer viewed as interchangeable. This is a drawback of the engle-granger cointegration approach in that the results can be dependent on which stock is chosen to be the dependent variable. People have constructed ways of dealing with this that aren't worth getting into here.

My point is that you are pretty much doing un-normalized engle-granger cointegration but you didn't explain how you obtained the variances, $\sigma^2_{A}$ and $\sigma^2_{B}$.

In the regression formulation that I wrote, $\sigma^2_{B}$ comes directly from the regression of log prices of A on the log prices of $B$. Also, the residuals of the regression model need to be checked for stationarity. It's not clear to me how you obtain the variances without running a regression relation of some kind. Do you run total least squares possibly ?

### ARBITRAGE IN THE 2 ASSET CASE

Note that, just because one arrives at the regression model, it's still not clear how one would do the arbitrage if the long term equilibrium relationship seems out of sync.

One way to about it is to write out the ECM implied by the cointegrating regression model:

Given (2), the corresponding ECM model is shown below: (See this for a heuristic derivation: https://users.wfu.edu/cottrell/ecn215/misc/extra/error_corr_2008.pdf )

$$\Delta log(S_{t}^{A}) =\gamma_1 \times \Delta log(S_{t}^{B}) + \gamma_2 \times (log(S_{t}^{A}) -\beta _{B} \times log(S_{t}^{B})) + \nu _{t}$$

So, the steps become:

- Estimate the model using Ordinary least squares:

$$log(S_{A}^{t}) = \beta \times log(S_{B}^{t}) +\epsilon _{t}$$

- Check that the estimated residuals, $\hat{\epsilon_t}$ are stationary. If they are, then the estimated residuals ${\hat{\epsilon _{t}}= log(S_{t}^{A}) - \beta _{B} \times log(S_{t}^{B}})$ from this regression are saved and used in the ECM representation.

- $ \Delta log(S_{t}^{A}) = \gamma_1 \times \Delta log(S_{t}^{B})+ \gamma_{2} \times \hat{\epsilon}_{t-1}+\eta _{t}$.

Now estimate (3) to find $\hat{\gamma_1}$ and $\hat{\gamma_2}$.

The ECM representation in Step 3) is useful because, once you have $\gamma_1$ and $\gamma_2$, you then have a way of measuring when it looks like the true current spread ${\epsilon}_{t}$ is going to either A) revert from it's currently large value or widen due to its currently small value. Without the ECM representation, there is no easy way to really gauge the long term dis-equilibrium for an arbitrage possibility.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.