Skip to content
All library documents

Choosing Hedge Ratios and Spread Statistics for Cointegration Pairs Trading

Article Quant Q&A · Author: Agustín Cugno

Summary

The document frames an out-of-sample estimation question for a cointegration pairs strategy. In sample, the proposed workflow applies the Engle–Granger two-step procedure, estimates a hedge coefficient for the spread, and standardizes that spread using its estimated mean and standard deviation. Signals are then generated when the spread’s z-score crosses a chosen threshold.

For each new observation, the trader must decide whether to refit the hedge ratio and spread statistics using updated data or retain the estimates from the original sample. Rolling means and standard deviations are also mentioned. The author recognizes that a truly stationary spread would have fixed population moments, but asks specifically how to handle an updated beta. The document poses the issue without an answer, evidence, or a recommended estimation policy, so it highlights a model-selection and adaptation problem rather than resolving it.

Key ideas

  • The in-sample workflow estimates a cointegrating hedge ratio, constructs a spread, and converts it to a z-score.
  • Out-of-sample signals depend on whether the hedge ratio and spread statistics are refit or held fixed.
  • Rolling estimates are another possible way to update spread normalization.
  • Stationarity implies constant population mean and variance, but estimated parameters may still vary.
  • The document raises the estimation question without prescribing a solution or reporting results.

Tags

Full text
# Which mean, standard deviation and beta should I use in a cointegration pairs trading strategy?


# Which mean, standard deviation and beta should I use in a cointegration pairs trading strategy?












I'm running a cointegration pairs trading strategy.

In sample, the implementation is very straightforward:

- run a Engle y Granger two-step procedure and reject $H_0$

- Calculate de hedge coefficient $\beta$ of the regression $s_t = y_t - \beta x_t$ (where $y$ and $x$ are my two assets)

- calculate the z-score of the spread $s_t$: $\displaystyle z_t = \frac{s_t - \mu_{s}}{\sigma_s}$ where $\mu_s$ and $\sigma_s$ are the mean and std of the spread in sample.

- Generate trading signals with some treshold (i.e. 2 std)

My problem is with out of sample signals: when a new observation (say, a pair ($y_{T+1}, x_{T+1}$)) comes into data, i have two alternatives:

- reestimate the parameters ($\beta, \mu_s$ and $\sigma_s$) with the new data in $T+1$ and calculate $z_{T+1}$,

- Use the original parameters calculated in - sample to calculate the $z_{T+1}$

Which should I use? I also seen in literature people calculating rolling means and standard deviations to calculate the spread.

Note: I understand the fact that if the spread is stationary it has constant mean and variance, so if I knew the "population" parameters, the problem of which estimator of $\mu$ and $\sigma$ (rolling, in sample, in sample + T+1) is inexistent. But my question is more related about the new "beta" that arises on the new regression with data in T+1.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.