Skip to content
All library documents

Choosing an Intercept in Engle–Granger Cointegration Regression

Article Quant Q&A · Author: Tom

Summary

The document compares three ordinary least squares specifications for estimating a hedge ratio before applying an augmented Dickey–Fuller test to the residual spread: regression through the origin, regression with an intercept, and the usual regression with an intercept where the slope is taken as the hedge ratio. It highlights that these choices can produce different slopes and stationarity results, reflecting a known ambiguity in the Engle–Granger approach.

The response suggests checking whether conclusions are consistent across specifications and choosing whether to include an intercept based on the assumed relationship between the series. It favors omitting the intercept in this context, reasoning that an intercept can allow more variation in the slope and may weaken its stability out of sample. This is presented as judgment rather than a universal rule. The document does not define formal model-selection criteria, explain residual-based test critical values, or provide evidence that one specification is generally superior; the example’s p-values are illustrative only.

Key ideas

  • The intercept choice changes the estimated slope and can change the residual stationarity result.
  • Regression through the origin and regression with an intercept generally yield different ordinary least squares slopes.
  • Whether to include an intercept depends on the assumed economic relationship between the two series.
  • The response favors omitting an intercept when a stable hedge ratio is the priority, but presents this as a preference rather than a rule.

Tags

Full text
# What value to put in lm() function when testing for cointegration (R)


# What value to put in lm() function when testing for cointegration (R)












I'm a CS student working on a financial computing project + have a question regarding cointegration testing using linear regression with the lm() function.

https://www.rdocumentation.org/packages/stats/versions/3.6.2/topics/lm

Data:

I've seen many examples through different strategies/notes online and was wondering which is the correct one to use under certain scenarios ( +0, +1, or nothing)

eg:

```
  m <- lm(series[[9]] ~ series[[1]] + 0)
  beta <- m$coefficients[1]
  cat ("Assumed hedge ratio is ", beta, "\n")
  sprd <- series[[9]] - beta * series[[1]]
  adf.test(sprd, alternative = 'stationary', k=0)$p.value #0.6647128

  m <- lm(series[[9]] ~ series[[1]] + 1)
  beta <- m$coefficients[1]
  cat ("Assumed hedge ratio is ", beta, "\n")
  sprd <- series[[9]] - beta * series[[1]]
  adf.test(sprd, alternative = 'stationary', k=0)$p.value #0.5656023

  model <- lm(series[[9]] ~ series[[1]])
  b <- model$coefficients[2]
  spreadp1 <- series[[9]] - b*series[[1]]
  adf.test(spreadp1, k=0)$p.value # 0.4339312
```

## Answer by mark leeds (score 1)

https://quant.stackexchange.com/a/53347

Hi: There are a couple of different issues in what you're doing.

A) One key question is what does "series" contain ?

B) 1) and 3) are always going to differ and it's never clear which is correct ( it's one of the pitfalls of the EG test ) I would do it both ways and see if the adf test result is consistent. Don't worry about the lack of consistency in the least squares estimate of the two approaches. The two procedures will only result in the same coefficient if you use total least squares regression instead of OLS.

C) Whether you use 2) versus (1 or 3) depends on whether you think that there is an intercept in the underlying model.

In the context of testing for cointegration, I would be inclined to not include an intercept because it sort of "locks" one series into being a specific amount higher ( or lower ) than the other series. Also, you want your least square estimate to be relatively stable over time ( i.e: when you go out of sample, you hope that your least squares estimate doesn't change ). By including an intercept, you're allowing for more flexibility in the non-intercept coefficient which is probably going to make it less stable out of sample.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.