Skip to content
All library documents

Avoiding False Cointegration from Misaligned Hedge-Ratio Multiplication

Article Quant Q&A · Author: Edwin Jose Palathinkal

Summary

The document examines an apparent mean-reverting basket built from several foreign-exchange rates. The proposed process fits a linear regression to estimate hedge ratios, forms a spread, then applies an augmented Dickey–Fuller test. The reported significance appears suspicious, and a backtest loses money.

The accepted explanation identifies a coding error: multiplying a coefficient vector by a data frame performs element-wise multiplication with recycling in column-major order, rather than applying each coefficient to its intended series. The resulting spread is therefore wrong, and manually matching coefficients to columns or using a column-wise operation removes the apparent signal. The answers also flag execution and validation issues: the data source supplies ask prices, which cannot be assumed available for selling, and model fitting should be separated from testing. The discussion explains the bug and practical caveats but gives no corrected out-of-sample performance evidence.

Key ideas

  • Element-wise multiplication can recycle coefficients across data-frame entries instead of matching each coefficient to a series.
  • A miscomputed spread can create a misleading stationarity result and apparent mean reversion.
  • Hedge ratios must be applied to the intended columns when constructing a basket spread.
  • Ask-only price data can misrepresent executable trading prices.
  • Separate training and testing data when evaluating a fitted cointegration strategy.

Tags

Full text
# Why does this Co-integrated basket look too good to be true?


# Why does this Co-integrated basket look too good to be true?












You need quantmod & tseries in R to run this:

```
library(quantmod)
library(tseries)

pairs <- c(
    "EUR/USD",
    "GBP/USD",
    "AUD/USD",
    "USD/CAD",
    "USD/CHF",
    "NZD/USD"
    )

name <- function(n) {
    gsub("/","",n,fixed=TRUE)
}

getFrame <- function(p) {
    result <- NULL
    as.data.frame(lapply(p, function(x) {
        if(!exists(name(x))) {
            getSymbols(x, src="oanda") 
        }
        if(is.null(result)) {
            result <- get(name(x))
        } else {
            result <- merge(result, get(name(x))) 
        }
    }))
}

isStationary <- function(frame) {
    model <- lm(frame[,1] ~ as.matrix(frame[,-1]) + 0)
    spread <- frame[,1] - rowSums(coef(model) * frame[, -1])
    results <- adf.test(spread, alternative="stationary", k=10)
    if(results$p.value < 0.05) {
			coefficients <- coef(model)
			names(coefficients) <- gsub("as.matrix.frame.......", "", names(coefficients))
			plot(spread[1:100], type = "b")
			cat("Minimum spread: ", min(spread), "\n")
			cat("Maximum spread: ", max(spread), "\n")
			cat("P-Value: ", results$p.value, "\n")
        cat("Coeficients: \n")
        print(coefficients)
    }
}

frame <- getFrame(pairs)
isStationary(frame)
```

I get FX daily data from Oanda, do a simple linear regression to find the hedging ratios, and then use the Augmented DF test to test for the P-value of mean reversion in the spread.

When I run it I get this:

```
Minimum spread:  -1.894506 
Maximum spread:  2.176735 
P-Value:  0.03781909 
Coeficients: 
        GBP.USD     AUD.USD     USD.CAD     USD.CHF     NZD.USD 
 0.59862816  0.48810239 -0.12900886  0.04337268  0.02713479
```

EUR.USD coefficient is 1.

When I plot the spread the first 100 days look like this:

Surely something must be wrong. The holy grail shouldn't be so easy to find.

Can someone help me find what is wrong?

I tried backtesting on Dukascopy with the above coefficients as lot sizes of a basket, but I run into loses. And the spread has a different order of magnitude in dukascopy. Why is that?

## Answer by Sergey (score 9, accepted)

https://quant.stackexchange.com/a/1709

The main problem in your code is this line:

`rowSums(coef(model) * frame[, -1])`

I'm not sure exactly what is does, perhaps some matrix multiplication, but definitely not what you expect it to do. Try to replace it with manual multiplication

`spread <- frame[,1] - (coef(model)[1]*frame[,2] + coef(model)[2]*frame[,3] + coef(model)[3]*frame[,4] + coef(model)[4]*frame[,5] + coef(model)[5]*frame[,6])`

And holy grail will disappear

I can see a couple of other errors as well:

- You cant sell on ASK price. With getSymbols.oanda you always get ASK

- You'd better separate testing and training data sets

## Answer by Joshua Ulrich (score 3)

https://quant.stackexchange.com/a/1715

@Sergey correctly identified the problem. The explanation is that `coef(model)` is a vector, `frame` is a data.frame, and element-by-element multiplication takes place in column-major order. The shorter vector (`coef(model)`) is recycled along the longer vector (each column in `frame`). For example:

```
frame <- data.frame(V1=1:5)
frame$V2 <- 2
frame$V3 <- 5
coef.model <- 1:3
frame * coef.model
#   V1 V2 V3
# 1  1  6 10
# 2  4  2 15
# 3  9  4  5
# 4  4  6 10
# 5 10  2 15
```

What you intended was something like this:

```
sweep(frame,2,coef.model,"*")
#   V1 V2 V3
# 1  1  4 15
# 2  2  4 15
# 3  3  4 15
# 4  4  4 15
# 5  5  4 15
```

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.