Skip to content
All library documents

Estimating Mean-Reversion Half-Life from a Discrete-Time Regression

Article Quant Q&A · Author: Vladimir Belik

Summary

The document explains why the shortcut half-life formula can give a misleading result when applied to a regression of changes on lagged levels. In a discrete autoregressive process, the regression slope on the lagged level is the autoregressive coefficient minus one, so the exact half-life uses the logarithm of one plus that slope. The simpler formula that divides by the slope approximates this expression when the slope is close to zero.

A simulated mean-zero Ornstein–Uhlenbeck process illustrates the correction: the shortcut gives an estimate notably different from the process's known half-life, while the logarithmic formula recovers a value close to it. The post also explains why a sine wave is a poor demonstration of simple mean reversion: its strong momentum and periodic reversals do not behave like the fitted OU process. The derivation assumes a stable process with a positive autoregressive coefficient below one; applying it outside that range can produce invalid or non-mean-reverting interpretations.

Key ideas

  • For a regression of changes on lagged levels, the slope equals the autoregressive coefficient minus one.
  • The discrete-time half-life is calculated from the logarithm of the autoregressive coefficient.
  • The slope-based shortcut is an approximation that works best when the slope is near zero.
  • A cyclical series with strong momentum may not be well described by a simple OU mean-reversion model.
  • A simulated process with known parameters can help check a half-life estimation method.

Tags

Full text
# Why is my mean-reversion half-life completely wrong?


# Why is my mean-reversion half-life completely wrong?












I am using a couple of resources (here and here) to calculate the mean reversion half-life of a time series. This method of calculating it is also presented in Ernest Chan's Algorithmic Trading on page 47.

The method, as I understand it, is to use the formula: halflife = -log(2)/G

Where G is the coefficient of a linear regression where the dependent variable is the differenced series (Yt - Y(t-1)) and the dependent variable is the lagged series (Y(t-1)).

Here is my code:

```
a <- sin(0.2*seq(1, 100, 0.25))
y <- diff(a)
x <- a[1:(length(a)-1)]
lm <- lm(y~x)
halflife <- -log(2)/lm$coefficients[2]
```

I wanted to try it out with a clearly cyclical series, so I used a sine function. Here is a graph of the function:

The result I get, however, is: -1064 .

How is this possible? First of all, the value should be positive. This means that my coefficient is not negative, which suggests the series is not mean reverting. But it clearly is, right?

I'd greatly appreciate any assistance with understanding what's going on here.

## Answer by Jamie Ballingall (score 3)

https://quant.stackexchange.com/a/70417

I think the main problem is that, as @KurtG says, your example sine function exhibits strong momentum (albeit with reversals) and fitting an Ornstein-Uhlenbeck process to it results in strange parameters.

However, it might be nice to check that your code is correct. To do so, let's simulate a process that we know to be OU with an obvious half-life:

```
e <- rnorm(n, 0, 1)
a <- array(data = NA, dim = n)
a[1] = 0
for (i in 2:n) { a[i] <- 0.5 * a[i - 1] + 0.1 * e[i] }
```

Now `a` contains a large sample from an OU process with mean 0 and half-life of 1.

We run your fitting code:

```
y <- diff(a)
x <- a[1:(length(a)-1)]
lm <- lm(y~x)
halflife <- -log(2)/lm$coefficients[2]
```

and get a halflife of (with my random numbers) about 1.38. Which is close, but wrong and too far from the true value to be explained by just randomness because of the large sample size.

Let's revisit the formulas for halflife. First, and easiest to derive, if we have the OU process (with mean zero) $$ dX_t = -\theta X_t dt + \sigma dW_t $$ with $\theta > 0$ and $\sigma > 0$ then the half-life is $$ h = \frac{\log \left( 2 \right)}{\theta} $$ If we write this in discrete time we have: $$ X_t = \phi X_{t-1} + \epsilon $$ with $0 < \phi < 1$. Then the half-life is $$ h = -\frac{\log \left( 2 \right)}{\log \left( \phi \right)} $$ But your regression uses the differenced form of: $$ X_t - X_{t-1} = \left(\phi - 1 \right) X_{t-1} + \epsilon $$ We can define $\beta = \phi - 1$ (and hence $\phi = 1 + \beta$) to get: $$ X_t - X_{t-1} = \beta X_{t-1} + \epsilon $$ which is the regression you are running. Specifically, `lm$coefficients[2]` is giving you $\beta$. Rewriting a previous formula we get: $$ h = -\frac{\log \left( 2 \right)}{\log \left( 1 + \beta \right)} $$ If we try that:

```
halflife <- -log(2)/log(1 + lm$coefficients[2])
```

we get a number that is almost exactly 1, as expected. So that formula is probably the best one to use.

Where did your original formula come from? Well, the function $\beta \mapsto \log \left( 1 + \beta \right)$ is approximately linear near zero, so replacing $\log \left( 1 + \beta \right)$ with $\beta$ is a reasonable approximation. Nevertheless, using the more exact formula is probably worth it most of the time since it isn't expensive to evaluate.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.