Skip to content
All library documents

Hurst Exponent Estimation: R/S Analysis Versus Lag-Scaling

Article Quant Q&A · Author: lalalal

Summary

The document asks why a rescaled-range (R/S) description of Hurst exponent estimation differs from a Python example. The described R/S approach divides a series into blocks at several sizes, centers each block by its mean, measures the cumulative-deviation range, scales that range by the block's standard deviation, and averages across blocks. It then estimates the exponent from how this statistic changes with block size.

The code instead computes standard deviations of lagged price differences over a range of lags, plots their logarithms against log lag, and doubles the fitted slope. These are distinct estimation approaches, so their calculations and assumptions do not match term by term. The document provides the formula descriptions and code as evidence but contains no answer evaluating their statistical properties. It also gives no data-based comparison, so it does not establish which estimate is more reliable or how choices such as price levels versus returns affect results.

Key ideas

  • R/S analysis estimates scaling by comparing averaged rescaled ranges across block sizes.
  • The code estimates scaling from the standard deviation of lagged differences.
  • The code doubles the log-log slope to obtain its Hurst estimate.
  • The document poses the methodological distinction but does not resolve estimator reliability.

Tags

Full text
# Why does the Hurst exponent pseudo code not match the Python implementation?


# Why does the Hurst exponent pseudo code not match the Python implementation?












I am working on understanding the Hurst exponent calculation by Ernest Chan; however, the description of the algorithm does not match the Python implementation.

Chan [Algorithmic Trading: Winning Strategies and Their Rationale] outlines following steps for Hurst Exponent computation using R/S method:

▪ Divides the time series of length 𝑁 into 2𝑘 (𝑘 = 0,1,2…) adjacent subperiods of the same length 𝑛, such that 2𝑘 × 𝑛 = 𝑁.

▪ For each sub period of length 𝑛, a partial time series is created based on various subperiods, 𝑋 = 𝑋1,𝑋2,…𝑋𝑛.

▪ Arithmetic mean (𝑥̅ = 1 𝑛 ∑ 𝑋𝑖 𝑛 𝑖=1 ) is calculated and a new mean-adjusted series (𝑌 𝑡 = 𝑥𝑡 − 𝑥̅) column is constructed.

▪ Cumulated deviations from the arithmetic mean are noted in a new series (𝑍𝑡 = ∑ 𝑌 𝑡 𝑡 𝑡=1 ).

▪ An adjusted range is calculated by 𝑅(𝑛) = max(𝑍1,𝑍2,…𝑍𝑛) − min (𝑍1,𝑍2,…𝑍𝑛)

▪ Standard deviation is computed 𝑆(𝑛) = √1 𝑛 ∑ (𝑋𝑖 − 𝑥̅)2 𝑛 𝑖=1

▪ Each adjusted range 𝑅(𝑛) is then standardized by corresponding standard deviation 𝑆(𝑛), to form the rescaled range 𝑅(𝑛)/𝑆(𝑛)

▪ An average rescaled range over all the partial time series of length 𝑛 is calculated The process is repeated iteratively using 𝑘 = 0,1,2… for each length 𝑛 = 𝑁/2𝑘 of subperiods.

```
from numpy import *
from pylab import plot, show
# first, create an arbitrary time series, ts
ts = [0]
for i in range(1,100000):
    ts.append(ts[i-1]*1.0 + random.randn())
# calculate standard deviation of differenced series using various lags
lags = range(2, 20)
tau = [sqrt(std(subtract(ts[lag:], ts[:-lag]))) for lag in lags]
# plot on log-log scale
plot(log(lags), log(tau)); show()
# calculate Hurst as slope of log-log plot
m = polyfit(log(lags), log(tau), 1)
hurst = m[0]*2.0
print 'hurst = ',hurst
```

Could someone please explain the difference of the description and the implementation in case my math is not good enough? Or are these different ways to implement the Hurst expoenent?

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.