Estimating the Hurst Exponent from Lagged Price Differences
Summary
The document discusses a log-log regression method for estimating the Hurst exponent from a price series. For each lag, it calculates differences between log prices separated by that lag, measures their variance or a related dispersion statistic, and fits a line against the logarithm of the lag. Under the scaling relation described, the slope is related to the Hurst exponent, with the conversion depending on whether variance, standard deviation, or a transformed statistic is used.
The answers offer competing explanations of a Python implementation’s square-root-of-standard-deviation step and note that an alternative variance-based implementation produces the same estimate after changing the slope conversion. The discussion helps connect the code to scaling behavior, but it does not fully settle the derivation of the original snippet. Results also depend on the precise statistic computed, the sampled lags, and the time series, so the code’s scaling convention should be checked before interpreting an estimate as evidence of mean reversion or momentum.
Key ideas
- The Hurst method estimates scaling by regressing a dispersion measure of lagged log-price differences against lag size on logarithmic axes.
- Variance scaling is related to the Hurst exponent through an exponent of twice its value.
- Using variance, standard deviation, or a square-root transformation changes the slope conversion.
- The discussion gives differing accounts of the original code’s transformation, so its exact derivation remains uncertain.
- An estimated Hurst exponent depends on the chosen lags, statistic, and data.
Tags
Full text
# Explanation of Standard Method Generalized Hurst Exponent
# Explanation of Standard Method Generalized Hurst Exponent
Apologies if this question is vague, I've gone over how to word it several times in my head, and I'm not sure it gets clearer each time.
I've been looking at this website article https://www.quantstart.com/articles/Basics-of-Statistical-Mean-Reversion-Testing and have been investigating the code for the Hurst exponent in Python. The article gives a code snippet of Python as follows (to calculate Hurst):
```
def hurst(ts):
"""Returns the Hurst Exponent of the time series vector ts"""
# Create the range of lag values
lags = range(2, 100)
# Calculate the array of the variances of the lagged differences
tau = [sqrt(std(subtract(ts[lag:], ts[:-lag]))) for lag in lags]
# Use a linear fit to estimate the Hurst Exponent
poly = polyfit(log(lags), log(tau), 1)
# Return the Hurst exponent from the polyfit output
return poly[0]*2.0
```
Which works great, but because of my annoying personality where I need to understand something before I use it, I've been driving myself nuts for a day and a half trying to understand how this code has been developed/derived (especially the sqrt(std) part). The article itself does have some brief steps, but I'm not able to follow them. It may be that I don't quite understand the <| ..... |> notation means and how it can relate to standard deviation. Attached here:
Can anyone provide a link to an article, website or paper that shows from what principles this calculation could have been derived? From Racine's paper I'm aware that Hurst's original method was the RS method, but I believe the method used in the code is from the generalized Hurst exponent or Standard method.
My mathematics isn't at pHd level, but I do have an Engineering degree, so it's not totally useless either. What I am having a huge problem understanding is why the code uses the square root of standard deviation, so if anyone could shed some light on that, it would be greatly appreciated.
Thanks for your time, apologies again if this isn't totally clear.
## Answer by vanguard2k (score 2, accepted)
https://quant.stackexchange.com/a/35516
I am unaware of the notation in this case but I am still trying to make sense of it. (maybe someone can jump in with an edit?) Here's what I have got:
There is the notion of quadratic variation out there that could apply here. You can also think about it as a scalar product (which is the same). You can see that the notation here is rather sloppy, because $\text{log}(t)$ denotes the log price process at time $t$. For discrete points in time, you can interpret it as the vector of log prices at the times $t_i$ and treat the bracket as a scalar product (now denoting $\log(t+\tau)$ and $\log(t)$ as a vector): $$ \langle (\text{log}(t+\tau)-\text{log}(t))^2\rangle \approx \langle \text{log}(t+\tau)-\text{log}(t),\text{log}(t+\tau)-\text{log}(t)\rangle $$
Now, I try to switch to a more precise notation.
Let $t_i, i = 1,\ldots,T$ be the discrete timestamps and $P_i$ the price at $t_i$. Then the term above means the following:
$$ \left(\sum_{i=1}^{T-\tau} (\log(P_{i+\tau})-\log(P_i))^2\right)^{1/2}. $$
Now to determine the hurst exponent $H$, we say that this is approximately $\tau^{2H}$:
$$ \left(\sum_{i=1}^{T-\tau} (\log(P_{i+\tau})-\log(P_i))^2\right)^{1/2} \approx \tau^{2H}$$
taking the logarithm on both sides gives: $$ \log(\ldots) \approx 2H \log(\tau)$$
So what the algorithmdoes is, take 99 values for $\tau$ and calculate the quadratic variation of the lagged differences. Then regress those onto $\tau$ to get an approximation for $2H$.
Now, if we come back to the code: I dont know what the polyfit function does but I assume it performs a polynomial regression with the degree as a parameter. What I dont get is the last line where the result is multiplied by two but you should be able to verify this if you clarify the function output (what is regressed on what, what coefficients are saved where).
## Answer by Wei (score 3)
https://quant.stackexchange.com/a/35536
@vanguard2k I haven't exactly been able to satisfy myself with the derivation of the original code, but I have been able to do the next best thing. I've looked at the source of the code, which is QuantStart which credits Dr Tom Starke, which uses a slightly different code and also credits Dr Ernie Chan. I've then gone to Dr Chan's blog and used his principles to come up with my own code. It uses variance instead of std, gets rid of the sqrt, and uses a 2.0 divisor instead of a 2.0 multiplier (which is the 1/4 that you mention). And the results it gives are the same as the original code. The difference is I understand (I think) the principles that are behind the code below and the final answer is the same. Thanks alot for all your help troubleshooting with me, in the end you were pretty much right, all the elements that didn't seem to make sense cancelled each other out in the end (I guess).
http://epchan.blogspot.fr/2016/04/mean-reversion-momentum-and-volatility.html
- Dr Chan states that if z is the log price, then volatility, sampled at intervals of τ, is volatility(τ)=√(Var(z(t)-z(t-τ))). To me another way of describing volatility is standard deviation, so std(τ)=√(Var(z(t)-z(t-τ)))
- std is just the root of variance so var(τ)=(Var(z(t)-z(t-τ)))
- Hence (Var(z(t)-z(t-τ))) ∝ τ^(2H)
- Taking the log of each side we get log (Var(z(t)-z(t-τ))) ∝ 2H log τ
- [ log (Var(z(t)-z(t-τ))) / log τ ] / 2 ∝ H (gives the Hurst exponent) where we know the term in square brackets on far left is the slope of a log-log plot of tau and a corresponding set of variances.
lags = range(2,100)
```
def hurst_ernie_chan(p):
variancetau = []; tau = []
for lag in lags:
# Write the different lags into a vector to compute a set of tau or lags
tau.append(lag)
# Compute the log returns on all days, then compute the variance on the difference in log returns
# call this pp or the price difference
pp = subtract(p[lag:], p[:-lag])
variancetau.append(var(pp))
# we now have a set of tau or lags and a corresponding set of variances.
#print tau
#print variancetau
# plot the log of those variance against the log of tau and get the slope
m = polyfit(log10(tau),log10(variancetau),1)
hurst = m[0] / 2
return hurst
```
## Answer by J.Melody (score 2)
https://quant.stackexchange.com/a/46652
I think the given python code snippet is composed according to the following steps:
\begin{equation} var(\tau) = \left< |z(t+\tau)-z(t)|^2 \right> \thicksim \tau^{2H} \end{equation}
\begin{equation} \Rightarrow var(...)\thicksim \tau^{2H} \end{equation}
\begin{equation} \Rightarrow std(...) \thicksim \tau^{H} \end{equation}
\begin{equation} \Rightarrow sqrt(std(...)) \thicksim \tau^{\frac H 2} \end{equation}
\begin{equation} \Rightarrow log(sqrt(std(...)) \thicksim \frac H 2 log(\tau) \end{equation}
thus, after applying ployfit(), we get the result H which is poly[0]*2.0
## Answer by sen_saven (score 1)
https://quant.stackexchange.com/a/35515
I fail to understand how this code works too.
In the past, I have used this matlab implementation for the generalized hurst exponent calculation and it was quite reliable.
Recently someone has translated this into python, but I haven't tested this yet.
Hope that helps.Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.