Hill Tail Estimates Can Mislead for Skewed Student-t Returns
Summary
The document examines why Hill estimates from simulated skewed Student-t samples can differ between the left and right tails even when the distribution is parameterized with a common degrees-of-freedom value. It attributes the observed discrepancy to skewness affecting extreme quantiles and notes that Hill estimation is sensitive to tail selection and sample behavior. The example compares those tail estimates with a maximum-likelihood fit of the skewed Student-t model.
It also raises the broader question of whether asset log returns have different tail exponents on each side, and points to distribution families that can represent asymmetry more flexibly. The cited empirical claim about heavier downside tails is not demonstrated with new data here. The simulation and parameter fit illustrate the issue but do not establish universal tail behavior or prove that the reported fit is robust.
Key ideas
- A Hill estimate can differ across tails in finite samples from a skewed Student-t distribution.
- Skewness changes the distribution of extreme observations and can affect empirical tail estimates.
- A model-based maximum-likelihood fit can help assess parameters, but the example does not establish estimator reliability in general.
- Separate left and right tail exponents may require a distribution that can represent asymmetric tail decay.
- Claims about financial-return tail asymmetry require empirical support beyond the displayed simulation.
Tags
Full text
# Hill Estimator for SkewStudentT is Biased
# Hill Estimator for SkewStudentT is Biased
Why Hill Estimator for Pareto Tail exponent produces incorrect results for Skewed Student T?
The true $\alpha = 4$ yet, yet it estimates left and right as 3 and 5.
Or, maybe it's other way around and SkewStudentT doesn't actually produce tails with the $\alpha = 4$ the skew distort the tails?
Code
```
import numpy as np
from arch.univariate import SkewStudent
import matplotlib.pyplot as plt
from scipy.optimize import minimize
from itertools import product
# SkewT --------------------------------
skewt = SkewStudent()
def skewt_pdf(η, λ, x):
ll = skewt.loglikelihood([η, λ], resids=np.asarray(x), sigma2=1, individual=True)
return np.exp(ll)
def skewt_cdf(η, λ, x):
return skewt.cdf(resids=x, parameters=[η, λ])
def skewt_quantile(η, λ, p):
return skewt.ppf(pits=p, parameters=[η, λ])
def skewt_sample(η, λ, n):
return skewt.simulate([η, λ])(n)
x = skewt_sample(4.0, 0.5, 10_000)
def estimate_hill(x):
x = -np.sort(-x) # sort descending
logx = np.log(x)
sumk_logx = np.cumsum(logx)
return 1 / ((sumk_logx / (np.arange(1, len(x)+1))) - logx)
# Hill ---------------------------
η, λ = 4.0, -0.5
x = skewt_sample(η, λ, 10_000)
ltail = -np.sort(x)[:500]
rtail = np.sort(x)[-500:][::-1]
la = estimate_hill(ltail)
ra = estimate_hill(rtail)
plt.plot(la, color='red', label='Left')
plt.plot(ra, color='green', label='Right')
plt.ylim(2, 6)
plt.legend()
plt.show()
# MLE ---------------------------
def fit_mle_(x, init):
x = np.asarray(x)
def nll(p):
η, λ = p
if η <= 2.01 or abs(λ) >= 1: return np.inf
pdf = skewt_pdf(η, λ, x)
if np.any(pdf <= 0) or np.isnan(pdf).any(): return np.inf
return -np.sum(np.log(pdf))
return minimize(nll, x0=init, bounds=[(2.01, 100), (-0.99, 0.99)])
def fit_mle(x, inits):
results = [fit_mle_(x, init) for init in inits]
valids = [r for r in results if r.success]
if not valids: raise RuntimeError("All MLE fits failed")
return min(valids, key=lambda r: r.fun).x
inits = list(product([3.0, 6.0], [-0.5, 0.0, 0.5]))
print("MLE:", fit_mle(x, inits)) # => MLE: [4.0401 0.5077]
```
P.S.
Is there any evidence that left and right tails of $log S_T/S_0, T \in [30...365]$ have different tail exponents for left and rights? And maybe extension to SkewStudentT that allows independent exponents for left and right tails?
## Answer by Joe King (score 4)
https://quant.stackexchange.com/a/83832
### Possible Issues with Hill Estimator for Skewed Student's T Distribution
The Hill estimator is commonly used to estimate the tail index of heavy-tailed distributions, assuming symmetric tail behavior. However, the Skewed Student T distribution, characterized by a skewness parameter $\lambda \neq 0$, exhibits asymmetric tails. For a Skewed Student T distribution with $\lambda = -0.5$ and a true tail exponent $\alpha = 4$, the left tail may appear heavier ($\hat{\alpha}_{\text{left}} \approx 3$) and the right tail lighter ($\hat{\alpha}_{\text{right}} \approx 5$). This discrepancy arises because skewness distorts the extreme quantiles, leading to different tail behaviors (Zhu & Galbraith, 2010).
### Empirical Evidence for Asymmetric Tails
> Is there any evidence that left and right tails of logST/S0,T∈[30...365] have different tail exponents for left and rights? And maybe extension to SkewStudentT that allows independent exponents for left and right tails?
Yes, financial log-returns, such as $\log S_T/S_0$, often exhibit asymmetric tail behavior. Empirical studies have shown that left tails (associated with market crashes) tend to be heavier than right tails (associated with market rallies), implying $\alpha_{\text{left}} < \alpha_{\text{right}}$ (Gabaix et al., 2006). The standard Skewed Student T distribution cannot model independent left and right tail exponents effectively. Alternatives like the Asymmetric Power Distribution (Zhu & Zinde-Walsh, 2009) or Generalised Hyperbolic distributions are more suitable for capturing such asymmetries.
### Recommendations
- Use Maximum Likelihood Estimation (MLE): MLE provides robust estimates for skewed data. For instance, the MLE estimate for the given parameters was $[4.04, 0.51]$, confirming its reliability.
### References
- Gabaix, X., Gopikrishnan, P., Plerou, V., & Stanley, H. E. (2006). A theory of power-law distributions in financial market fluctuations. Nature, 423(6937), 267–270.
- Zhu, D., & Galbraith, J. W. (2010). A tale of tails: An empirical analysis on the distribution of daily stock returns. Journal of Econometrics, 157(2), 189–205.
- Jones, M. C., & Faddy, M. J. (2003). A skew extension of the t-distribution, with an application to robust estimation of location. Journal of the Royal Statistical Society: Series B (Statistical Methodology), 65(1), 31-44.
- Branco, M. D., & Dey, D. K. (2001). A general class of multivariate skew-elliptical distributions. Journal of Multivariate Analysis, 79(1), 99-113.
### Appendix
```
import numpy as np
from arch.univariate import SkewStudent
import matplotlib.pyplot as plt
from scipy.optimize import minimize
from itertools import product
skewt = SkewStudent()
def skewt_pdf(eta, lambda_, x):
"""
Calculate the PDF for the Skewed Student T distribution.
Parameters:
eta (float): Degrees of freedom parameter for the distribution.
lambda_ (float): Skewness parameter for the distribution.
x (array-like): Array of quantiles.
Returns:
numpy.ndarray: The PDF values for the given quantiles.
"""
ll = skewt.loglikelihood([eta, lambda_], resids=np.asarray(x), sigma2=1, individual=True)
return np.exp(ll)
def skewt_sample(eta, lambda_, n):
"""
Generate random samples from the Skewed Student T distribution.
Parameters:
eta (float): Degrees of freedom parameter for the distribution.
lambda_ (float): Skewness parameter for the distribution.
n (int): Number of samples to generate.
Returns:
numpy.ndarray: Random samples from the specified Skewed Student T distribution.
"""
return skewt.simulate([eta, lambda_])(n)
def estimate_hill(x):
"""
Estimate the tail index using the Hill estimator.
Parameters:
x (array-like): Array of data points.
Returns:
numpy.ndarray: Estimated tail indices.
"""
x = -np.sort(-x) # Sort in descending order
logx = np.log(x)
sumk_logx = np.cumsum(logx)
return 1 / ((sumk_logx / (np.arange(1, len(x)+1))) - logx)
def fit_mle_(x, init):
"""
Fit the Skewed Student T distribution using Maximum Likelihood Estimation (MLE).
Parameters:
x (array-like): Array of data points.
init (list): Initial guess for the parameters [eta, lambda_].
Returns:
scipy.optimize.OptimizeResult: Result of the MLE fit.
"""
x = np.asarray(x)
def nll(p):
eta, lambda_ = p
if eta <= 2.01 or abs(lambda_) >= 1:
return np.inf
pdf = skewt_pdf(eta, lambda_, x)
if np.any(pdf <= 0) or np.isnan(pdf).any():
return np.inf
return -np.sum(np.log(pdf))
return minimize(nll, x0=init, bounds=[(2.01, 100), (-0.99, 0.99)])
def fit_mle(x, inits):
"""
Fit the Skewed Student T distribution using multiple initializations for MLE.
Parameters:
x (array-like): Array of data points.
inits (list): List of initial guesses for the parameters [eta, lambda_].
Returns:
numpy.ndarray: Best fit parameters [eta, lambda_].
"""
results = [fit_mle_(x, init) for init in inits]
valids = [r for r in results if r.success]
if not valids:
raise RuntimeError("All MLE fits failed")
return min(valids, key=lambda r: r.fun).x
# Parameters
eta, lambda_ = 4.0, -0.5
x = skewt_sample(eta, lambda_, 10000)
ltail = -np.sort(x)[:500]
rtail = np.sort(x)[-500:][::-1]
la = estimate_hill(ltail)
ra = estimate_hill(rtail)
plt.plot(la, color='red', label='Left')
plt.plot(ra, color='green', label='Right')
plt.ylim(2, 6)
plt.legend()
plt.show()
inits = list(product([3.0, 6.0], [-0.5, 0.0, 0.5]))
print("MLE:", fit_mle(x, inits))
```Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.