Robust Volatility Estimation for Small Samples
Summary
The discussion considers alternatives to estimating daily stock volatility with the standard deviation of close-to-close log returns, which can be strongly affected by outliers in a small sample. Suggested robust dispersion measures include the interquartile range and semi-interquartile range; these reduce sensitivity to extremes, though standard deviation remains preferable under a normal distribution according to one answer.
Other responses recommend using price-range estimators such as Yang–Zhang, which incorporates open, high, low, and close data, and Parkinson, which uses high and low prices. Rank or indicator transforms are also proposed to limit outlier influence. For scarce observations, one answer suggests a regression-based meta-estimate using related variables such as time of day to expand the information available. These approaches involve tradeoffs: robustness or additional price information does not guarantee superiority in every setting, and model-based estimates may lose properties such as unbiasedness and require empirical validation. The discussion offers suggestions rather than comparative test results.
Key ideas
- Interquartile and semi-interquartile ranges can reduce sensitivity to return outliers.
- Standard deviation is still favored when returns are normally distributed.
- Range-based estimators can use high and low prices, or open, high, low, and close data.
- Transforms and regression-based estimates are alternatives, but their performance needs empirical assessment.
Tags
Full text
# better estimator of volatility for small samples
# better estimator of volatility for small samples
One commonly used sample estimator of volatility is the standard deviation of the log returns.
It is indeed a very good estimator (unbiased, ...) when the sample is large.
But I don't like it for small sample as it tends to overweight outliers in log returns.
Do you know if any other statistical dispersion measure that can be use to estimate the volatility of a stock? (I don't care about statistical properties; I just want it to estimate differently / better the daily risk of this stock.)
PS: I have already tried to use the norm 1 instead of the Euclidean norm. Any other idea / remark?
## Answer by Ralph Winters (score 5, accepted)
https://quant.stackexchange.com/a/872
You could use something like the interquartile or semi-interquartile range, which is somewhat more insensitive to extremes. This is a better measure to use if your data is skewed, however if your data is normally distributed it is still better to use the standard deviation.
http://en.wikipedia.org/wiki/Interquartile_range
## Answer by ast4 (score 7)
https://quant.stackexchange.com/a/874
You mention "daily" risk, so I'm assuming you're looking at a daily frequency. Yang-Zhang Volatility (Drift-independent Volatility Estimation Based on High, Low, Open and Close Prices) fits the bill for what you're asking, it takes into account intraday fluctuations as well.
## Answer by user98 (score 1)
https://quant.stackexchange.com/a/876
In order to suppress the effect of outliers, you can use indicator transform or rank transform. Once you convert your data in those form then you can find the volatility. Personally, I like indicator transform as it is easy, more popular in many applications and gives good results.
## Answer by JoseOrtiz3 (score 1)
https://quant.stackexchange.com/a/45346
There are waaaaayyy better estimators than $Var(log(Close_{t+1}/Close_{t}))$. This close-close estimator is unbiased, and has a data efficiency defined as $1$.
The Parkinson estimator uses high and low prices only (useful when you don't trust your open and close prices, or don't have them). It has a data efficiency of about $4$. The expected variance from "true" is four times smaller than close-close.
The Yang-Zhang estimator use open, high, low, and close prices, and has a data efficiency of about $14$.
Many other estimators exist with different benefits and drawbacks. Nice math-based discussion of the different approaches in the Yang-Zhang paper.
However there is no magic bullet.
For small sample size, you are going to want to artificially inflate your sample size using machine learning hacks. You can turn it into a regression problem, where you think of things that generally should correlate with volatility (time of day, for instance), and use a regressor to estimate the estimate (meta-estimate?), which essentially means estimating the volatility using more samples than what you previously were limited to. However you will lose mathematical promises like unbiased-ness, and quality of estimate will need to be checked empirically.Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.