Evaluating Intraday Volatility Estimates Without a Known True Value
Summary
The document addresses how to assess a realized-volatility estimator when the underlying day's true volatility is unknown. It argues that financial volatility does not provide a uniquely observable ground truth: the value depends on the model, including whether it allows jumps, and one-minute observations are affected by market microstructure noise. A conventional mean squared error against a known target therefore cannot simply be computed from the same observed day.
Instead, the response frames estimator evaluation as a model and prediction problem. It recommends testing predictive capability and controlling for key intraday features before comparing procedures: seasonal patterns, volatility clustering, and path dependence. One suggested approach is to remove seasonality and compare estimators within clusters, though the answer does not specify a complete protocol or a universal benchmark. Its guidance is conceptual, and conclusions will depend on the target definition, data frequency, and assumptions used in evaluation.
Key ideas
- There is no model-independent, directly observed true volatility value for a trading day.
- The chosen structural model, including its treatment of jumps, changes what volatility means.
- One-minute data can contain market microstructure noise that affects realized-volatility estimates.
- Intraday seasonality, clustering, and path dependence complicate estimator comparisons.
- Evaluation can focus on predictive ability and compare methods after accounting for these features.
Tags
Full text
# What is the realized volatility's estimation error? # What is the realized volatility's estimation error? Given an estimation procedure and real data, how would one compute the mean squared error? What value represents the "true" realized volatility in the case of calculating the Mean Squared Error in estimation? I'm specifically interested in intraday estimation error (one minute trade data for example)? So for example: - I want to estimate the realized volatility statistic on 1 day's worth of real 1 minute data - I'm going to use block bootstrapping as my estimation procedure - I run the block bootstrapping estimation algorithm and get an estimated realized volatility value - To compute the estimation error, I need a true realized volatility value for this 1 day of 1 minute data. What can I use for this "true" realized volatility value for the specific data that I am estimating from? ## Answer by lehalle (score 3, accepted) https://quant.stackexchange.com/a/9930 Unfortunately, financial markets are not like physical measures, where you know the "true" value of a physical variable but you just access to it thanks to noised sensors. We do not know the "true" volatility, just because there is not such one value... In statistics you have two kinds of modelling procedures: - the ones dedicated to estimate the unknown values of parameters of a known structural model. Here you have the usual "confidence intervals" approach. - the ones for simultaneously inferring from data the shape (i.e. "class") of the model and its parameters. Here you have the usual issues coming from "overfitting", and the usual approaches are "cross validation", "regularization - penalization", "Vapnik-Chervonenkis dimensions", etc. In finance you are very often in the second case, and especially for volatility: for instance its value is not the same if your underlying model includes jumps or not... what is the "true" one? Moreover (as I commented) at your time scale (1 min) you face the microstructure noise, see my answer here for a brief. But come back to my generic answer on "knowledge discovery" via statistical modelling: what can you test if you believe you have a good new estimation procedure? You can test your prediction capability of course, but you will have to face a lot of ugly features of intraday volatility: - it is not iid, even not ergodic, since it has a seasonality (see Market Microstructure in Practice for intraday seasonalities). - once you removed the seasonality, it is clustered (we cannot ignore it since Robert Engle's Nobel prize). - moreover, it is path dependent (I mean, even inside a volatility cluster as defined by Engle)... Thus if you wan to challenge existing estimation procedure, you will have to remove the first two features and demonstrate on the deseasonalized data, cluster by cluster. Of course you could alternatively try to perform a change of state space to estimate something else than volatility. Like use it to estimate the seasonality or the switching probabilities themselves...
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.