Choosing Realized Volatility Proxies to Evaluate Forecasts
Summary
The document explains why realized volatility used to evaluate forecasts is not directly observable. Researchers must choose a proxy, such as one derived from returns, and recognize that the resulting measure depends on assumptions about prices, sampling, and the estimator. Even a familiar close-to-close calculation is only one possible estimate.
The example compares a forecast based on returns before the forecast date with a future-window estimate, while the answer emphasizes that evaluation depends on an agreed definition of the target. Alternative estimators can produce different realized volatility values, so forecast assessments may vary with the proxy. The discussion also relates latent volatility to stochastic-volatility models. It gives conceptual guidance and references research and textbooks, but does not prescribe one universally correct estimator or quantify how choices affect forecast rankings.
Key ideas
- True return variance is latent, so forecast evaluation relies on an estimated proxy.
- Price selection and return construction involve assumptions that affect the volatility estimate.
- Close-to-close volatility is common for simplicity, but other estimators can produce different values.
- Teams should define the realized-volatility target and its estimator before comparing forecasts.
- Stochastic-volatility models represent variance as an unobserved part of return dynamics.
Tags
Full text
# On the correct definition of ex-post variables
# On the correct definition of ex-post variables
I recently had an interview with a quant from a hedge fund and we were discussing about properly defining ex-post variables when backtesting the forecasting ability of certain market variables.
For example, if I were to forecast the future realized volatility over the next month using historical volatility:
- For the Ex-Ante/Forecast: I would use the daily returns (in business days) of t-20 to t (for a total of 21 business days) to deduce daily volatility and multiply it by $\sqrt{21}$ to deduce monthly volatility.
- For the Ex-Post/Realized: I would use the daily returns (in business days) of t+1 to t+21 (for a total of 21 business days) to deduce daily volatility and multiply it by $\sqrt{21}$ to deduce monthly volatility.
I have always taken the above to be correct, but what the quant mentioned about ex-post volatility being subjective and having to be agreed upon by a team seems to make sense too.
Anybody has any comments/thoughts? I wish to discuss correlation too, and a solution to this question would answer it as well.
## Answer by AKdemy (score 4, accepted)
https://quant.stackexchange.com/a/84140
Historical (also called realized) volatility is inherently unknown and not directly observable. Therefore, you need to agree on the assumptions and understand the limitations of the estimate you choose to use before assessing a forecast.
See, for example, Ruey Tsay, Analysis of Financial Time Series) (P.110 3rd edition, the link brings you directly to the relevant section) or Hull and many other textbooks.
Pretty much any statistician working with volatility mentions this feature of volatility somewhere in their papers (frequently using the term latent) because of the challenge of having to rely on a (noisy) proxy to assess forecast quality. A well-cited paper by Andersen, Diebold et al. can be found on duke.edu. It was published in December 2006 in The Handbook of Economic Forecasting, Vol. 1, pp. 777–878. I'll copy-paste the relevant section:
> "As discussed at some length in Sections 1 and 5, the 'true' variance, or volatility, is inherently unobservable, and we are faced with the challenge of having to rely on a proxy in order to assess the forecast."
There are also several frequently used estimators. If you look at returns, you do not observe them. You need to compute them and in doing so, you need assumptions. E.g., take the mid of bid and ask, decide to use the close price each day, and compute the log difference between consecutive trading days. This is not directly observable but rather something one decides to do with observables.
Now, using already derived returns data, you apply some statistic of your choice to estimate volatility (it's in fact called an estimate in the literature for the very reason that vol itself is unobservable, see R for some examples).
What you have in mind is most likely the close-to-close method (computing the annualized standard deviation of log returns). As already mentioned, there are numerous others. Each will yield a different number for "volatility." Most are theoretically better than the simple close-to-close number, which is so commonly used because it's easy to compute, and easy to grasp.
An interesting side remark: this latency is also an aspect that makes stochastic vol (SV) models appealing, because they explicitly include an unobserved (non-measurable) shock to return variance in the volatility dynamics. Therefore, variance becomes inherently latent, implying that the volatility process is not measurable with respect to observable (past) information.Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.