Choosing Loss Functions to Evaluate Volatility Forecasts
Summary
Volatility forecast evaluation depends on what quantity is forecast and what observation serves as the benchmark. Possible targets include volatility, variance, or log volatility; because true volatility is latent, evaluation may use implied volatility or a realized measure, each with its own limitations. The choice of penalty also matters: errors can be measured with L1 or L2 losses, in levels or percentages.
The discussion identifies mean absolute error and a squared log error as examples, and notes that studies often combine L1 and L2 measures with a Mincer–Zarnowitz regression. It also mentions Diebold–Mariano tests for comparing forecasts. In practical settings, L1 error on volatility, volatility percentages, or vega-scaled volatility is reported. The document does not establish a single best metric or provide a systematic comparison of their performance; selection should reflect the model, application, target, and available benchmark.
Key ideas
- Volatility forecast evaluation may target volatility, variance, or log volatility.
- Because volatility is latent, the chosen proxy—such as implied or realized volatility—affects the evaluation.
- L1 and L2 losses can be applied in levels or percentages, with different interpretations.
- Mincer–Zarnowitz regressions and Diebold–Mariano tests offer complementary evaluation tools.
- No single loss function is best for every model and application.
Tags
Full text
# evaluation of volatility models using loss functions
# evaluation of volatility models using loss functions
This question has two parts, What is the state of the art in an academic or public knowledge sense of volatility forecast model evaluation?
Since there are many methods out in the wild, and do correct me if I am wrong but I haven't been able to find a complete list or any list for that matter of methods used so I'd like to start that here. It is very popular to use many of the different loss functions with arma-garch style models.
Some methods commonly used include:
Mean Absolute Error
$MAE = n^{-1} \sum_{t=1}^n | \sigma_t - h_t|$
$ \boldsymbol{R^2 \ log}$
$R^2LOG = n^{-1} \sum_{t=1}^n (log(\sigma_t^2 h_t^{-2}))^2 $
## Answer by onlyvix.blogspot.com (score 1)
https://quant.stackexchange.com/a/9246
In Forecasting Financial Market Volatility Ser-Huang Poon dedicated entire chapter to the question, so the issue is far from simple. I don't believe there one single best way because of many questions that depend on model form and application such as
Should one evaluate volatility or variance, or perhaps ln(vol)? What is the benchmark - volatility is latent, should one use implied, or realized, and which realized measure should be used? What penalty should be used? L1 / L2? Points or percents? In addition there are specific tests to compare one model again another, s.a. Diebold & Mariano tests.
Usually what I see in academic literature is several metrics, L1 and L2 norm on volatility, as well as standard Mincer-Zarnowitz regression. In practical applications I most often see L1 on volatility, volatility %, or Vega * volatility.Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.