Skip to content
All library documents

Why Expected Shortfall Is Not Elicitable on Its Own

Article Quant Q&A · Author: user74093

Summary

The document raises questions about evaluating forecasts of expected shortfall (ES), a tail-risk measure, when ES alone is not elicitable. Unlike value-at-risk (VaR), ES does not have a standalone scoring loss whose expected value is minimized by the correct forecast. The questioner considers cases where a test sample has no losses beyond forecast VaR thresholds, or only some models’ thresholds are exceeded; in such cases, VaR forecasts may still be compared using the pinball loss, while the ES forecasts are difficult to rank directly.

It also distinguishes model rejection from model ranking: an ES backtest may use a p-value to reject a forecast, but that does not by itself provide a principled ordering of competing forecasts. The document points toward joint evaluation of VaR and ES, which is possible even though ES alone is not elicitable. It presents these as questions rather than supplying an answer, so it does not explain scoring rules, backtest design, or practical comparison methods.

Key ideas

  • Expected shortfall does not have a standalone scoring rule that is minimized by its correct forecast.
  • VaR forecasts can be evaluated with the pinball loss even when ES forecasts cannot be ranked directly.
  • An ES backtest p-value can support rejection without establishing a ranking among models.
  • VaR and ES are jointly elicitable and can be evaluated as a forecast pair.

Tags

Full text
# Understanding intuition behind in-elicitability "problem" of expected shortfall


# Understanding intuition behind in-elicitability "problem" of expected shortfall












Keeping related questions in mind (ES not elicitable), I am trying to understand the intuition behind the "problem" driven by the expected shortfall (ES) not being elicitable with four short questions described below.

As there exists no loss function that is minimized by the ES, as in the case of the value-at-risk (VaR), there are practical examples (to my understanding) providing modelled values for ES, which cannot be ranked based on a test dataset. These examples would be if the test dataset either shows (1) no losses beyond the underlying VaRs (as for computing ES, the VaR is required) or (2) losses beyond only some, but not all, underlying VaRs. In these cases, based on the "pin ball" loss function, the VaRs themselves can, however, be ranked.

Question #1: Is my understanding of the elicitability "problem" of ES indeed correct?

Question #2: Are there other practical examples leading to the inability to rank modelled values for ES?

Next, there are, however, backtests for ES, such that based on p-values modelled values for ES are possibly rejected based on a test dataset; however, in theory, this does not necessarily allow for ranking multiple modelled values for ES.

Question #3: Why is ranking based on p-values in this case inappropriate and would only allow for model rejection (understanding this is a more theoretical and/or fundamental statistics question, but, if possible, a short explanation suffices, because I am still struggling to wrap my head around this concept after doing some research)?

Finally, theory states that the VaR and ES are, however, joint-elicitable, such that (to my understanding) the pair of these modelled values can indeed be ranked as opposed to the ES on itself.

Question #4: Since a modelled value for VaR is required to compute the corresponding modelled ES, why would there be any "problem" in the firsts place due to the in-elicitability of ES itself, since the pair of modelled values for VaR and ES can be ranked together (unless a value for the ES is modelled directly of course)?

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.