Out-of-Sample Backtesting for Historical VaR
Summary
The document explains why counting losses beyond Value at Risk on the same sample used to estimate the quantile cannot validate a VaR model. For a historical VaR estimate based on past profit and loss observations, the proposed approach is to use a rolling window: calculate VaR from prior data, then compare it with the next realized outcome. Repeating this process creates an out-of-sample series of VaR forecasts and violations.
The answer says violations should be assessed as binomial events and suggests a runs test to detect excessive clustering. It also warns that a model can produce too many exceptions during volatile periods, while noting that historical simulation may represent tails better than a Gaussian model. The exchange does not give detailed test formulas, significance thresholds, or a full procedure for model failure and remediation.
Key ideas
- Backtesting VaR requires comparing forecasts with outcomes that were not used to estimate them.
- A rolling historical window produces sequential out-of-sample VaR forecasts.
- Violation counts can be evaluated as binomial outcomes, and a runs test can assess clustering.
- Historical simulation can still underperform when volatility rises sharply.
Tags
Full text
# How to backtest the VaR model? # How to backtest the VaR model? I have a sorted historical P&L vector of 250 days and say, I want to calculate the 90% VaR on this distribution. I will look for the 225 element (90% * 250 = 225) and this will be my Value at Risk. Now how to back test the VaR model ? If I look for the number of days in last year that the Loss exceeded the VaR, it will always be 25 days, since by construction, it's the number which corresponds to the quantile of the VaR... And in case the VaR model is not valid, what is usually done ? Am I missing something please ? Thanks ## Answer by Richi Wa (score 1) https://quant.stackexchange.com/a/10470 you should backtest in the future. Thus you calculate your VaR based on the last 250 business days and then look at the return tomorrow. You have to do this in a rolling/sliding fashion. Your approach is in-sample and what you should do is out-of-sample. The number of violations should be binomial. Furthermore you could do a runs test to test whether your violations do not cluster too much. Having said this: the model will be ok but not good. In volatile times you will have too many violations. But it will be better than a Gaussian model as you capture the tails better.
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.