Evaluating VaR Backtests Across Full and Subperiod Samples
Summary
The document asks whether a VaR model’s testing window can be divided into consecutive subperiods to compare its performance in volatile and calmer market conditions. In the example, the overall test slightly understates risk, while the earlier period has more violations and the later period’s count is close to expectation. The author also asks whether gaps between subperiods are needed to ensure independence.
The answer recommends examining both the full backtest and subperiod results. A subperiod can expose when a model performs poorly, but a favorable violation count in another period does not explain the difference by itself. The analyst should investigate what changed and whether indicators can identify conditions that call for a different risk estimate. The response does not prescribe a formal test for comparing periods or resolve statistical dependence concerns; it offers guidance on interpreting results and explaining model failures.
Key ideas
- Report results for the full backtest alongside results for selected subperiods.
- Subperiod analysis can reveal periods when a VaR model’s estimates were less reliable.
- Explain failures by investigating market conditions rather than treating a severe decline alone as an excuse.
- The document does not specify a formal method for testing independence between adjacent periods.
Tags
Full text
# Can I split my backtesting into multiple consecutive sub-periods? # Can I split my backtesting into multiple consecutive sub-periods? I'm testing a model to estimate the VaR of a portfolio with different stocks. I used 1500 data to estimate some parameters, and now I have other 1500 data for backtesting purposes (for a total of 3000 data on return series). The model slightly underestimates VaR, i.e. the number of violations exceeds a bit the expected number of violations. However, if I split the 1500-days testing window in two 750-days periods, I get moderate violations in the first testing period (which spans approximately from 2007 to 2009) and a 'perfect performance' in the second 750-days period (number of violations matches the expected number of violations almost exactly). Therefore, I wanted to consider this splitting procedure in order to highlight the model's advantages and drawbacks in volatile vs. more stable times (performing unconditional and conditional coverage tests separately for two periods). Question: Are there problems related to the idea of splitting the testing window in two consecutive periods? More precisely: - Is it okay to consider period 1 from $T$ to $T+749$ and period 2 from $T+750$ to $T+1499$? - Or is it better to allow for "some space", i.e. consider non-immediately consecutive periods, such as from $T$ to $T+499$ for period 1 and then from $T+750$ to $T+1249$ for period 2 (with 500 data each)? I guess my question is: are there any issues related to "independence" between consecutive testing windows, or can I continue with my approach of two consecutive periods? ## Answer by SRKX (score 0, accepted) https://quant.stackexchange.com/a/27949 Ideally you would like to look at both the global backtested period and sub-periods as well, there is nothing wrong with that. No backtesting framework is perfect and no risk ex-ante estimate is perfect either. So you can look at the results over the global period, which conclude that your approach is decent, and then highlight that it particularly didn't work in a given period. Then, your job is to explain what happened in that period and why your estimate failed. You could claim that there was a very specific event such as government intervention, votes results, etc... Obviously, stating that "market went down very severely that period" is not a good excuse - you're trying to predict risk - but you could claim that market condition change from time to time, and then try to find indicators which detect these changes, and which will allow you to change you estimate to another one more specific to these circumstances.
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.