Validating Financial Models Through Checks, Rationale, and Use
Summary
The discussion treats model validation as an ongoing process rather than a point at which a model becomes definitively correct. Suggested checks include testing simple cases and boundaries, investigating discrepancies, comparing outputs with historical outcomes, guarding against overfitting, and trying to break the model. Predictive checks can compare observed data with simulations generated under the model’s assumptions.
Validation should also examine why a model’s results make economic sense, who bears the cost of any apparent returns, and whether the mechanism can persist. For pricing and risk models calibrated to vanilla options, the answers highlight how well the model tracks later market prices, recalibration frequency and its effect on PnL, parameter interpretability and observability, and implementation speed. The appropriate standard varies by domain and intended use. Passing mathematical and software checks alone does not establish that assumptions, calibration, data, or application are suitable; evidence and expert scrutiny accumulate over time.
Key ideas
- Model validation is iterative and depends on the model’s domain and intended use.
- Test simple cases, boundaries, historical fit, predictive behavior, and resistance to overfitting.
- Simulated data and predictive checks can reveal whether model-generated outcomes resemble observed data.
- A credible return or pricing result needs an economic rationale and a plausible explanation of its persistence.
- For option-calibrated models, assess market tracking, recalibration needs, parameter qualities, and implementation speed.
Tags
Full text
# Model Validation Criteria # Model Validation Criteria Let's say I have a brand new fancy model on some asset class (calibration porcedure included over a set of vanilla options) in which I truly believe I made a step forward comparing to existing literature (a phantasmatic situation I recognize). Which criteria should be in order so I can claim that I am in a position to "validate" this model? PS: We will assume that the math are OK and that I can prove that there is no bug in the model implementation ## Answer by Andrey Taptunov (score 12) https://quant.stackexchange.com/a/187 I don't think that there is a precise point in time when we can say that model is valid (well, it's a model not a law). For example, E. Derman in his article on Model risk describes the verification of model as a iterative process: > It is impossible to avoid errors during model development, especially when they are created under trading floor duress .... So, after the model is built, the developer tests it extensively. Thereafter, other developers “play” with it too. Next, traders who depend on the model for pricing and hedging use it. Finally, it’s released to salespeople. After a suitably long period during which most wrinkles are ironed out, it’s given to appropriate clients. This slow diffusion helps eliminate many risks, slowly but steadily. He also describes 6 types possibilities which constitutes model risk here: - Incorrect model - Correct model, incorrect solution - Correct model, inappropriate use - Badly approximated solution - Software and hardware bugs - Unstable data In addition to that he provides some tips that can be used to avoid model risk and its consequences (to name just few): - Test complex models in simple cases first - Test the model’s boundaries - Don’t ignore small discrepancies ## Answer by vonjd (score 8) https://quant.stackexchange.com/a/10242 The most important questions in my opinion (besides all the mathematical questions and robustness tests that have to be performed anyway) are: - Why is this model working in the first place? - What is the economic rationale behind it? - Why have the returns not been arbitraged away? - Who is paying for your returns and why? - Is this situation likely to be sustainable? When you cannot satisfactorily answers these questions chances are that your results are too good to be true. ## Answer by John (score 5) https://quant.stackexchange.com/a/10311 I agree with many of the other comments and answers. In addition, I would recommend the paper by Gelman and Shalizi on "Philosophy and the practice of Bayesian Statistics", even if you're not a Bayesian. I would emphasize Section 4 on Model checking. They note the importance of posterior predictive checks, i.e. if I forecast from a model, would I get sensible results. They also recommend simulating fake data assuming the model is true and comparing it to the actual data. I regularly use these tools. ## Answer by Sason Torosean (score 4) https://quant.stackexchange.com/a/10244 Not knowing what type of model you have, I am going to guess that these general steps will make the user of the model feel confident/comfortable about using your model. In my opinion, you have shown that your model does reasonably well: 1). If you input historical data and see if the output is reasonably close to what the real historical output was. 2). Not fall into the trap of over-fitting. 3). Not use obscure reasons for calibration i.e there should be at least some mathematical justification. 4). The simple test cases are the most overlooked, make sure they are passed successfully. 5). Try to break it and explain why it did. ## Answer by Arshdeep (score 4) https://quant.stackexchange.com/a/54965 If the model you're talking about is something that prices and risk manages an exotic (since you mentioned you calibrated to vanillas), I'd like to see: - How does the evolution of the volatility surface / correlations look like. When I vega hedge, my ability to recover future prices of vanillas is important for me to not leak PnL by rebalancing/recalibrating. - How often does your model need to be 'recalibrated' to hit the market price over the life of the exotic? This is again problematic due to PnL leakage. - Since admittedly your model can price and risk manage things better, can you explain what features make it better than the second best model for the instrument? -Parameter markings - If your model needs to mark parameters that are difficult to risk manage, lack intuition, or simply cannot be marked independently of others (i.e. say when one feature of the market changes, you need to change the whole family of markings to accommodate it, so you don't really know what parameter is doing which job. Moreover, are these parameters market observable? Or can I estimate them historically? - Implementation speed of course is important. ## Answer by RndmSymbl (score 2) https://quant.stackexchange.com/a/10265 What constitutes valid is often - a part from model risk discussed above - an agreement between subject matter experts in a specific field or company. And hence, different fields require different standards. I used to work in social science, and top journals would find things perfectly valid if Crombach Alpha, Power, R-Square, F-score and factor significance on the regressions were within reasonable ranges. Within risk management the bar is much higher, there is even a journal on risk model validation. And if you follow the "Cutting Edge" section in Risk you will find that academics even within one field may occasionally disagree on what is considered valid - re FVA. To cite an other example, spatial econometrics faces different criteria in model validation.
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.