In-Sample and Out-of-Sample Tests for Option Model Calibration
Summary
The document asks how researchers should compare Lévy and stochastic-volatility models calibrated to vanilla option prices, especially when the models will be used to price exotics such as arithmetic Asian options. It distinguishes two purposes: fitting a volatility surface and producing prices when exotic-market quotes are unavailable. The author questions whether model comparisons should hold out some vanilla data for evaluation, instead of assessing fit on the same quotes used in calibration.
The text provides no empirical comparison or answer, but it frames a useful validation issue. In-sample error measures how closely a calibrated model reproduces its inputs; held-out quotes can indicate performance on unseen observations, provided the split reflects the intended use. For exotic pricing, however, using all available vanilla quotes may be appropriate for calibration, while out-of-sample fit on vanilla options does not by itself establish accuracy for exotic payoffs. The document leaves the choice of split, test instruments, and evaluation criteria open.
Key ideas
- Calibration fit on vanilla options and validation on held-out quotes answer different questions.
- A data split can help compare a model’s ability to fit unseen vanilla option prices.
- Using all available market data may be sensible when the goal is to calibrate for exotic pricing.
- Good vanilla-option fit does not alone establish accurate exotic-option prices.
- The document poses the validation question but supplies no empirical answer or testing procedure.
Tags
Full text
# Calibration of Levy models and Stochastic Volatility Models - Data used # Calibration of Levy models and Stochastic Volatility Models - Data used I'd like to ask a question regarding something I often come across in research papers. It's about how authors calibrate Levy Models and Stochastic Volatility Models to compare models' performance, and to determine prices for exotic options like Arithmetic Asian Options. Generally, these models are calibrated primarily with data from Vanilla option prices (when joint calibration with exotic option markets is not possible). My two questions are: 1) Shouldn't we split the data into a training set and a testing set when evaluating how well different models fit the volatility surface? In other words, isn't it more appropriate to use one set of data for model calibration, and another to assess how accurately the model reflects the volatility surface? 2) Frequently, authors do not separate the data, instead they use the whole data set for both calibration (training) and performance evaluation (testing). Why? Thanks for considering this. P.S. I do understand that for the only purpose of pricing exotic options, using all available data for calibration makes sense. However, my concern is about comparing the effectiveness of different models.
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.