Deep Learning Risks in Volatility Surface Calibration and Pricing
Summary
The document discusses potential drawbacks of using deep learning to calibrate and price a full implied-volatility surface. It notes that neural methods can produce very fast calibration at application time, while training may involve substantial computation. The main answer focuses on overfitting: as model complexity grows, a network may fit historical data by learning patterns that do not generalize to new market conditions.
The warning is that strong in-sample fit or rapid pricing does not establish reliable forecasts or robust live performance. This concern applies to financial models broadly, including conventional approaches, and cannot be resolved simply by increasing network depth or tuning architecture. A second response mentions training time and suggests that changing market conditions could be addressed with later data, but gives no supporting analysis; using future observations to evaluate a model would also require careful time-aware validation. The document provides qualitative cautions, not comparative experiments or a specific network design for calibration.
Key ideas
- Fast inference does not guarantee that a deep learning model will generalize beyond its training sample.
- Increasing model complexity can fit historical noise and create overconfidence in live use.
- Training time and market regime changes are practical concerns raised in the discussion.
- The document offers broad cautions rather than empirical comparisons or a validated architecture.
Tags
Full text
# Theoretical and practical drawbacks of using Deep Learning for calibration and pricing
# Theoretical and practical drawbacks of using Deep Learning for calibration and pricing
I am investigating the suitability of using deep learning for pricing and calibration for the full implied volatility surface. Such examples of their application are in papers here and here. During examples, the latter achieved the calibration task in mere milliseconds, which were orders of magnitude faster than numerical approximations and monte carlo methods.
However, I was wondering what some of the drawbacks of using Deep Learning for this task could be, and whether there were any papers or research that highlight this issue. The latter paper cited 'cumbersome computations' during training, but I was wondering if there were any drawbacks regarding, perhaps, the architecture of DL networks as well?
## Answer by demully (score 7)
https://quant.stackexchange.com/a/69148
The essence of the problem is the "bias-variance" problem in machine learning. Which you can wiki (or find dozens of papers on; it's a famous issue).
You can, with ever greater complexity, create a model that ever-better explains history. But the "shortcuts" it uses to do this can sacrifice its ability to forecast new and unseen data. The model should be more uncertain and more prepared to allow historical mistakes to prevent overconfidence and reduce the chances of future ones.
It's the same problem, but just with newer tech, than the infamous "the backtests stop working once the structured product goes live" problem that has dogged finance for decades. The models are not "optimised"; they are "over-optimised" (for the past, and thus under-optimised for the future).
If shallow learning cannot solve the problem, it does not follow that the problem can be solved. Applying deeper and deeper and deeper learning until you can solve the bit of the problem you have in your sample is just ultra-sophisticated cheating. And the boss who pays your bonus doesn't understand the model anyway, so so long as the product is selling well, you both still get paid.
This, albeit put very bluntly, is a much bigger problem than arguing about the hyperparameters of this model or that. Both will torture the data to force it to "confess" that noise is signal if you want them to...
hope this helps, DEM
## Answer by Johnny (score 0)
https://quant.stackexchange.com/a/61904
Deep learning has drawbacks in that it can take a long time to train. In application, if the market conditions change, you can always use future market data that hasn't happened yet to feed in to your model to get an edge. Proven by the filtration $F_{t}$.Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.