Skip to content
All library documents

Fast Option Pricing Versus Deep Hedging Calibration

Article Quant Q&A · Author: dasfobia

Summary

The document separates machine learning used to approximate an already computed option pricing function from methods that learn both a claim’s value and its optimal hedge. For fast pricing, a model is trained on examples generated by slower numerical pricing or simulation, then used to interpolate prices for nearby market and model inputs. The response notes that gradient boosting can serve this approximation task, alongside neural networks.

Deep pricing, as described here, solves a stochastic control problem associated with the claim’s value and hedging strategy. The answer characterizes this as requiring coordinated learning of a control and a value function, often with two neural networks. It suggests that synchronizing neural network training may explain their use in this setting, while offering no comparative benchmark or general proof that boosting performs poorly. The discussion also relates the challenge to reinforcement learning and mean field games, where interacting agents or market impact add further coupled dynamics.

Key ideas

  • Fast pricing learns to approximate outputs from slower numerical pricing or simulation.
  • Gradient boosting can be used for fast pricing on tabular inputs.
  • Deep pricing aims to learn both a contingent claim’s value and its optimal hedge.
  • The described control approach coordinates learning of a policy and a value function.
  • The proposed explanation for neural networks’ use is their convenient coordination, not a demonstrated universal advantage.

Tags

Full text
# Deep vs "shallow" calibration of option pricing models


# Deep vs "shallow" calibration of option pricing models












I am currently investigating the application of deep learning in calibrating option pricing models, specifically, models of rough volatility, such as rBergomi. While there is a lot of research on using deep neural nets as fast option pricing engines with market and model parameters as input features, there seems to be practically nothing on using gradient boosting to approximate the pricing function in some option pricing model. Given that gradient boosting usually outperforms deep learning on tabular data, why is the research concentrated on deep calibration?

## Answer by lehalle (score 2)

https://quant.stackexchange.com/a/78144

There is a difference between

- "fast pricing" that is mimicking the function that maps characteristics of the observables (what you call market data and parameters in your question) to the value of a derivative contract that has been already computed (but a way that is slow)

- "deep pricing" that is computing the value and the optimal hedging strategy of a contingent claim.

The fact that the second one is called "deep" is compatible with your question: deep neural networks are used that for, whereas for the first one (learning a mapping) a lot of other options are used.

For fast pricing boosting methods are used too. See for instance

- Davis, Jesse, Laurens Devos, Sofie Reyners, and Wim Schoutens. "Gradient boosting for quantitative finance" Journal of Computational Finance 24, no. 4 (2020)

- Ferraz, João Diogo Marques, Master Thesis "Pricing options using the XGBoost Model" , Instituto Superior de Economia e Gestão, 2022.

Fast pricing is about using a database of slow numerical simulation over a set $\Omega$ of examples of mappings $X\in \Omega\mapsto Y=\mathbb{E}(\mbox{payoff})$.

The advantage is that if you have made simulations for two examples $X_1$ and $X_2$ and you face in real-time a new configuration $X_{new}$ that is not "too different"from them, the ML also will interpollate and give quickly a result.

Deep pricing is different: you want to solve the HJB characterising the claim; it is made of two component: the value of the claim (that is what you will interpolate later with fast pricing) and the optimal control that is the hedging portfolio that replicates the exposure of the claim.

This is commonly called the Howard (or policy-value iteration) method. When you use machine learning you have to use two algorithms:

- one to emulate the optimal control, solving the optimum part of the HJB

- and once the control is set, the HJB boils down to a PDE for a value function, that you need to emulate using another neural net.

Globally, when you use this process you have to get 2 neural nets to converge simultaneously. And in a sense it is like a larger "global neural net", it works decently well. People who tried other algos (like boosting) did not obtained (at my knowledge) good result because you have to control simultaneously the convergence of the two algos. For neural nets it is not that difficult (you can synchronize their learning rates). It somehow looks like Q-learning of Reinforcement Learning.

A way to understand this effect is to look at what happens when you introduce another convergence: when you want to solve a Mean Field Game of stochastic controlers. Imagine you want your hedging process to take into account the fact that other banks are hedging similar products and that all the hedging trades create a coupling via market impact. Recently Angiuli, Andrea, Jean-Pierre Fouque, and Mathieu Laurière. "Unified reinforcement Q-learning for mean field game and control problems" Mathematics of Control, Signals, and Systems 34, no. 2 (2022): 217-271. If you are interested in this kind of question I recommend Capponi, Agostino, and C.-A. L, eds. Machine Learning and Data Sciences for Financial Markets: A Guide to Contemporary Practices. Cambridge University Press, 2023. There is a full part: New Frontiers for Stochastic Control in Finance with 6 chapters on it.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.