Skip to content
All library documents

Automatic Differentiation for Calibration and Risk Sensitivities

Article Quant Q&A · Author: Olaf

Summary

The document distinguishes automatic differentiation (AD) from finite differences and symbolic differentiation. AD propagates derivative information through the operations in a coded function, applying the chain rule as the function is evaluated. In principle, this gives derivatives accurate up to floating-point rounding, while finite differences introduce approximation error. Unlike symbolic systems that produce a derivative expression, AD commonly computes derivative values alongside the original function’s result.

A simple implementation uses dual numbers and overloaded arithmetic: each value carries a derivative component, and operations propagate that component. The document notes that the approach can support optimization, model calibration, and risk sensitivities, and can be extended to vector inputs and outputs to compute Jacobians. It does not compare AD modes or quantify performance for Monte Carlo applications; efficiency depends on input and output dimensions, and numerical stability still matters.

Key ideas

  • Automatic differentiation propagates derivatives through a program’s operations using the chain rule.
  • Its derivative is theoretically exact apart from floating-point rounding, unlike finite-difference estimates.
  • Dual numbers provide one implementation by carrying derivative information alongside each value.
  • AD can support optimization, calibration, sensitivities, and Jacobian calculations.
  • The most efficient AD approach depends on the dimensions of the inputs and outputs.

Tags

Full text
# How does Algorithmic Differentiation work and where can it be applied?


# How does Algorithmic Differentiation work and where can it be applied?












The title says it all, but let me expand on it. Algorithmic differentiation seems to be a method that allows a program / compiler to determine what the derivative is of a function. I imagine it's a great tool for things like optimization, since many of these algorithms work better if we know the cost function and its derivative.

I know that algorithmic differentiation is not the same as numerical differentiation, since the derivative is not exact. But is it then the same as symbolic differentiation, as implemented in e.g. SymPy, Matlab or Mathematica?

Second, where can it actually be used in quantitative finance? I mentioned optimization, so calibration comes to mind. Greeks are also a natural application, since these are derivatives by definition, but how does that work in the case of Monte Carlo? And are there any other applications?

## Answer by Tyler Olsen (score 8, accepted)

https://quant.stackexchange.com/a/21898

Automatic Differentiation (aka AD) is a family of methods that are used to evaluate the derivative of a coded function. These methods are far more accurate than finite differences, since they are theoretically exact in the absence of floating point roundoff error.

AD is, however, subtly different than symbolic differentiation. The key difference here is that computer algebra systems (like Mathematica) will return a function that can be evaluated for a derivative. In many implementations of Automatic differentiation, however, the derivative is computed as a side-effect of evaluating the original function.

There are a number of different ways to implement AD, and they are outlined in a reasonable amount of detail on the wikipedia page.

https://en.wikipedia.org/wiki/Automatic_differentiation

Simple Implementation Method

Perhaps the most straightforward method to implement AD is using "Dual Numbers" and operator overloading. The key idea is to introduce a new type of scalar number, called a Dual Number. The dual number augments the algebra of real numbers by replacing every number $x$ with $x + x' \epsilon$, where $\epsilon$ has the property $\epsilon^2 = 0$.

Using this, you can define the algebra of dual numbers as you might expect:

\begin{align} (u+u'\epsilon) + (v + v'\epsilon) &= (u+v) + (u' + v')\epsilon \\ (u+u'\epsilon) - (v + v'\epsilon) &= (u-v) + (u' - v')\epsilon \\ (u+u'\epsilon) * (v+v'\epsilon) &= (uv) + (uv' + vu')\epsilon \\ (u + u'\epsilon) / (v+v'\epsilon) &= (u/v) + \frac{u'v - uv'}{v^2}\epsilon \end{align}

See the wikipedia page for extensions of this to other functions ($\sin$, $\cos$, $\exp$, ...)

The astute reader will notice that the $\epsilon$ component of the result of any operation between dual numbers implements the chain rule for that operation. Due to this, any sequence of operations (i.e. a coded function) will simultaneously compute the function value and its derivative. The only loss of accuracy that one incurs using this method is floating-point roundoff errors (so you still need to use numerically stable algorithms).

On the software side, the primary burden is the implementation of a Dual Number class and to overload any operators/functions that you want to use. In addition, any functions that you wish to differentiate should be (in C++ parlance) templated on the input type. Eg, rather than:

```
double f(double x) { ... }
```

you would write

```
template<typename T>
T f(T x) {...}
```

Doing so will allow you to re-use the same code to do normal function calls or to do the AD function calls.

Example Code

Here is a link to my C++ implementation of the method that I use in my finite element library. The code is free to take/use/modify if you'd like. You'll have to strip out the non-standard header and the namespace open/close lines, but it shouldn't have any dependencies to the rest of the library.

https://github.com/tjolsen/YAFEL/blob/master/include/utils/DualNumber.hpp

Final Thoughts

The ideas described here can be generalized to vector-valued inputs/outputs. By doing so, you can compute Jacobians without having to derive what each partial derivative should be. Depending on the dimension of the inputs and outputs, other methods than the Dual Number method described here may be more efficient, so it's worth reading about in more detail.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.