Skip to content
All library documents

Least Squares Monte Carlo for Pricing American Options

Article Quant Q&A · Author: Nasser Bin

Summary

Least Squares Monte Carlo (LSM) estimates the continuation value of an American option by simulating risk-neutral price paths and fitting regressions as it works backward through possible exercise dates. At each date, the estimated value of waiting is compared with the payoff from exercising, and the path’s future cash flow is updated accordingly. This provides a way to handle early exercise when ordinary forward Monte Carlo does not directly provide conditional values at intermediate times.

The explanation contrasts LSM with lattice and finite-difference methods, which use backward induction but become costly as the number of underlying assets grows. Nested simulations can estimate continuation values but may be impractical. LSM instead regresses discounted future values on chosen basis functions and state variables. The document notes that finite samples and limited basis functions can introduce bias, and that results depend on representing the information available at each time adequately. It describes convergence in general terms but supplies no numerical example or implementation guidance.

Key ideas

  • American option valuation requires comparing exercise value with the expected discounted value of continuing.
  • LSM simulates paths forward and estimates continuation values with regressions while stepping backward through exercise dates.
  • The regression uses selected basis functions and state variables to represent conditional future value.
  • Lattices and finite-difference methods become more costly as the number of underlying assets increases.
  • Finite simulations and an inadequate basis or state representation can bias estimates.

Tags

Full text
# Least Squares Monte Carlo


# Least Squares Monte Carlo












Could you explain to me in words (no formulas) the concept of the Least Squares Monte Carlo method to price an American style option?

## Answer by Quantuple (score 14, accepted)

https://quant.stackexchange.com/a/46083

To compute the price of an American option or a callable instrument in general, at each potential exercise date, one is required to compare its continuation value (discounted risk-neutral expectation of what the option would pay off if it was not exercised) to the relevant exercise value/early redemption price.

By construction, lattice and finite difference methods allow a straightforward computation of the former continuation values, since they work by backward induction (starting from the terminal condition at expiry and working backwards up to inception computing risk-neutral expectations). However, these methods suffer from the curse of dimensionality (the computational burden increases rapidly with the number of underlying assets).

At the other end of the spectrum, the standard Monte Carlo method is forward in nature: one simulates realisations of the underlying price process under the risk-neutral measure, applies the payoff function and takes the discounted expectation of those paths' payouts to obtain the option price. By construction, computing continuation values at future times is then less straightforward. One could do it with nested simulations but this wouldn't be practical.

An alternative, first proposed by Longstaff and Schwartz in a celebrated paper, consists in simulating the paths and then working backwards in time to estimate the continuation values through least squares regression over a set of so-called basis functions. At each time step, the paths' payouts are updated by comparing the estimated continuation values to the exercise values and then repeating the operation. The method has some known biases when using a finite number of simulations and basis functions but is shown to converge otherwise.

More info on the regression step as asked in the comments.

In a regression problem, you face a noisy data set and ask yourself the question of what process could have generated this data. In its simplest form you could see this data-generating process as a black box, taking some inputs $x$ and generating outputs $y = f(x) + \epsilon$ where $\epsilon$ is a zero-mean noise term. Your goal is then to estimate $$ f(x) = \Bbb{E}\left[ y \vert x \right] $$

There are obviously many ways to approach the problem. One of the most intuitive ( discriminative modelling) is to suppose a parametric form $f(x) = \sum_{i=1}^N \alpha_i \phi_i(x)$ where you decide of the $\phi_i(x)$ yourself (basis functions) and the problem then boils down to estimating the $\alpha_i$ that best fit the data. This is called discriminative modelling (versus generative modelling).

Now you might ask yourself the question, what the real $f(x)$ of the DGP was $f(x) = sin(x)$ and I only selected one basis function $\phi_1(x)=x$. Surely in that case I'll never have a good estimate of $f(x)$: this is why the choice of basis functions is important: they should ideally allow you to represent a wide set of functions (notion of complete set).

How does that relate to the problem at hand? The price of a derivative at $t$ (which you are looking for) is the expectation of its discounted value at $t+\delta$ conditional on all the information you have at $t$, mathematically: $$ V_t = \Bbb{E} \left[ P(t,t+\delta) V_{t+\delta} \vert \mathcal{F}_t \right] $$

This is of the same form as the previous problem: you observe $P(t,t+\delta) V_{t+\delta}$ conditional on $\mathcal{F}_t$ (simulated using MC) and you are looking for the data-generating function $V_t$. More specifically, if you assume that "all the information available" at $t$ boils down to the knowledge of the current spot price $S_t$ you get the following problem $$ V_t = f(S_t) = \Bbb{E} \left[ P(t,t+\delta) V_{t+\delta} \vert S_t \right] $$ similar to $$ f(x) = \Bbb{E} \left[ y \vert x \right] $$ Now again the fact of choosing to represent all the information available at $t$ ($\mathcal{F}_t$) by just the knowledge of $S_t$ (so basically the choice of regressor on top of the choice of basis functions), might be a too stringent assumption in some cases.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.