Skip to content
All library documents

Why GBM Drift Is Hard to Estimate from Short Samples

Article Quant Q&A · Author: Futurist

Summary

The discussion concerns joint maximum likelihood estimation of the drift and volatility of a geometric Brownian motion from simulated log prices. Although the likelihood can identify volatility clearly, it may appear nearly flat in the drift direction when observations cover only a short total time horizon. The replies explain that this is an expected statistical limitation, rather than necessarily an error in the likelihood or data generation.

Adding observations while holding the horizon fixed does not materially improve drift precision, because the extra observations subdivide the same overall period. A longer horizon provides more information about drift. Rescaling parameters through preconditioning may help numerical optimization when the objective has directions with very different curvature, but it does not resolve the underlying lack of information. A Bayesian maximum a posteriori approach with a prior for drift is mentioned as one possible alternative; the note does not provide a full derivation or empirical comparison.

Key ideas

  • Joint GBM likelihood estimation can be much less informative about drift than volatility over a short horizon.
  • Increasing sampling frequency over a fixed horizon does not substantially improve drift accuracy.
  • A longer observation horizon improves the information available for estimating drift.
  • Preconditioning can help optimization when parameter directions have very different scales.
  • A prior on drift can be incorporated through maximum a posteriori estimation.

Tags

Full text
# How to get around flat likelihood function when calibrating GBM parameters?


# How to get around flat likelihood function when calibrating GBM parameters?












I want to calibrate jointly the drift mu and volatility sigma of a geometric brownian motion,

$$\log(S_t) = \log(S_{t-1}) + (\mu - 0.5*\sigma^2) \Delta t + \sigma*\sqrt{\Delta t}*Z_t$$

where $Z_t$ is a standard normally distributed random variable, and am testing this by generating data $x = \log(S_t)$ via

```
x(1) = 0;
for i = 2:N
  x(i) = x(i-1) + (mu-0.5*sigma^2)*Deltat + sigma*sqrt(Deltat)*randn;
end
```

and my (log-)likelihood function

```
function LL = LL(x, pars)
  mu    = pars(1);
  sigma = pars(2);
  Nt = size(x,2);
  LL = 0;
  for j = 2:Nt
    LH_j = normpdf(x(j), x(j-1)+(mu-0.5*sigma^2)*Deltat, sigma*sqrt(Deltat));
    LL = LL + log(LH_j);
  end
```

which I maximize using `fmincon` (because sigma is constrained to be positive), with starting values 0.15 and 0.3, true values 0.1 and 0.2, and `N = Nt = 1000` or 100000 generated points over one year (i.e. $\Delta t$ = 0.0001 or 0.000001).

Calibrating the volatility alone yields a nice likelihood function with a maximum around the true parameter, but for small Deltat (less than say 0.1) calibrating both mu and sigma persistently shows a (log-)likelihood surface being very flat in mu (at least around the true parameter); I would expect also a maximum there; for a reason I think it should be possible to calibrate a GBM model to a data series of 100 stock prices in a year, making the average of Deltat = 0.01.

Any sharing of experience or help is greatly appreciated (thoughts passing through my mind: the likelihood function is not right / this is a normal behaviour / too few data points / data generation is not correct / ...?).

## Answer by H. Arponen (score 2)

https://quant.stackexchange.com/a/16432

This is what often happens in optimization problems, i.e. some direction is almost flat. Google 'preconditioning'. Basically the idea is to rescale the variables, so that the Hessian has approx. same order of magnitude values on the diagonals.

Also, that's not a stationary process, so estimation of mu can be difficult.

BTW not sure if it's a very good idea to use the 'normpdf' function here. Instead you should probably just write out the complete loglikelihood expression (not sure if that will help though). It seems about right to me though.

## Answer by bcf (score 2)

https://quant.stackexchange.com/a/20999

The log likelihood function is indeed rather flat in the $\mu$-direction, for small time horizons (you used $T = 1$ it looks like). As you may have noticed, increasing the number of observations but keeping the time horizon the same DOES NOT IMPROVE the accuracy of the estimate of $\mu$ - this is a bit counterintuitive, if you ask me. But, increasing the time horizon $T$ DOES improve the accuracy of the estimate. To see this, see my question here.

Merton (1980) proposed using a Bayesian technique, in which you first specify a prior distribution for $\mu$ and use Bayes' rule to derive a new log likelihood function. This is nowadays known as "maximum a-posteriori (MAP)" estimation and is one way to try to get better estimate of $\mu$. But the fact remains: MLE estimation of $\mu$ is notoriously inaccurate for short time horizons.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.