Fitting and Testing Geometric Brownian Motion Against Returns
Summary
The document discusses how to interpret and empirically assess Geometric Brownian Motion (GBM). Under GBM, log returns follow a normal distribution whose mean combines the drift and a volatility adjustment. The drift is therefore a model parameter: its meaning depends on the assumed process, rather than being a model-free statistic that can simply be read from a return sample. Estimation fits parameters conditional on the chosen model, while diagnostic tests can assess whether its assumptions are plausible.
The questioner reports a discrepancy between an input drift and a simulated estimate from terminal prices. The response does not diagnose the simulation in detail; it emphasizes that parameter estimates belong to the model used and that similarly named parameters in different distributions or processes can represent different quantities. The document thus offers a conceptual distinction between estimation and validation, but not a complete testing procedure or explanation of the reported numerical mismatch. Its discussion is limited to the stated GBM setup and does not establish that real asset returns satisfy its assumptions.
Key ideas
- Under GBM, the mean of log returns relates to drift through a volatility adjustment.
- Drift is defined relative to an assumed model and is not a model-independent quantity.
- Fitting a model estimates its parameters, while diagnostic tests evaluate whether its assumptions are plausible.
- Parameters with the same symbol can have different meanings in different models.
- The reported simulation discrepancy is not fully investigated, so the document does not identify its cause.
Tags
Full text
# Empirically validating GBM assumptions
# Empirically validating GBM assumptions
I'm trying to get a better grasp on the suitability of Geometric Brownian Motion as a model for asset prices, and I've stumbled across a hurdle which I can't seem to get over. Assuming GBM, the asset price is given by
$ S_t = S_0 \exp\{ (\mu - \frac{1}{2}\sigma^2)t + \sigma W_t \} $,
which implies that daily log-returns are normally distributed:
$ \log(\frac{S_{t+1}}{S_t}) \sim \mathcal{N}(\mu - \frac{1}{2}\sigma^2, \sigma^2) $.
Now my problem is: how would you go about verifying this with empirical data? If I compute the sample mean of log-returns with real data, then I have to already assume the model in order to extract the drift $ \mu $ by subtracting half the variance. Is there a way to get the drift term directly from the data, without assuming the model? Or is the drift simply a theoretical quantity postulated in the SDE
$ dS_t = S_t (\mu dt + \sigma W_t) $
which defines GBM? I.e. does the drift correspond to any computable statistical quantity?
EDIT: here is my attempt to gain clarity by simulation:
```
T = 1
N = 1000
S0 = 1
mu = -0.00028117*252
sigma = 0.008723*252
seed = 1
f, g = f_g_black_scholes(lamda = mu, mu = sigma)
numpaths = 10000
paths = pd.DataFrame()
for k in range(1, numpaths):
t, S, W = euler_maruyama(seed=seed+k, X0=X0, T=T, N=N, f=f, g=g)
S = pd.DataFrame(S)
paths = pd.concat([paths, S], axis = 1)
final_price = paths.iloc[-1, :]
sigma_sim = np.std(np.log(final_price))
mu_sim = np.mean(np.log(final_price)) + 0.5*sigma_sim**2
```
The drift and volatility I estimated from S&P 500 data, and I made use of this implementation of the Euler-Maruyama solver to numerically generate sample paths. The results I obtain are:
```
mu = -0.070855
mu_sim = -0.010407
sigma = 2.198196
sigma_sim = 2.221657
```
As you can see there is still a significant discrepancy between the input drift and the result of the simulation. How should I make sense of this?
## Answer by KT8 (score 1)
https://quant.stackexchange.com/a/70351
All you can do is fit the values of the parameters of a model you have already assumed.
> Now my problem is: how would you go about verifying this with empirical data?
As Kermittfrog mentioned, after fitting a particular model, you can use test to check whether your initial assumptions make sense or not.
> If I compute the sample mean of log-returns with real data, then I have to already assume the model in order to extract the drift $\mu$ by subtracting half the variance. Is there a way to get the drift term directly from the data, without assuming the model? Or is the drift simply a theoretical quantity postulated in the SDE
Parameters are quantities associated with a model in this case. For example consider normal and lognormal distributions, in both you have a $\mu$ and a $\sigma$, but depending on the model you assume you'll get different quantities from your fit, because this parameters mean different things (no matter on the label or greek letter you use to describe them).Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.