Passer au contenu
Tous les documents de la bibliothèque

Estimation par maximum de vraisemblance des modèles normaux et exponentiels

Notebook Cours Quantopian

Résumé

Ce tutoriel présente l’estimation par maximum de vraisemblance à l’aide de distributions normales et exponentielles. Pour un échantillon normal, il établit les estimations de la moyenne et de l’écart-type et les compare aux estimations d’une bibliothèque logicielle. Pour un échantillon exponentiel, il explique la convention de paramétrage utilisée par la bibliothèque numérique et estime l’échelle de la distribution à partir de la moyenne de l’échantillon. Des graphiques de densités de probabilité ajustées permettent de les comparer visuellement aux observations simulées.

L’exemple final ajuste une loi normale aux rendements quotidiens d’une seule action, puis utilise le test de Jarque–Bera pour déterminer si l’échantillon de rendements est compatible avec la normalité. Le tutoriel souligne qu’ajuster une distribution ne suffit pas à établir qu’elle est appropriée ; il convient de vérifier la qualité de l’ajustement. Les démonstrations sont introductives et s’appuient sur des échantillons simulés ainsi que sur un exemple limité de rendements historiques. Une loi normale ajustée peut ne pas rendre compte de caractéristiques importantes des rendements financiers, et le document n’établit ni la normalité des rendements de l’exemple ni la stabilité des paramètres ajustés.

Idées clés

  • Le maximum de vraisemblance choisit les paramètres de distribution qui rendent l’échantillon observé le plus plausible selon le modèle.
  • Pour un échantillon normal, les estimations du maximum de vraisemblance utilisent la moyenne de l’échantillon et un écart-type calculé en divisant par le nombre d’observations.
  • Avec la convention de loi exponentielle présentée, l’échelle estimée est la moyenne de l’échantillon.
  • Une densité ajustée peut être comparée visuellement aux données observées, mais une concordance visuelle ne constitue pas une validation formelle.
  • Le test de Jarque–Bera peut évaluer la normalité ; ajuster une loi normale ne démontre pas à lui seul que les rendements sont normaux.

Étiquettes

Texte intégral
# Maximum Likelihood Estimates (MLEs)


<a href="https://www.quantrocket.com"><img alt="QuantRocket logo" src="https://www.quantrocket.com/assets/img/notebook-header-logo.png"></a>

© Copyright Quantopian Inc.<br>
© Modifications Copyright QuantRocket LLC<br>
Licensed under the [Creative Commons Attribution 4.0](https://creativecommons.org/licenses/by/4.0/legalcode).<br>
<a href="https://www.quantrocket.com/disclaimer/">Disclaimer</a>

***
[Quant Finance Lectures (adapted Quantopian Lectures)](Introduction.ipynb) › Lecture 13 - Maximum Likelihood Estimation
***

# Maximum Likelihood Estimates (MLEs)

By Delaney Granizo-Mackenzie and Andrei Kirilenko developed as part of the Masters of Finance curriculum at MIT Sloan.


In this tutorial notebook, we'll do the following things:
1. Compute the MLE for a normal distribution.
2. Compute the MLE for an exponential distribution.
3. Fit a normal distribution to asset returns using MLE.

First we need to import some libraries

```python
import math
import matplotlib.pyplot as plt
import numpy as np
import scipy
import scipy.stats
```

## Normal Distribution
We'll start by sampling some data from a normal distribution.

```python
TRUE_MEAN = 40
TRUE_STD = 10
X = np.random.normal(TRUE_MEAN, TRUE_STD, 1000)
```

Now we'll define functions that, given our data, will compute the MLE for the $\mu$ and $\sigma$ parameters of the normal distribution.

Recall that

$$\hat\mu = \frac{1}{T}\sum_{t=1}^{T} x_t$$

$$\hat\sigma = \sqrt{\frac{1}{T}\sum_{t=1}^{T}{(x_t - \hat\mu)^2}}$$

```python
def normal_mu_MLE(X):
    # Get the number of observations
    T = len(X)
    # Sum the observations
    s = sum(X)
    return 1.0/T * s

def normal_sigma_MLE(X):
    T = len(X)
    # Get the mu MLE
    mu = normal_mu_MLE(X)
    # Sum the square of the differences
    s = sum( np.power((X - mu), 2) )
    # Compute sigma^2
    sigma_squared = 1.0/T * s
    return math.sqrt(sigma_squared)
```

Now let's try our functions out on our sample data and see how they compare to the built-in `np.mean` and `np.std`

```python
print("Mean Estimation")
print(normal_mu_MLE(X))
print(np.mean(X))
print("Standard Deviation Estimation")
print(normal_sigma_MLE(X))
print(np.std(X))
```

Now let's estimate both parameters at once with scipy's built in `fit()` function.

```python
mu, std = scipy.stats.norm.fit(X)
print("mu estimate:",  str(mu))
print("std estimate:", str(std))
```

Now let's plot the distribution PDF along with the data to see how well it fits. We can do that by accessing the pdf provided in `scipy.stats.norm.pdf`.

```python
pdf = scipy.stats.norm.pdf
# We would like to plot our data along an x-axis ranging from 0-80 with 80 intervals
# (increments of 1)
x = np.linspace(0, 80, 80)
plt.hist(X, bins=x, density='true')
plt.plot(pdf(x, loc=mu, scale=std))
plt.xlabel('Value')
plt.ylabel('Observed Frequency')
plt.legend(['Fitted Distribution PDF', 'Observed Data', ]);
```

## Exponential Distribution
Let's do the same thing, but for the exponential distribution. We'll start by sampling some data.

```python
TRUE_LAMBDA = 5
X = np.random.exponential(TRUE_LAMBDA, 1000)
```

`numpy` defines the exponential distribution as
$$\frac{1}{\lambda}e^{-\frac{x}{\lambda}}$$

So we need to invert the MLE from the lecture notes. There it is

$$\hat\lambda = \frac{T}{\sum_{t=1}^{T} x_t}$$

Here it's just the reciprocal, so

$$\hat\lambda = \frac{\sum_{t=1}^{T} x_t}{T}$$

```python
def exp_lamda_MLE(X):
    T = len(X)
    s = sum(X)
    return s/T
```

```python
print("lambda estimate:", str(exp_lamda_MLE(X)))
```

```python
# The scipy version of the exponential distribution has a location parameter
# that can skew the distribution. We ignore this by fixing the location
# parameter to 0 with floc=0
_, l = scipy.stats.expon.fit(X, floc=0)
```

```python
pdf = scipy.stats.expon.pdf
x = range(0, 80)
plt.hist(X, bins=x, density='true')
plt.plot(pdf(x, scale=l))
plt.xlabel('Value')
plt.ylabel('Observed Frequency')
plt.legend(['Fitted Distribution PDF', 'Observed Data', ]);
```

## MLE for Asset Returns

Now we'll fetch some real returns and try to fit a normal distribution to them using MLE.

```python
from quantrocket.master import get_securities
from quantrocket import get_prices

aapl_sid = get_securities(symbols="AAPL", vendors='usstock').index[0]

prices = get_prices('usstock-free-1min', data_frequency='daily', sids=aapl_sid, fields='Close', start_date='2014-01-01', end_date='2015-01-01')
prices = prices.loc['Close'][aapl_sid]

# This will give us the number of dollars returned each day
absolute_returns = np.diff(prices)
# This will give us the percentage return over the last day's value
# the [:-1] notation gives us all but the last item in the array
# We do this because there are no returns on the final price in the array.
returns = absolute_returns/prices[:-1]
```

Let's use `scipy`'s fit function to get the $\mu$ and $\sigma$ MLEs.

```python
mu, std = scipy.stats.norm.fit(returns)
pdf = scipy.stats.norm.pdf
x = np.linspace(-1,1, num=100)
h = plt.hist(returns, bins=x, density='true')
l = plt.plot(x, pdf(x, loc=mu, scale=std))
```

Of course, this fit is meaningless unless we've tested that they obey a normal distribution first. We can test this using the Jarque-Bera normality test. The Jarque-Bera test will reject the hypothesis of a normal distribution if the p-value is under a c.

```python
from statsmodels.stats.stattools import jarque_bera
jarque_bera(returns)
```

```python
jarque_bera(np.random.normal(0, 1, 100))
```

---

**Next Lecture:** [Regression Model Instability](Lecture14-Regression-Model-Instability.ipynb) 

[Back to Introduction](Introduction.ipynb) 

---

*This presentation is for informational purposes only and does not constitute an offer to sell, a solicitation to buy, or a recommendation for any security; nor does it constitute an offer to provide investment advisory or other services by QuantRocket LLC ("QuantRocket"). Nothing contained herein constitutes investment advice or offers any opinion with respect to the suitability of any security, and any views expressed herein should not be taken as advice to buy, sell, or hold any security or as an endorsement of any security or company.  In preparing the information contained herein, the authors have not taken into account the investment needs, objectives, and financial circumstances of any particular investor. Any views expressed and data illustrated herein were prepared based upon information believed to be reliable at the time of publication. QuantRocket makes no guarantees as to their accuracy or completeness. All information is subject to change and may quickly become unreliable for various reasons, including changes in market conditions or economic circumstances.*
![notebook output](figures/p1_1.png)
![notebook output](figures/p1_2.png)
![notebook output](figures/p1_3.png)

Reproduit dans son intégralité avec attribution, conformément à la licence de la source. Licence: CC BY 4.0

Ce résumé a été rédigé par l’agent de recherche de Stratmill à partir de la source originale ; il n’en est pas une copie.