Notions de base des notebooks Jupyter pour l’analyse quantitative des marchés
Résumé
Ce tutoriel d’introduction montre comment utiliser les notebooks Jupyter pour l’analyse quantitative. Il explique la différence entre les cellules de code et de texte, l’exécution des cellules et l’affichage des résultats, l’importation de bibliothèques courantes d’analyse et de visualisation, ainsi que l’utilisation de la complétion automatique et de la documentation intégrée pour explorer les fonctions disponibles. Les exemples montrent comment générer des données aléatoires, tracer des séries, légender des graphiques et calculer des statistiques élémentaires.
Le tutoriel applique ensuite ces outils aux cours des actions : rechercher un identifiant de titre, récupérer les cours de clôture quotidiens, calculer les rendements en pourcentage, tracer un histogramme des rendements et comparer la distribution observée à un échantillon normal fondé sur sa moyenne et son écart-type estimés. Il montre aussi comment calculer et tracer une moyenne glissante, qui n’est disponible qu’après l’accumulation d’un nombre suffisant d’observations. Il s’agit d’exemples de base sur les méthodes de travail, pas d’une stratégie de trading ni d’un test rigoureux du comportement des rendements ; le tutoriel indique explicitement que la forme de l’échantillon de rendements peut différer d’une distribution normale.
Idées clés
- Les cellules d’un notebook peuvent contenir du code exécutable ou du texte explicatif ; l’exécution affiche les résultats sous la cellule.
- L’importation de bibliothèques, la complétion automatique et la documentation intégrée facilitent l’analyse exploratoire.
- Les séries de prix peuvent être résumées et visualisées, et les variations en pourcentage permettent de calculer les rendements.
- Un histogramme des rendements peut être comparé à une distribution normale construite à partir des statistiques estimées de l’échantillon.
- Les moyennes glissantes nécessitent une fenêtre initiale complète avant que leurs valeurs soient disponibles.
Étiquettes
Texte intégral
# Introduction to Notebooks
<a href="https://www.quantrocket.com"><img alt="QuantRocket logo" src="https://www.quantrocket.com/assets/img/notebook-header-logo.png"></a>
© Copyright Quantopian Inc.<br>
© Modifications Copyright QuantRocket LLC<br>
Licensed under the [Creative Commons Attribution 4.0](https://creativecommons.org/licenses/by/4.0/legalcode).<br>
<a href="https://www.quantrocket.com/disclaimer/">Disclaimer</a>
***
[Quant Finance Lectures (adapted Quantopian Lectures)](Introduction.ipynb) › Lecture 1 - Introduction to Notebooks
***
# Introduction to Notebooks
<a href="https://youtu.be/W-TlWzwM208?t=28" target="_blank">Quantopian video for this lecture ↗</a>
Jupyter notebooks allow one to perform a great deal of data analysis and statistical validation. We'll demonstrate a few simple techniques here.
## Code Cells vs. Text Cells
As you can see, each cell can be either code or text. To select between them, choose from the 'Markdown' dropdown menu on the top of the notebook.
## Executing a Command
A code cell will be evaluated when you press play, or when you press the shortcut, shift-enter. Evaluating a cell evaluates each line of code in sequence, and prints the results of the last line below the cell.
```python
2 + 2
```
Sometimes there is no result to be printed, as is the case with assignment.
```python
X = 2
```
Remember that only the result from the last line is printed.
```python
2 + 2
3 + 3
```
However, you can print whichever lines you want using the `print` statement.
```python
print(2 + 2)
3 + 3
```
## Knowing When a Cell is Running
While a cell is running, a `[*]` will display on the left. When a cell has yet to be executed, `[ ]` will display. When it has been run, a number will display indicating the order in which it was run during the execution of the notebook `[5]`. Try on this cell and note it happening.
```python
#Take some time to run something
c = 0
for i in range(10000000):
c = c + i
c
```
## Importing Libraries
The vast majority of the time, you'll want to use functions from pre-built libraries. Here I import numpy and pandas, the two most common and useful libraries in quant finance. I recommend copying this import statement to every new notebook.
Notice that you can rename libraries to whatever you want after importing. The `as` statement allows this. Here we use `np` and `pd` as aliases for `numpy` and `pandas`. This is a very common aliasing and will be found in most code snippets around the web. The point behind this is to allow you to type fewer characters when you are frequently accessing these libraries.
```python
import numpy as np
import pandas as pd
# This is a plotting library for pretty pictures.
import matplotlib.pyplot as plt
```
## Tab Autocomplete
Pressing tab will give you a list of Python's best guesses for what you might want to type next. This is incredibly valuable and will save you a lot of time. If there is only one possible option for what you could type next, Python will fill that in for you. Try pressing tab very frequently, it will seldom fill in anything you don't want, as if there is ambiguity a list will be shown. This is a great way to see what functions are available in a library.
Try placing your cursor after the `.` and pressing tab.
```python
np.random.
```
## Getting Documentation Help
Placing a question mark after a function and executing that line of code will give you the documentation Python has for that function. It's often best to do this in a new cell, as you avoid re-executing other code and running into bugs.
```python
np.random.normal?
```
## Sampling
We'll sample some random data using a function from `numpy`.
```python
# Sample 100 points with a mean of 0 and an std of 1. This is a standard normal distribution.
X = np.random.normal(0, 1, 100)
```
## Plotting
We can use the plotting library we imported as follows.
```python
plt.plot(X)
```
### Squelching Line Output
You might have noticed the annoying line of the form `[<matplotlib.lines.Line2D at 0x7f72fdbc1710>]` before the plots. This is because the `.plot` function actually produces output. Sometimes we wish not to display output, we can accomplish this with the semi-colon as follows.
```python
plt.plot(X);
```
### Adding Axis Labels
No self-respecting quant leaves a graph without labeled axes. Here are some commands to help with that.
```python
X = np.random.normal(0, 1, 100)
X2 = np.random.normal(0, 1, 100)
plt.plot(X);
plt.plot(X2);
plt.xlabel('Time') # The data we generated is unitless, but don't forget units in general.
plt.ylabel('Returns')
plt.legend(['X', 'X2']);
```
## Generating Statistics
Let's use `numpy` to take some simple statistics.
```python
np.mean(X)
```
```python
np.std(X)
```
## Getting Real Pricing Data
Randomly sampled data can be great for testing ideas, but let's get some real data. In QuantRocket, all securities are referenced by sid (short for "security ID") rather than by symbol since symbols can change. So, first, we'll use the `get_securities` function to look up the sid for MSFT.
(Notice the use of `vendors='usstock'` in the `get_securities` function call. This limits the query to securities from the US Stock dataset. This filter isn't necessary if you've only collected US Stock data, but is a best practice when looking up securities by symbol in case you've also collected data from other global exchanges where the same ticker symbols are re-used.)
```python
from quantrocket.master import get_securities
securities = get_securities(symbols='MSFT', fields=['Sid','Symbol','Exchange'], vendors='usstock')
securities
```
This returns a pandas dataframe, where sids are stored in the dataframe's index.
Then we use `get_prices` to query our data bundle. Although the bundle contains minute data, here we use the `data_frequency` parameter to request the data at daily frequency:
```python
MSFT = securities.index[0]
from quantrocket import get_prices
data = get_prices("usstock-free-1min", data_frequency='daily', sids=MSFT, start_date='2012-01-01', end_date='2015-06-01', fields="Close")
```
Our data is now a dataframe. You can see the datetime index and the colums with different pricing data.
```python
data.head()
```
This is a pandas dataframe, so we can index in to just get the closing price for MSFT like this. For more info on pandas, please [click here](http://pandas.pydata.org/pandas-docs/stable/10min.html).
```python
X = data.loc['Close'][MSFT]
```
Because there is now also date information in our data, we provide two series to `.plot`. `X.index` gives us the datetime index, and `X.values` gives us the pricing values. These are used as the X and Y coordinates to make a graph.
```python
plt.plot(X.index, X.values)
plt.ylabel('Price')
plt.legend(['MSFT']);
```
We can get statistics again on real data.
```python
np.mean(X)
```
```python
np.std(X)
```
## Getting Returns from Prices
We can use the `pct_change` function to get returns. Notice how we drop the first element after doing this, as it will be `NaN` (nothing -> something results in a NaN percent change).
```python
R = X.pct_change()[1:]
```
We can plot the returns distribution as a histogram.
```python
plt.hist(R, bins=20)
plt.xlabel('Return')
plt.ylabel('Frequency')
plt.legend(['MSFT Returns']);
```
Get statistics again.
```python
np.mean(R)
```
```python
np.std(R)
```
Now let's go backwards and generate data out of a normal distribution using the statistics we estimated from Microsoft's returns. We'll see that we have good reason to suspect Microsoft's returns may not be normal, as the resulting normal distribution looks far different.
```python
plt.hist(np.random.normal(np.mean(R), np.std(R), 10000), bins=20)
plt.xlabel('Return')
plt.ylabel('Frequency')
plt.legend(['Normally Distributed Returns']);
```
## Generating a Moving Average
`pandas` has some nice tools to allow us to generate rolling statistics. Here's an example. Notice how there's no moving average for the first 60 days, as we don't have 60 days of data on which to generate the statistic.
```python
# Take the average of the last 60 days at each timepoint.
MAVG = X.rolling(window=60).mean()
plt.plot(X.index, X.values)
plt.plot(MAVG.index, MAVG.values)
plt.ylabel('Price')
plt.legend(['MSFT', '60-day MAVG']);
```
---
**Next Lecture**: [Introduction to Python](Lecture02-Introduction-to-Python.ipynb)
[Back to Introduction](Introduction.ipynb)
---
*This presentation is for informational purposes only and does not constitute an offer to sell, a solicitation to buy, or a recommendation for any security; nor does it constitute an offer to provide investment advisory or other services by QuantRocket LLC ("QuantRocket"). Nothing contained herein constitutes investment advice or offers any opinion with respect to the suitability of any security, and any views expressed herein should not be taken as advice to buy, sell, or hold any security or as an endorsement of any security or company. In preparing the information contained herein, the authors have not taken into account the investment needs, objectives, and financial circumstances of any particular investor. Any views expressed and data illustrated herein were prepared based upon information believed to be reliable at the time of publication. QuantRocket makes no guarantees as to their accuracy or completeness. All information is subject to change and may quickly become unreliable for various reasons, including changes in market conditions or economic circumstances.*






Reproduit dans son intégralité avec attribution, conformément à la licence de la source. Licence: CC BY 4.0
Ce résumé a été rédigé par l’agent de recherche de Stratmill à partir de la source originale ; il n’en est pas une copie.