עבור לתוכן
כל מסמכי הספרייה

שימוש בהיסטוגרמות, תרשימי פיזור ותרשימי קו בניתוח פיננסי

מחברת הרצאות Quantopian

סיכום

שיעור מבוא זה מסביר כיצד תרשימים נפוצים יכולים לסייע לחוקרים לבדוק נתונים פיננסיים ולזהות מבנה אפשרי או בעיות בנתונים. באמצעות מחירי יומיים של שתי מניות US כדוגמאות, הוא מדגים היסטוגרמות להתפלגויות אמפיריות, היסטוגרמות מצטברות לשכיחויות מצטברות, תרשימי פיזור לקשרים בין תצפיות מזווגות ותרשימי קו לערכים לאורך זמן. הוא גם מראה מדוע תשואות בדרך כלל מועילות יותר מרמות מחירים בתרשימי התפלגות: המחירים אינם נייחים ועשויים לנוע בטווח רחב.

הדוגמאות מבחינות בין חקירה חזותית לבין בדיקה פורמלית. תרשים יכול לסייע בגיבוש השערה, אך דפוסים נראים לעין עשויים לנבוע מהטיית אישוש, והתפלגות תשואות היסטורית אינה מבטיחה התנהגות זהה בעתיד. תרשימי קו מחברים תצפיות שנדגמו ועלולים להסתיר תנועות ביניהן; תרשימי פיזור משמיטים את התאריכים של הערכים המזווגים. השיעור אינו מציע אסטרטגיה או מבחן סטטיסטי לאימות קשר, ולכן טענות על מודלים דורשות ניתוח זהיר יותר מבדיקה חזותית בלבד.

רעיונות מרכזיים

  • השתמשו בהיסטוגרמות כדי לבדוק את התפלגות התדירויות של תצפיות, ובהיסטוגרמות מצטברות כדי לראות שכיחויות מצטברות.
  • בתרשימי התפלגות של סדרות פיננסיות, תשואות שימושיות לעיתים קרובות יותר מרמות מחיר שאינן נייחות.
  • תרשימי פיזור מציגים ערכים מזווגים אך משמיטים את התאריכים המשויכים לכל תצפית.
  • תרשימי קו מקלים על מעקב אחר התפתחות לאורך זמן, אך עלולים להסתיר שינויים בין נקודות שנדגמו.
  • השתמשו בתרשימים לחקירה ולגיבוש השערות, ולא כהוכחה שדפוס או מודל ימשיכו להתקיים.

תגיות

הטקסט המלא
# Graphical Representations of Data


<a href="https://www.quantrocket.com"><img alt="QuantRocket logo" src="https://www.quantrocket.com/assets/img/notebook-header-logo.png"></a>

© Copyright Quantopian Inc.<br>
© Modifications Copyright QuantRocket LLC<br>
Licensed under the [Creative Commons Attribution 4.0](https://creativecommons.org/licenses/by/4.0/legalcode).<br>
<a href="https://www.quantrocket.com/disclaimer/">Disclaimer</a>

***
[Quant Finance Lectures (adapted Quantopian Lectures)](Introduction.ipynb) › Lecture 5 - Plotting Data
***

# Graphical Representations of Data
By Evgenia "Jenny" Nitishinskaya, Maxwell Margenot, and Delaney Granizo-Mackenzie

<a href="https://youtu.be/nKq_wz3Qk8w?si=gtDPpJ3fxhfEyY4d&t=91" target="_blank">Quantopian video for this lecture ↗</a>

Representing data graphically can be incredibly useful for learning how the data behaves and seeing potential structure or flaws. Care should be taken, as humans are incredibly good at seeing only evidence that confirms our beliefs, and visual data lends itself well to that. Plots are good to use when formulating a hypothesis, but should not be used to test a hypothesis.

We will go over some common plots here.

```python
# Import our libraries

# This is for numerical processing
import numpy as np
# This is the library most commonly used for plotting in Python.
# Notice how we import it 'as' plt, this enables us to type plt
# rather than the full string every time.
import matplotlib.pyplot as plt
```

## Getting Some Data

If we're going to plot data we need some data to plot. We'll get the pricing data of Apple (AAPL) and Microsoft (MSFT) to use in our examples. First, we will look up the sids for AAPL and MSFT:

```python
from quantrocket.master import get_securities
securities = get_securities(symbols=["MSFT", "AAPL"], fields=["Sid", "Symbol"], vendors='usstock')
securities
```

### Data Structure

Knowing the structure of your data is very important. Normally you'll have to do a ton of work molding your data into the form you need for testing. QuantRocket has done a lot of cleaning on the data, but you still need to put it into the right shapes and formats for your purposes.

In this case the data will be returned as a pandas dataframe object. The rows are timestamps, and the columns are the sids for AAPL and MSFT.

```python
from quantrocket import get_prices

start = '2014-01-01'
end = '2015-01-01'

data = get_prices("usstock-free-1min", data_frequency="daily", sids=securities.index.tolist(), fields='Close', start_date=start, end_date=end)
data = data.loc["Close"]
data.head()
```

For convenience, let's rename the columns to be symbols instead of sids.

```python
sids_to_symbols = securities.Symbol.to_dict()
data = data.rename(columns=sids_to_symbols)
data.head()
```

Indexing into the 2D dataframe will give us a 1D series object. The index for the series is timestamps, the value upon index is a price. Similar to an array except instead of integer indices it's times.

```python
data['MSFT'].head()
```

## Histogram

A histogram is a visualization of how frequent different values of data are. By displaying a frequency distribution using bars, it lets us quickly see where most of the observations are clustered. The height of each bar represents the number of observations that lie in each interval. You can think of a histogram as an empirical and discrete Probability Density Function (PDF).

```python
# Plot a histogram using 20 bins
plt.hist(data['MSFT'], bins=20)
plt.xlabel('Price')
plt.ylabel('Number of Days Observed')
plt.title('Frequency Distribution of MSFT Prices, 2014');
```

### Returns Histogram

In finance rarely will we look at the distribution of prices. The reason for this is that prices are non-stationary and move around a lot. (For more info on non-stationarity please see Lecture 43: *Integration, Cointegration, and Stationarity*.) Instead we will use daily returns. Let's try that now.

```python
# Remove the first element because percent change from nothing to something is NaN
R = data['MSFT'].pct_change()[1:]

# Plot a histogram using 20 bins
plt.hist(R, bins=20)
plt.xlabel('Return')
plt.ylabel('Number of Days Observed')
plt.title('Frequency Distribution of MSFT Returns, 2014');
```

The graph above shows, for example, that the daily returns of MSFT were above 0.03 on fewer than 5 days in 2014. Note that we are completely discarding the dates corresponding to these returns. 

IMPORTANT: Note also that this does not imply that future returns will have the same distribution.

### Cumulative Histogram (Discrete Estimated CDF)

An alternative way to display the data would be using a cumulative distribution function, in which the height of a bar represents the number of observations that lie in that bin or in one of the previous ones. This graph is always nondecreasing since you cannot have a negative number of observations. The choice of graph depends on the information you are interested in.

```python
# Remove the first element because percent change from nothing to something is NaN
R = data['MSFT'].pct_change()[1:]

# Plot a histogram using 20 bins
plt.hist(R, bins=20, cumulative=True)
plt.xlabel('Return')
plt.ylabel('Number of Days Observed')
plt.title('Cumulative Distribution of MSFT Returns, 2014');
```

## Scatter plot

A scatter plot is useful for visualizing the relationship between two data sets. We use two data sets which have some sort of correspondence, such as the date on which the measurement was taken. Each point represents two corresponding values from the two data sets. However, we don't plot the date that the measurements were taken on.

```python
plt.scatter(data['MSFT'], data['AAPL'])
plt.xlabel('MSFT')
plt.ylabel('AAPL')
plt.title('Daily Prices in 2014');
```

```python
R_msft = data['MSFT'].pct_change()[1:]
R_aapl = data['AAPL'].pct_change()[1:]

plt.scatter(R_msft, R_aapl)
plt.xlabel('MSFT')
plt.ylabel('AAPL')
plt.title('Daily Returns in 2014');
```

## Line graph

A line graph can be used when we want to track the development of the y value as the x value changes. For instance, when we are plotting the price of a stock, showing it as a line graph instead of just plotting the data points makes it easier to follow the price over time. This necessarily involves "connecting the dots" between the data points, which can mask out changes that happened between the time we took measurements.

```python
plt.plot(data['MSFT'])
plt.plot(data['AAPL'])
plt.ylabel('Price')
plt.legend(['MSFT', 'AAPL']);
```

```python
# Remove the first element because percent change from nothing to something is NaN
R = data['MSFT'].pct_change()[1:]

plt.plot(R)
plt.ylabel('Return')
plt.title('MSFT Returns');
```

## Never Assume Conditions Hold

Again, whenever using plots to visualize data, do not assume you can test a hypothesis by looking at a graph. Also do not assume that because a distribution or trend used to be true, it is still true. In general much more sophisticated and careful validation is required to test whether models hold. Plots are mainly useful when initially deciding how your models should work.

---

**Next Lecture:** [Means](Lecture06-Means.ipynb) 

[Back to Introduction](Introduction.ipynb) 

---

*This presentation is for informational purposes only and does not constitute an offer to sell, a solicitation to buy, or a recommendation for any security; nor does it constitute an offer to provide investment advisory or other services by QuantRocket LLC ("QuantRocket"). Nothing contained herein constitutes investment advice or offers any opinion with respect to the suitability of any security, and any views expressed herein should not be taken as advice to buy, sell, or hold any security or as an endorsement of any security or company.  In preparing the information contained herein, the authors have not taken into account the investment needs, objectives, and financial circumstances of any particular investor. Any views expressed and data illustrated herein were prepared based upon information believed to be reliable at the time of publication. QuantRocket makes no guarantees as to their accuracy or completeness. All information is subject to change and may quickly become unreliable for various reasons, including changes in market conditions or economic circumstances.*
![notebook output](figures/p1_1.png)
![notebook output](figures/p1_2.png)
![notebook output](figures/p1_3.png)
![notebook output](figures/p1_4.png)
![notebook output](figures/p1_5.png)
![notebook output](figures/p1_6.png)
![notebook output](figures/p1_7.png)

מוצג במלואו בציון המקור ובהתאם לרישיון שלו. רישיון: CC BY 4.0

הסיכום נכתב בידי סוכן המחקר של Stratmill על סמך המקור; הוא אינו העתק של המקור.