Estimating Realized Volatility from Log Returns
Summary
The document considers how to estimate and plot volatility from a long history of FTSE closing prices. Its code calculates log returns, but the questioner’s approach appears to treat the deviation of each individual return from the overall mean as a volatility observation. The accepted response advises against taking standard deviations of pairs of consecutive returns and instead introduces realized variance: sum squared log returns over a chosen period. Taking the square root gives realized volatility for that period, such as a week or month.
The evidence is a conceptual explanation rather than a comparison against market data or alternative estimators. The note does not provide implementation details for rolling windows, annualization, or adjustments for sampling frequency. In particular, the proposed sum depends on the interval over which returns are aggregated, so results for different periods should not be confused. The brief answer points to realized variance’s mathematical properties but does not explain them or address every detail in the original plotting code.
Key ideas
- Realized variance can be estimated by summing squared log returns over a selected period.
- Taking the square root of that sum gives realized volatility for the period.
- The response cautions against calculating volatility from only pairs of consecutive returns.
- The chosen aggregation period determines the horizon represented by the volatility estimate.
Tags
Full text
# Calculate and plot historical volatility with Python
# Calculate and plot historical volatility with Python
I have downloaded historical data for FTSE from 1984 to now. What I would like to do is to graph volatility as a function of time. What I have written is:
```
import matplotlib.pyplot as plt
import datetime as dt
import numpy as np
import math
lines = [line.rstrip('\n') for line in open("Data.txt")]
a = list(range(len(lines)))
adjClose = [float(i) for i in lines]
adjClose.reverse()
dates = [line.rstrip('\n') for line in open("Date.txt")]
dates.reverse()
x = [dt.datetime.strptime(d,'%Y-%m-%d').date() for d in dates]
dailyVolatility = np.std(np.diff(np.log(adjClose))).round(4)
# Calculate returns
returns = []
for i in range(len(adjClose[1:])):
element = adjClose[i]/adjClose[i-1]
element = math.log(element)
returns.append((element))
returns = returns[1:]
mean_returns = np.mean(returns)
vol = []
for i in range(len(returns)):
element = (returns[i]-mean_returns)**2
element=math.sqrt((element))
vol.append(element)
for i in range(len(vol)):
vol[i]*=math.sqrt(252)
x = x[2:]
plt.plot(x,vol)
plt.show()
```
So I first load the data and then calculate the log returns and also take the average; moreover, I calculate the standard deviation for every pair of numbers in my log returns. Is my reasoning correct? In this case I haven't averaged at all for the standard deviation formula, since N-1 = 2-1=1.
## Answer by Matt Brigida (score 2, accepted)
https://quant.stackexchange.com/a/20645
I don't have enough reputations points to comment, so I'll put this into an answer.
I am not sure what this means, "standard deviation for every pair of numbers in my log returns." but it sounds like you are taking the standard deviation of two consecutive returns. If that is the case, I would not do that.
I think you want "realized variance". This is just the sum of squared log returns. You can then take the square root of this sum to get realized volatility. If you sum over a week or month, you get the realized volatility over that week or month.
See the Wikipedia article for the nice mathematical properties of realized variance.Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.