Skip to content
All library documents

How Sampling Frequency Affects Estimated Volatility

Article Quant Q&A · Author: u23489638

Summary

The document compares price-change volatility for TLT across several sampling intervals, from one minute to daily, using two years of closing-price data. It filters observations to regular US market hours and trading days, calculates the standard deviation of price changes, and annualizes it using the square root of the number of periods in a year. The reported annualized estimates are similar across intraday intervals, while the daily estimate is much lower when the author treats a trading day as 24 hours.

The author then recalculates using a 6.5-hour trading session and 252 trading days per year; these adjustments bring the estimates closer together. The exercise raises a useful distinction between calendar time and trading time when scaling volatility. It does not establish whether the remaining differences are statistically meaningful, and the document offers no formal test. Its results may also depend on the instrument, sampling choices, missing observations, and treatment of overnight price changes.

Key ideas

  • Volatility estimates can vary with the sampling interval and the annualization convention.
  • Square-root-of-time scaling depends on how periods per year are defined.
  • Restricting prices to regular trading hours changes the interpretation of daily versus intraday returns.
  • The reported estimates do not include a statistical test of whether their differences are significant.

Tags

Full text
# Instrument volatility scaling as a function of sample rate


# Instrument volatility scaling as a function of sample rate












## Question:

I did some experimentation today to see if volatility changes as a function of the sample frequency.

Prior to starting this experiment, I believed that volatility was a function of sample frequency and that the scaling should be like $\sqrt T$. It might be that I have totally misunderstood something.

What I found by performing my experiment is that volatility is approximatly constant.

I will now explain what I did:

- I obtained data for the closing price of `TLT` which is an instrument which gives exposure to US Treasury Bonds

- I have two years of data, from 2022 and 2023

- The data native sample rate is 1 minute bins, although some bins for very quiet trading days are missing (24th November after 13:00 being an obvious example)

- I resampled this data at intervals of 1 minute, 5 minute, 15 minute, 60 minute and 1440 minute (24 hours)

- I discarded any data which was not within NASDAQ trading hours (09:30 - 16:00) and I discarded any data which was not within NASDAQ trading days, which included both weekends and holidays as determined by `pandas_market_calendars`

- I calculated the change in closing price from the closing price data

- I calculated the standard deviation of changes in closing price

- I calculated the annualized standard deviation of changes in closing price by scaling by $\sqrt N$ where N is the number of periods

For example:

```
annualized_std_24_hour = math.sqrt(256) * std_24_hour
```

because there are (approximately) 256 blocks of 24 hours in a year. (Trading days.)

For the final resampling, data was resampled each day at `15:00 America/New_York`, so that a closing price is taken from within trading hours.

The following dataframe contains my results.

```
   time_period   count       std  annualized_std
0            1  283865  0.054267       32.948647
1            5   57353  0.119761       32.518706
2           15   19601  0.203873       31.960574
3           60    5081  0.403402       31.620142
4         1440     725  1.036977       16.591634
```

There is one obvious outlier here. The resampled at daily frequency data is notably different to the others.

Why is this the case? Am I calculating the volatility (std) here correctly?

## Possible Hypothesis 1:

I had one hypothesis:

- Perhaps because more volume of this instrument is traded between the hours of `09:30` and `16:00` (6.5 hours), and because almost no trades are made during the night (far outside of trading hours) the scaling factor to convert between days and hours is wrong.

- The scaling factor I was using was 24 hours / day. Perhaps it should be 6.5 hours / day.

The dataframe below shows the results for 6.5 hours / day.

```
   time_period   count       std  annualized_std
0            1  283865  0.054267       17.147019
1            5   57353  0.119761       16.923271
2           15   19601  0.203873       16.632810
3           60    5081  0.403402       16.455644
4         1440     725  1.036977       16.591634
```

All of the values for periods shorter than 1 day are over-estimated compared to the value for 1 day.

I understand that this could just be by chance and might not be statistically significant. Looking at the values I don't think it is just random noise.

## Possible Hypothesis 2:

It then occurred to me that NASDAQ has 252 trading days in a year. The difference might become significant at this level of precision.

The same table of values calculated for 252 trading days.

```
   time_period   count       std  annualized_std
0            1  283865  0.054267       17.012531
1            5   57353  0.119761       16.790538
2           15   19601  0.203873       16.502355
3           60    5081  0.403402       16.326578
4         1440     725  1.036977       16.461502
```

I would suggest it is now even harder to tell if there is any significance here.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.