Sampling Schemes and Bias in Intraday Realized Volatility
Summary
The document discusses how to estimate realized volatility from irregularly spaced tick data over a rolling window. It distinguishes calendar-time sampling, which uses fixed time intervals, from tick-time sampling, which updates on each transaction or at a chosen tick frequency. A cited comparison reports that tick-time sampling can estimate realized variance well in some settings, particularly when microstructure noise is low, but its noise may be strongly autocorrelated and can bias estimates.
The discussion also explains why sampling times chosen in response to price movements can distort variance estimates, while random observation times that are independent of the price path need not invalidate the estimator. For multiple assets, it outlines previous-tick, interpolation, refresh-time, and generalized synchronization approaches, noting tradeoffs in frequency and use of future observations. The practical recommendation is to begin with sparse calendar-time sampling to obtain synchronized returns and limit microstructure effects. The document is an overview rather than a rigorous derivation and does not establish a single best method for every market or sampling objective.
Key ideas
- Tick-time sampling measures returns over a number of observations rather than a fixed calendar duration.
- Tick-time realized variance can be useful, but autocorrelated microstructure noise may bias estimates.
- Sampling times selected in response to price movements can alter measured variance.
- Previous-tick, refresh-time, and other synchronization methods address asynchronous observations across assets.
- Sparse calendar-time sampling is presented as a practical starting point for reducing noise and aligning returns.
Tags
Full text
# Measuring the rolling volatility of intraday tick data (unevenly spaced time series)
# Measuring the rolling volatility of intraday tick data (unevenly spaced time series)
I am wanting to measure the rolling volatility with about a 15 minute window from tick observations and then update it periodically as ticks come in.
I am aware of the formula:
$$\sigma^2 = \frac{1}{T_N - T_0}\sum_i^M \frac{\log\left(\frac{S_{t_{i+1}}}{S_{t_i}}\right)}{t_{i+1} - t_i}$$
So in the case of 15 minutes, we would normalise the price ticks to 1-minut, and let $T_N=15$ and $T_0=0$, and sum up to $M=\text{len}(t_m - t_i), \quad s.t. \quad t_m =T_N$.
Supposedly this is only affective if the time spaces area deterministic, but in my case, $t_{i+1} - t_i$ is random. Why can’t I just update $\sigma^2$ as a new tick comes in using that formula and remove any $t_i$ not in the 15 minute window. Why do the spaces being random suddenly make this equation not the maximum likelihood estimate of the variance?
I’m not trying to do any forecasting, just merely trying to measure the realised volatility that just occurred.
## Answer by Pleb (score 2, accepted)
https://quant.stackexchange.com/a/80860
My response will focus on the choice of sampling scheme rather than the objectives of your statistical study.
## High-Frequency Sampling Schemes, briefly:
Yes, you can update your formula based on the number of ticks instead of calendar time. Hence, the measure will be the average over the last 15 ticks instead of the last 15 minutes. This approach is known as tick-time sampling (or raw tick-time sampling), which operates asynchronously in relation to "clock" time. Tick-time sampling can also be applied sparsely, such as by sampling every 5th tick instead of every single one.
From the article of Griffin, Jim E., and Roel CA Oomen (2008), they conduct a comparison study on tick-time sampling (against transaction-time sampling) and find the following:
- Tick-time sampling tends to outperform in estimating Realized Variance (RV), especially when there is lower microstructure noise, fewer ticks or less frequent efficient price movements.
- However, they emphasize the need to account for the strong autocorrelation in the noise when estimating RV in tick-time. This rules out noise-robust estimators with independent noise assumptions.
Therefore, while tick-time sampling can be effective, it is important to note that it may introduce bias into your volatility or variance estimator, especially if the microstructural noise is not properly accounted for.
### Synchronizing the Data
When working in a multivariate setting, it is essential to synchronize the tick data across assets. Below are brief descriptions of some documented synchronization techniques:
- Previous-Tick Method (L. Zhang (2011)): In this method, for each asset, the most recent tick is used to fill in any missing data when synchronizing across different assets. This approach ensures that the latest available data point is used across all assets at each sampling point.
- Linear Interpolation Method (Dacorogna, M.M et al. (2001) pp. 37 - 39) This technique constructs a linear interpolation between the previous tick and the future tick. This is not used much in literature because it involves using data from the future to construct the interpolation. The authors also argue that the difference between the Interpolation method and Previous-Tick method are often negligible, making Previous-Tick the overall preferred choice of the two.
- Refresh-Time Sampling (Barndorff-Nielsen, Ole E., et al. (2011)): A new common time axis is constructed where all assets have at least one tick before each sampling point. This ensures simultaneous sampling across assets, but the downside is that the sampling frequency is determined by the least liquid asset, which could reduce the overall frequency of data points. An illustrative example is provided below.$^\star$
- Generalized Synchronization Method (Aït-Sahalia, Yacine et al. (2010)): This method allows for more flexibility by sampling an arbitrary observation for asset $i$ in the time interval $(\tau_{j-1}, \tau_j]$. Depending on how the observations are sampled, this method generalizes both the Previous-Tick Method and Refresh-Time Sampling: If you sample the last observed tick for all assets before $\tau_j$, it behaves like Refresh-Time Sampling, while selecting the last observed tick for each asset mimics the Previous-Tick Method. Further details can be found in the referenced paper.
### Final Remarks
In this context, I would still recommend sparse sampling based on a calendar-time scheme to obtain synchronous returns while reducing or eliminating microstructure noise (atleast to begin with). This approach avoids the need to use bias-corrected volatility estimators, thus simplifying the statistical analysis.
In conclusion, no matter the sampling scheme, it is important to keep it independent from the statistical analysis in order to ensure consistency. Furthermore, by maintaining a uniform sampling scheme across the analysis, you retain the flexibility to switch between schemes if necessary in the future.
$^\star$ Previous Tick and Refresh-Time sampling can seem identical. To understand the differences consider a bivariate setup: $X$ ticks at $t_1,t_3,t_5$, and Y at $t_3,\: t_7$. For Previous-tick, we would sample whenever one of the assets ticks, ie. $PTS = [t_3,t_5,t_7]$, however for Refresh-Time we sample when all assets have ticked atleast once, hence $RTS = [t_3, t_7]$.
## Answer by Rylan (score 0)
https://quant.stackexchange.com/a/80813
I don't think the time step needs to be determinsitic per se. You can imagine performing this experiment with fixed time steps, and halfway through, flipping a coin to determine whether you want to partition the remaining $m$ time steps into $2m$ steps by adding an additional time in the middle of each of them.
On the other hand, if the selection of time steps is allowed to depend on the movement of the stock itself, you can imagine a few ways of choosing the time steps that will impact variance. For example:
- Choosing $t_i$ to be the ith time the stock hits a price $c \in \mathbb{R}$
- If you observe the whole path of the stock before choosing time steps, you can likely pick out the time steps to artifically inflate/deflate the variance
et cetera.
I'm only using intuitive arguments here; if anyone is interested I invite them to edit my answer to fill out the details and rigour.Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.