Choosing Realized Kernel Bandwidth for High-Frequency Volatility
Summary
The document asks how to select the optimal bandwidth for a realized-kernel volatility estimate when the desired output uses one-minute bins. It gives the bandwidth formula from the cited reference and notes that its noise-to-signal input uses integrated variance estimated from 20-minute realized variances. The answer explains that this estimate compares dense and sparse realized variance: sparse sampling is intended to limit microstructure noise, while the denser observations capture it for the noise estimate.
The response says that estimates at intervals below five minutes may be unreliable because variance estimates can become unstable, and points to pre-averaging followed by kernel estimation as an alternative for rolling measures. Another practical option is to accept that a one-minute metric uses information from several preceding hours. These are recommendations cited in the response, not results demonstrated with the questioner's data. The suitable bandwidth and sampling choices depend on the noise characteristics and the target horizon, which the document does not quantify.
Key ideas
- The bandwidth formula depends on a noise-to-signal estimate and the number of observations.
- The noise estimate can compare realized variance from dense sampling with sparse sampling.
- Sparse sampling is used to reduce the influence of microstructure noise when estimating integrated variance.
- The response cautions that very short sampling intervals can destabilize variance estimates.
- Pre-averaging and kernel methods are suggested for rolling estimates, though no data-specific bandwidth is calculated.
Tags
Full text
# Optimal bandwidth for Realized Kernel
# Optimal bandwidth for Realized Kernel
If I want to estimate Realized Kernel for 1 min bins, is there a way to compute the optimal bandwidth? In the reference paper: Realised Kernels in Practice: Trades and Quotes (Ole Barndoff-Nielsen et al.), it provides the following formula for the optimal bandwidth $H^*$:
$$ \hat{H}^* = c^* \hat{\xi}^{4/5} n^{3/5}$$ with $$ \xi^2 \approx \frac{\omega^2}{IV}$$ and $IV$ is estimating by averaging 20 minutes realized variances.
What bothers me here is that $IV$ is calculated by using 20-min Realized Variances, does it mean that we can not determine the optimal bandwidth if we are interested in the Realized Kernel for time bins smaller than 20 minutes?
## Answer by kurtosis (score 1)
https://quant.stackexchange.com/a/57645
This issue with using a kernel to estimate a quantity for a one-minute bin is that you can write $\xi^2$ as $$ \hat\xi^2 = \frac{\overline{RV}_{\text{dense}}}{\overline{RV}_{\text{sparse}}}. $$
The estimator $\overline{RV}_{\text{sparse}} \approx\widehat{IV}\approx \sqrt{T\int_0^Y \sigma_u^4 du}$ samples returns sparsely. The idea is that the sampling period is sufficiently long as to justify a claim that there are no "microstructure noise" (bid-ask bounce) effects at that time scale. Hence Barndorff-Nielsen, Hansen, Lunde and Shephard (2008) use 20-minute returns.
You are also likely to run into issues sampling at periods below five minutes since variance estimators explode with sampling periods that are shorter than five minutes, as shown in Andersen, Bollerslev, Diebold, and Labys (2003).
Perhaps the best advice would be to do some pre-averaging of your data, as advised in Podolskij and Vetter (2006), and then use kernels to give you a rolling estimation. Your best bet for that would be to look at Figueroa-López and Wu (2020, WP). Alternatively, you could just accept that your one-minute metrics include a quantity computed using data from the past few hours.Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.