Skip to content
All library documents

Hayashi–Yoshida Covariance and Realized Correlation

Article Quant Q&A · Author: apocalypsis

Summary

The document asks how the Hayashi–Yoshida estimator for asynchronously sampled asset returns relates to realized variance and Pearson correlation. It describes realized variance as the sum of squared returns and the Hayashi–Yoshida covariance estimator as a sum of return products for observation intervals that overlap. Dividing this covariance estimate by the square root of the two realized variances gives a normalized realized correlation measure.

The question reports an empirical comparison between realized variance per observation and a sample standard deviation, but does not resolve the scaling or interpretation. In particular, realized variance is an estimate of integrated variance over a time interval, while a standard deviation of sampled returns depends on the sampling frequency and normalization; the two are not directly interchangeable without specifying those conventions. The document provides no derivation, estimator properties, or empirical results for correlation, so it leaves open how the normalized estimator compares with Pearson correlation under asynchronous sampling and finite-sample conditions.

Key ideas

  • Realized variance is formed by summing squared returns over the observation period.
  • The Hayashi–Yoshida covariance pairs returns whose observation intervals overlap.
  • Normalizing the covariance estimate by the two realized variances yields realized correlation.
  • Comparisons with sample standard deviation require clear sampling and normalization conventions.

Tags

Full text
# realized correlation estimation


# realized correlation estimation












I'm trying to implement the Hayashi - Yoshida estimator for correlation (T. Hayashi, N. Yoshida: On covariance estimation of non-synchronously observed diffusion processes, 2005) and there's something I'm missing with respect to realized correlation and realized variance. Assuming that $Y = \log P$, and that I have access to $N$ discrete (asynchronous) observations for two assets on a time grid $[0,t]$, the HY estimator is as follows:

$$RV^{(i)}_{[0,t]}=\sum_{\tau \in[0,t]}\left(Y^{(i)}_{\tau}-Y^{(i)}_{\tau-1}\right)^2$$

$$RC^{(1,2)}_{[0,t]}=\frac{\sum_{\tau_1 \in[0,t]}\sum_{\tau_2 \in[0,t]}\left(Y^{(1)}_{\tau_1}-Y^{(1)}_{\tau_2-1}\right)\left(Y^{(2)}_{\tau_2}-Y^{(2)}_{\tau_2-1}\right)\mathbb{I}_{[\tau_1-1, \tau_1]\cap[\tau_2-1, \tau_2]}}{\sqrt{RV^{(1)}_{[0,t]}\cdot RV^{(2)}_{[0,t]}}}$$

where $RV$ estimates the realized variance and $RC$ the realized correlation over the period. Now, both these estimators should be unbiased but I was wondering, how do they relate exactly to the variance and correlation (Pearson)? I found empirically that $\sqrt{RV/n}\approx \sigma$ (sample std), which I guess makes sense (I'm not completely sure why), but I can't seem to find any similar scaling relationship between $RC$ and $\rho$.

The data is tick-by-tick over one day (not sure if it's relevant).

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.