Skip to content
All library documents

Estimating Covariance from Asynchronously Sampled Financial Series

Article Quant Q&A · Author: L. Bachelier

Summary

The document asks how to extract principal components and correlations from multiple data sets organized at different weekly or monthly periods, with observations taken on several days around an event. The response reframes the main challenge as estimating covariance across variables sampled at different times. Since principal component analysis diagonalizes a covariance or correlation matrix, reliable covariance estimates are a prerequisite for meaningful components.

For asynchronous observations, the response points to the Epps effect, where measured correlations can be distorted when series are not observed simultaneously. It describes a basic approach of calculating over overlapping observation periods and notes that this introduces bias. A cited covariance-estimation method for nonsynchronously observed diffusion processes is said to adjust for and compensate for that bias. The document does not provide formulas or show how the method applies to event-based weekly or monthly data, so the proposed reference may require adaptation to the specific sampling design.

Key ideas

  • Principal component analysis depends on a covariance or correlation matrix.
  • Asynchronous sampling can distort estimates of correlation, a problem associated with the Epps effect.
  • Using only overlapping observation periods is a basic way to estimate covariance from nonsynchronous data.
  • Overlap-based estimation can introduce bias, so adjustments may be needed.
  • The cited method concerns nonsynchronously observed diffusion processes and is not worked through for event-based data.

Tags

Full text
# Can we model components in a set of multivariate multi-period time-series data?


# Can we model components in a set of multivariate multi-period time-series data?












There are N data sets in periods occurring weekly/monthly, across a 10-year historical timeline.

In each period, five dates are observed (labelled a to e), where a denotes the day the period starts/an event occurs (T=0), while b to e denotes subsequent days following the event (T = 2 to 4).

--

An illustration is created to better understand how the components are structured and fed into the formulae for statistical inference.

Question: Is there a method to elicit principal components from the N data sets, and also find correlation?

P.S. This model intends to observe events/numbers (e.g. in the economic calendar) occurring weekly and monthly that affects changes in market prices.

## Answer by lehalle (score 0, accepted)

https://quant.stackexchange.com/a/10292

Your question is more about "how to estimate correlations between variables sampled at different frequencies?" than about PCA. After all, PCA is just diagonalization of the covariance (or correlation) matrix, aiming to obtain principal vectors driving the joint dynamics of your variables in an $L^2$ sense.

Since data are by construction not synchronized at the high frequency rate (i.e. you never have simultaneous trades or quote updates on two different instruments), this issue is celebrated as the Epps effect in intra-day finance (one day I will create a "tag wiki page" to describe it in detail). But it apply to any other covariance computations.

The "simplest" way to counter the Epps effect has been proposed by Ayashi and Yoshida in "On Covariance Estimation of Non-synchronously Observed Diffusion Processes". Thus you should have a look at this paper. They propose adjustments to the following simple principle: "do computations on overlapping periods only". Of course it introduces bias, and they explain how to compensate them.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.