Skip to content
All library documents

Estimating Autocorrelation: Choosing Means for the Sample Formula

Article Quant Q&A · Author: Fiatpanda2000

Summary

The document explains why common sample autocorrelation formulas can differ in how they estimate means and variances. It starts from the probabilistic definition: autocorrelation is the correlation between a process and its lagged version. In a finite sample, each expectation in that definition must be replaced with an estimator, and the choice of observations used for those estimates affects the resulting formula.

One approach estimates the mean using the observations that form the lagged pairs; another can use all available observations, including the extra observations at the sample edge. The answer argues that using more data may be suitable when the process is stationary, while a restricted estimate may be preferable when it is not. Both formulas estimate the same population quantity, but neither is universally best: the appropriate choice depends on the process and assumptions. The exchange does not give a comparative empirical evaluation, and stationarity remains a key caveat.

Key ideas

  • Autocorrelation is the correlation between a process and a shifted version of itself.
  • Sample formulas differ because they estimate the means and variances with different sets of observations.
  • Using all available observations can help under stationarity assumptions.
  • For a nonstationary process, estimating from the paired observations may be more appropriate.
  • The best estimator depends on the underlying distribution and assumptions.

Tags

Full text
# What's the right autocorrelation formula?


# What's the right autocorrelation formula?












I'm trying to see the influence of autocorrelation in my processes and to do so I have to compute it, however it seems to be hard to find a coherent formula over the web. I found pretty much two different ones which are those :

Even tough those two formulas are very close between each others, they seem to disagree on the way of computing the means used at various places within the formulas.

I struggle to figure out which one is the correct one even if I might feel like the first one is correct, compared to the second one, but it is just pure intuition and not mathematically justified.

Can anyone help me figure out which one is the correct one and why ? Thanks everyone for your help.

## Answer by lehalle (score 1)

https://quant.stackexchange.com/a/70498

the autocorrelation is the correlation of a process $X$ and its lagged version, hence you have to consider it from a probabilistic viewpoint. Use the $\sigma_\ell$ notation for the operator that shifts a process: $${\cal A}(X;\ell):=\frac{\mathbb{E}\big((X - \mathbb{E}X) \cdot (\sigma_\ell\circ X - \mathbb{E}\sigma_\ell\circ X)\big)}{\sqrt{\mathbb{E}(X - \mathbb{E}X)^2 \cdot \mathbb{E}(\sigma_\ell\circ X - \mathbb{E}\sigma_\ell\circ X)^2}}.$$

This is the true definition. Now you need to use empirical estimators for all these quantities.

Usually we take $\frac{1}{N}\sum_{n=1}^N X_n$ as an estimator for $\mathbb{E}X$ over a sample of size $N$.

I let you replace and you will get the first formula you snapshotted in your question.

Now think about the set of information you have: you know $X_t$ from $t=1$ to $t=N+\ell$. When it is about estimating quantities that are not a function of the dependencies between $X$ and $\sigma_\ell\circ X$, isn't it better to use all the available information?

i.e. $\frac{1}{N+\ell}\sum_{n=1}^{N+\ell} X_n$ may be a better estimator of $\mathbb{E}X$, no? simply because you use more observations (of course it is submitted to some stationarity assumptions, like it has been underlined in one remark).

Here we talk about $\mathbb{E}X$, $\mathbb{E}\sigma_\ell\circ X$, $\mathbb{E}(X - \mathbb{E}X)^2$ and $\mathbb{E}(\sigma_\ell\circ X - \mathbb{E}\sigma_\ell\circ X)^2$ that are all concerning $X$ only. You can estimate them using as many observation as possible. Of course if you care about outliers, you can also use a bootstrap method (especially for the variance terms, since bootstrap is designed that for), or any method you like.

If you do so, then you recover the last formula of your question:

- at the numerator you use $\bar X=1/N\sum_n X_n$ for $\mathbb{E}X$ and for $\mathbb{E}\sigma_l\circ X$: $$\mathbb{E}\big((X - \mathbb{E}X) \cdot (\sigma_\ell\circ X - \mathbb{E}\sigma_\ell\circ X)\big)\simeq \mathbb{E}\big((X - \bar X) \cdot (\sigma_\ell\circ X - \bar X)\big).$$

- the denominator boils down to $$\sqrt{\mathbb{E}(X - \mathbb{E}X)^2 \cdot \mathbb{E}(\sigma_\ell\circ X - \mathbb{E}\sigma_\ell\circ X)^2}\simeq \mathbb{E}(X - \bar X)^2.$$

- When you put them together $${\cal A}(X;\ell)\simeq \frac{\sum_n (X_n - \bar X) (X_{n+\ell}-\bar X)}{\sum_n (X_n - \bar X)^2}.$$

What is important to understand is that: both formula are ok since they are estimators of the same statistic ${\cal A}(X;\ell)$, that is the true formula. It depends how you want to build the estimators. Which one is the best? it depends on the true (not known) distribution of the $X_t$.

I would say that

- if $X$ is stationary, the second one is the best,

- whereas if $X$ is not stationary, the first one may be better.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.