Estimating Return Distributions from Overlapping Windows
Summary
The question examines empirical distributions of rolling stock returns calculated over overlapping windows. Because adjacent returns share observations, the series is serially correlated. The author considers thinning the data by sampling every window length and asks how to choose the starting offset, or whether the offset-specific samples can be combined while avoiding autocorrelation effects.
The replies do not provide a settled method for combining these samples. One suggests using autocorrelation diagnostics to identify a lag for thinning. Another argues that the issue depends on the statistical framework and highlights structural changes such as shifts in rates, capital structure, dividends, taxes, or market conditions. That reply also gives a particular heavy-tailed return distribution under restrictive assumptions, but its broad claims about autocorrelation and inference are not substantiated in the exchange. The document is best read as a question about dependence and regime changes, not as validated guidance for distribution estimation.
Key ideas
- Rolling returns share observations when their windows overlap, creating serial dependence.
- Sampling at intervals equal to the window length removes direct overlap but leaves a choice of starting offset.
- Autocorrelation diagnostics can help assess dependence at different lags, though the exchange does not establish a complete procedure.
- Structural breaks may affect empirical return distributions independently of overlap.
Tags
Full text
# Estimating distribution of rate of return
# Estimating distribution of rate of return
Let $f[t]$ be the price of a stock at time $t$. We can calculate the rolling rate of return of the stock in a window of length $n$ by computing: $$r[t] = \frac{f[t] - f[t-n]}{f[t-n]}$$ $r[t]$ is serially correlated, since neighboring values overlap by $n$ samples. I want to estimate a distribution for $r[t]$ empirically that is unaffected by autocorrelation. One way to do this is to thin the series (i.e., sample every $n$th value). This effectively means the windows used to create the $r[t]$'s that get sampled don't overlap.
- How do I decide whether to select samples $1, n+1, 2n+1, \ldots$ or $2, n+2, 2n+2, \ldots$ or $3, n+3, 2n+3, \ldots$, and so on?
- Is there a way to make use of all the samples (e.g., by creating distributions from each of the sets of samples above and then combining them) that is not sensitive to the autocorrelation?
## Answer by Deno (score 0)
https://quant.stackexchange.com/a/69499
If you are interested in finding out non correlated sequence/items in your time series , why don't u just apply ACF/PACT to find out lagged number (p). And then, I reckon, this p +1 would be your n, as there won't be much autocorrelation between these items.
## Answer by Dave Harris (score -1)
https://quant.stackexchange.com/a/63492
I have done that. The distribution if there are no dividends, mergers or bankruptcy and if liquidity costs are ignored is $$\Pr(r|r^*;\gamma)=\left[\frac{\pi}{2}+\tan^{-1}\left(\frac{r^*}{\gamma}\right)\right]^{-1}\frac{\gamma}{\gamma^2+(r-r^*)^2}.$$
It has no expected value. You can find a reduced form discussion here https://youtu.be/R3fcVUBgIZw.
If you attempt to tackle it directly as a ratio distribution, you end up needing data we never collected.
You can find a general solution here.
Autocorrelation is irrelevant for this discussion. Autocorrelation is an artifact of certain time series methods but not others. For example, Bayesian methods are not impacted by autocorrelation, so it is all but ignored. It is only important in Frequentist statistics because it interferes with inferences.
Your formula is not a time series. If it were, then a convolution would first have to happen in the numerator for the difference. However, as the right side of the numerator is also the denominator, it is no different than subtracting one. It is a shift variable and the difference of the values in the numerator can be ignored.
Your largest problem will not be autocorrelation, which can be ignored, but structural breaks due to interest rate, capital structure, dividend, tax and market changes. I strongly recommend a Bayesian method as you are almost only restricted to Theil's regression or quantile regression if you find that too computationally expensive.
There is no expectation so you cannot use squares minimizing routines, so also, no variance.Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.