Skip to content
All library documents

Choosing Samples for Annual Return Distribution Analysis

Article Quant Q&A · Author: qwer

Summary

The document compares two ways to study annual return distributions: calculating a rolling one-year return for every trading day, or sampling one return per year from a fixed annual date. The example uses long-run S&P data and reports broadly similar central estimates and volatility for the two approaches, while the rolling daily series has a wider observed range. The question is whether the extra daily observations add useful information or mostly repeat overlapping periods.

The response says the choice depends on the analysis goal. Daily rolling observations are strongly autocorrelated, while annual samples leave few observations. Suggested approaches include correcting uncertainty estimates for autocorrelation, using simulation or block bootstrapping with weekly or monthly blocks, applying extreme value methods to study tails, and using hidden-state models to represent crisis and recovery regimes. A further option is to fit a daily return distribution and simulate its annual aggregation. Each method depends on assumptions about the data and the target quantity, such as average returns, volatility, or extreme losses.

Key ideas

  • Daily rolling annual returns overlap and introduce autocorrelation into the sample.
  • Sampling one return per year avoids that overlap but yields few observations.
  • The appropriate sampling method depends on the statistic and market behavior being studied.
  • Block bootstrapping, autocorrelation adjustments, tail models, regime models, and aggregation simulations are possible alternatives.

Tags

Full text
# Looking at distribution of yearly returns of time series


# Looking at distribution of yearly returns of time series












For S&P, or any time series for that matter. When doing analysis on the distribution of the yearly returns, should I be looking at 1) the daily year over year values, 2) pick some starting point like December 31st of year YYYY and then look at YoY, or 3) something else.

My intuition is that for 1) you are artificially creating many "YoY" values that look the same because the previous days YoY will not change much given a one day move, but at the same time 2) seems a little arbitrary because youre just selecting a single point as a reference point (you could have chosen first of year or middle of year) and in fact 1) contains 2)

Doing both for S&P returns from Yahoo from 1950s doesnt seem like it shows a big difference in terms of mean or vol parameters (the daily YoY has wider min/max as you would expect):

```
> smry
          dailyYoY     yearly
Min.    -0.6699000 -0.4499000
1st Qu. -0.0120600 -0.0061320
Median   0.0937500  0.0834800
Mean     0.0724200  0.0736000
3rd Qu.  0.1783000  0.1681000
Max.     0.5222000  0.3274000
Std Dev  0.1555526  0.1487778
```

It seems like one could argue that the yearly is just a sampled version of the dailyYoY, but which should I be using? I'm leaning more towards dailyYoY, because the other method ignores information I already know

## Answer by Bram (score 1)

https://quant.stackexchange.com/a/34939

In my view there isn't a good answer to this question; the daily method introduces autocorrelation between your returns and the yoy samples leave you with little data.

I would let whatever you choose (and I'll mention some of the other things that you could think about below) mainly be guided by what you are trying to achieve with you analysis, e.g. do you care about the mean, standard deviation, worst loss, or still something else? Do you believe that recent returns are more relevant than those from 50 years ago?

So some ideas for what you could do beyond what you mentioned (and depending on your goals these could be reasonable or highly inappropriate):

As hinted at in the comments on the question, for the daily points there are some methods available to adjust e.g. the estimator for the standard deviation of yearly returns to correct for the autocorrelation. So that's one way to go. If there isn't an analytical correction, you could come up with one based on simulation from what you consider to be an appropriate idealized distribution.

Another way to go could be bootstrapping. If daily bootstrapping isnt useful for your application, a good trade-off might be to divide your data in weekly or monthly blocks and just sample 52 cq 12 of those.

If you're really interested in the tails, you could also look at something like extreme value theory for estimating returns.

Yet another way to go, if you're interested in crises (as people looking at very long time series often are) could be to model the data as bring emitted from a hidden Markov model where markets can be in a few states (e.g. {normal, crisis, recovery}) and estimate return distributions and transition probabilities from them and then sample from you this model.

Yet another alternative could be looking at what distribution fits your daily data well and simulate how that aggregates from daily to yearly for certain percentiles of the distribution.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.