Skip to content
All library documents

Regression with Overlapping Returns: Stationarity and Standard Errors

Article Quant Q&A · Author: rocketman

Summary

The document considers regression when both predictor and outcome are monthly rolling annualized returns, which can produce serial correlation because adjacent observations share much of their data. It contrasts differencing both series with fitting a levels regression and using heteroskedasticity and autocorrelation consistent standard errors. The replies emphasize that these choices address different problems: Newey–West adjusts inference for serial dependence but does not make nonstationary variables suitable for regression or prevent spurious relationships.

Suggested alternatives include using nonoverlapping observations, bootstrap standard errors, or a state-space model. One reply recommends checking whether changes in the two series are related and testing for cointegration if a shared long-run relationship is plausible. These are conditional suggestions rather than a settled procedure. The appropriate method depends on the research question, overlap length, and time-series properties; the document provides no data analysis or comparative results.

Key ideas

  • Overlapping rolling returns induce serial correlation in adjacent observations.
  • HAC standard errors address serial correlation in inference but do not fix nonstationarity.
  • Differencing can test whether changes in the predictor relate to changes in the outcome.
  • Nonoverlapping observations, bootstrapping, and state-space methods are alternatives to consider.
  • Cointegration may matter when nonstationary series plausibly share a long-run relationship.

Tags

Full text
# How to adjust regression for rolling returns?


# How to adjust regression for rolling returns?












I have a predictor variable (x) and dependent variable (y). Both are monthly rolling annualized returns, which naturally induces significant autocorrelation in x and y. They both also fail to be stationary under ADF tests (is this a natural consequence of their construction of being rolling returns?) In terms of regression, is it more appropriate to 1) difference both x and y first or 2)run the normal regression (y~x) and simply adjust standard errors by using heteroskedasticity and autocorrelation (HAC) consistent covariance matrix estimation methods(such as NeweyWest)? Under what circumstances is each of the methods more appropriate? Please leave any references should you have any.

## Answer by Kiwiakos (score 1)

https://quant.stackexchange.com/a/30182

It depends how large the overlapping interval is. Conceptually an infinite rolling window is equivalent to the level, and no one would suggest to 'regress on levels and apply Newey West'.

I think NW is 'robust' in the presence of relatively mild autocorrelation, not a panacea that will give the correct standard errors.

If you use, say dailly returns aggregated to montly rolling ones, then there is 90% overlap in your consecutive observations (and expected autocorrelation).

My ordered preferences would be: 4. Newey West or similar 3. Bootstrap standard errors 2. Cast as state space and apply Kalman Filter 1. Drop overlapping observations and use only the non-overlapping ones

## Answer by Andrea (score 0)

https://quant.stackexchange.com/a/29704

What is your goal? I assume it is to find if/how y is caused by x. You really ought to make sure that y and x are stationary. There are only a few cases in which linear regression makes sense when they are not stationary (e.g. x is a deterministic trend term) and even in those cases, the statistical properties of the coefficient estimators are different than usual.

If the regressor is not stationary, the results can just be wrong, as in the case of spurious regression. The Newey-West method is not used to solve problems of stationarity but rather of serial correlation which is not the main issue in your case. It just changes the estimated errors, but if x is not stationary, the coefficient estimates may well be biased.

My hunch is that it'd be enough to inspect if changes in x are related to changes in y. Do this by differencing the two series. However, if you suspect a long term common trend, do a test of cointegration too (if x is just one variable and not a set of many variables, look for the Engle-Granger test, it's very simple).

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.