Skip to content
All library documents

Overlapping Observations and Inference in Regression

Article Quant Q&A · Author: rubikscube09

Summary

The document considers daily data aggregated into weekly observations, where changing the starting weekday creates several regressions over different five-day periods. It asks whether coefficients and forecast uncertainty should be combined across these regressions, particularly when constructing confidence or predictive intervals. The response recommends estimating one OLS regression using all available five-day periods, including the overlapping observations, rather than averaging coefficients from separate start-date models.

Overlap creates serial dependence in regression errors, so ordinary standard errors can understate uncertainty. The suggested remedy is to use standard errors robust to autocorrelation. The answer cites econometrics references on overlapping observations and inference, but provides no implementation details or worked example. The recommendation is specific to the described linear regression setup; more complex models may require methods suited to their own dependence structure, and the document does not explain how to construct predictive intervals beyond the standard-error adjustment.

Key ideas

  • Use all available overlapping periods in a single OLS regression for the point estimate.
  • Overlapping observations create autocorrelation that must be accounted for in inference.
  • Use standard errors robust to autocorrelation instead of averaging standard errors across start-date regressions.
  • The response does not provide a specific robust estimator or a full procedure for predictive intervals.

Tags

Full text
# Averaging Results Across Regressions due to Periodicity/Overlaps


# Averaging Results Across Regressions due to Periodicity/Overlaps












Given data that arrives at a daily frequency, I aggregated it to a weekly frequency, and estimated an OLS regression on it. Given that there are roughly 5 trading days per week, I can construct 5 different OLS models using 5 different starting points. For example - one model uses returns from Monday-Monday, the next Tuesday-Tuesday, and so on.

Assuming I believe there are no seasonal effects (e.g. models trained using Monday-Monday returns should be no different than Tuesday-Tuesday), is there a correct way to combine the predictions/coefficients of these 5 (or in the general case, N) models? I am inclined to think quick and simple averaging of coefficients would work. In that case, is there a proper way to combine the standard errors and residual standard errors across models? I ask because I am interested in constructing confidence/predictive intervals for forecasts. I hesitate to estimate the model using the full dataset, because this will cause overlaps in my endogenous variable, and I am not well equipped/don't know how to deal with that.

Of course this question could be asked more generally for any (non-linear) kind of model, but it seems like OLS/linear models would have the most hope for a theoretically sound procedure/heuristic.

## Answer by Richard Hardy (score 1)

https://quant.stackexchange.com/a/73659

The efficient point estimator would be OLS on all 5-day periods, even though there will be a lot of overlapping. You would need to adjust the standard errors for autocorrelation by using robust standard errors. No model averaging is needed. Here are some references:

- Hayashi, Fumio. Econometrics. Princeton University Press (2011). See sections 6.6-6.8.

- Britten‐Jones, Mark, Anthony Neuberger, and Ingmar Nolte. "Improved inference in regression with overlapping observations." Journal of Business Finance & Accounting 38.5‐6 (2011): 657-683.

- Harri, Ardian, and B. Wade Brorsen. "The overlapping data problem." Available at SSRN 76460 (1998).

- Hansen, Lars Peter, and Robert J. Hodrick. "Forward exchange rates as optimal predictors of future spot rates: An econometric analysis." The Journal of Political Economy (1980): 829-853.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.