Testing Return Predictability with Autocorrelated Data
Summary
The document considers whether one asset’s return can predict another’s at a future horizon, using VIX and S&P returns as an example. It explains why overlapping daily observations can produce serial dependence and make ordinary regression t-tests unreliable. Suggested approaches include Newey–West standard errors, regression models that account for autocorrelated errors, and block bootstrapping, which can preserve some of the dependence that simple resampling loses.
The answers also caution that a univariate relationship may reflect omitted common drivers, and suggest adding market or established risk factors. Other proposals include vector autoregressions, shorter forecast horizons, and rolling out-of-sample evaluation to examine coefficient stability and forecasting performance. These are methodological suggestions rather than evidence that VIX predicts equity returns: the document reports no empirical test results. Forecasting a return far ahead may be especially difficult, and the appropriate model and inference depend on the data’s dependence structure and the research design.
Key ideas
- Overlapping return pairs can be serially dependent, making ordinary regression inference unreliable.
- Newey–West standard errors can adjust inference for autocorrelation and heteroskedasticity.
- Block bootstrapping can preserve dependence better than resampling individual observations.
- Adding common market and risk factors can help address omitted-variable bias.
- Rolling out-of-sample evaluation can assess forecast performance and coefficient stability.
Tags
Full text
# Using linear regression on (lagged) returns of one stock to predict returns of another
# Using linear regression on (lagged) returns of one stock to predict returns of another
Suppose I want to build a linear regression to see if returns of one stock can predict returns of another. For example, let's say I want to see if the VIX return on day X is predictive of the S&P return on day (X + 30). How would I go about this?
The naive way would be to form pairs (VIX return on day 1, S&P return on day 31), (VIX return on day 2, S&P return on day 32), ..., (VIX return on day N, S&P return on day N + 30), and then run a standard linear regression. A t-test on the coefficients would then tell if the model has any real predictive power. But this seems wrong to me, since my points are autocorrelated, and I think the p-value from my t-test would underestimate the true p-value. (Though IIRC, the t-test would be asymptotically unbiased? Not sure.)
So what should I do? Some random thoughts I have are:
- Take a bunch of bootstrap samples on my pairs of points, and use these to estimate the empirical distribution of my coefficients and p-values. (What kind of bootstrap do I run? And should I be running the bootstrap on the coefficient of the model, or on the p-value?)
- Instead of taking data from consecutive days, only take data from every K days. For example, use (VIX return on day 1, S&P return on day 31), (VIX return on day 11, S&P return on day 41), etc. (It seems like this would make the dataset way too small, though.)
Are any of these thoughts valid? What are other suggestions?
## Answer by Richard Herron (score 11)
https://quant.stackexchange.com/a/1130
A few thoughts.
Yes, your return series are autocorrelated (i.e., stocks don't exactly follow a random walk), so you should use Newey-West standard errors.
If you do this as a univariate regression $$R_{i,t} = \alpha_i + \beta_i R_{j,t-1} + \epsilon_{i,t}$$ then there's almost certainly an omitted variable inside $\epsilon$ that is moving both $R_i$ and $R_j$. So make sure you throw in at least some of the normal "predictors", like market return, default premium, term premium, or the other standard factors (Fama and French's SMB, HML, and UMD). In the hedge fund research you also see a set of seven or so factors.
## Answer by Vishal Belsare (score 6)
https://quant.stackexchange.com/a/1300
Have you considered fitting ARIMA with exogenous regressors model? Linear regression with autocorrelated errors might be appropriate.
R can do this with the arima() function via specifying the xreg argument.
## Answer by shabbychef (score 4)
https://quant.stackexchange.com/a/1303
Van Belle describes a basic correction for autocorrelation in a t-test, although it may be hard to wedge it into the regression t-test. For the 1-sample t-test of the mean, the correction is to multiply the t-statistic by $\sqrt{\frac{1 - \rho}{1 + \rho}}$, where $\rho$ is the 1-period autocorrelation (or estimate thereof).
## Answer by Frank_M (score 3)
https://quant.stackexchange.com/a/1305
Well vix is a measure of volatitity which would make it an estimate of a second moment for S&P 500 so you might try an arch/garch in the mean type model on S&P.
A good starting place for a project like this is to just do Vector Autoregressions on industry groups that you think might be related and see what comes up.
N+30 is a long way in the future, especially to estimate a return. I would use weekly data and try and estimate a mean.
## Answer by Felix (score 1)
https://quant.stackexchange.com/a/14464
This is a vector autoregression (VAR) model with restrictions. I would also use a shorter lag and add conventional factors to the VIX return.
Simple bootstrapping destroys the serial correlation in your model, but you might consider block bootstrapping.
I would also try the following test: use a subset to calibrate your model and then perform an out-of-sample step with the next data point of the regressor. Repeat this procedure by shifting the subset by one step (i.e. apply the regression to a rolling time window). This gives you an idea of the stability of the coefficients, and the forecasting ability of the model.Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.