How Return Frequency and Sample Size Affect Lagged-Return Significance
Summary
The document examines an apparently significant regression slope when current S&P 500 returns are regressed on the previous observation. The reported test uses daily closing prices over a long historical span and finds a positive lag-one coefficient with a small p-value, despite a near-zero R-squared. An answer explains that the result does not directly contradict a claim about annual returns: the sampling frequency differs, and daily returns can show positive serial correlation at short lags.
The large number of observations also gives the test substantial power to detect small associations, so statistical significance need not imply an economically important relationship. A second response notes that daily or monthly series can be serially correlated while aggregation over a year may obscure or offset short-term effects. The discussion is an informal interpretation, not a full re-estimation; it does not establish the coefficient's stability or address data choices and inference adjustments.
Key ideas
- The interpretation of lagged-return regression depends on the sampling frequency.
- Daily returns may show positive short-term serial correlation even when annual returns do not.
- A large sample can make a small relationship statistically significant.
- A very low R-squared indicates that the lagged return explains little variation in current returns.
- The discussion does not establish whether the reported relationship is stable or economically useful.
Tags
Full text
# Why do I have a statistically significant slope regressing R(t) on R(t-1)
# Why do I have a statistically significant slope regressing R(t) on R(t-1)
I am reading Cochrane's lecture note here
He mentioned that when you regress annual return on time t on that of time t-1, you will have neither statistically significant nor economically significant slope.
I performed a quick test with python as follows:
```
import statsmodels.formula.api as smf
import pandas as pd
import pandas.io.data as web
import datetime as dt
ts_spy = web.get_data_yahoo("^GSPC", start="1/1/1929")
ts_ret = ts_spy.Close.pct_change()
df_reg = pd.concat([ts_ret.shift(1), ts_ret], axis=1)
df_reg.columns =["prev", "cur"]
results = smf.ols("cur ~ prev", data=df_reg).fit()
print results.summary()
```
The result I got was not as claimed in the note.
```
OLS Regression Results
==============================================================================
Dep. Variable: cur R-squared: 0.001
Model: OLS Adj. R-squared: 0.001
Method: Least Squares F-statistic: 12.60
Date: Mon, 28 Apr 2014 Prob (F-statistic): 0.000387
Time: 08:50:08 Log-Likelihood: 52035.
No. Observations: 16180 AIC: -1.041e+05
Df Residuals: 16178 BIC: -1.041e+05
Df Model: 1
==============================================================================
coef std err t P>|t| [95.0% Conf. Int.]
------------------------------------------------------------------------------
Intercept 0.0003 7.64e-05 4.306 0.000 0.000 0.000
prev 0.0279 0.008 3.549 0.000 0.012 0.043
==============================================================================
Omnibus: 4891.255 Durbin-Watson: 1.998
Prob(Omnibus): 0.000 Jarque-Bera (JB): 298584.422
Skew: -0.614 Prob(JB): 0.00
Kurtosis: 24.009 Cond. No. 103.
==============================================================================
```
## Answer by user2763361 (score 6, accepted)
https://quant.stackexchange.com/a/11088
Why do you have 16180 observations? Is this daily data over 64 years or higher frequency data? I am guessing so by the magnitude of the intercept. At any rate, your test power would be huge with this large sample size, meaning small relationships will be statistically significant.
What Cochrane said is contingent on data frequency. At a high frequency it is untrue that you would find only statistical insignificance, as returns are very positively autocorrelated at these high sampling frequencies.
## Answer by user12348 (score 1)
https://quant.stackexchange.com/a/11094
You are right. Daily or even monthly financial series have serial correlation and lag 1 is generally the most correlated. Over the year, many competing forces may act on the market to randomize the returns.
Not sure the purpose of this exercise you are doing, but you can remove this auto-correlation if you want, using ARCH/GARCH(1,1).Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.