Why AR(1) R-Squared Differs for Prices and Returns
Summary
The document asks why fitting an autoregressive model of order one to stock prices can yield a high test-set R-squared while fitting the same form to returns can yield a negative score. An AR(1) regression predicts the current observation from its lag, but prices and returns have different statistical behavior. A price series can be highly persistent, so a lagged price may track much of its level variation. Returns are typically much less persistent, and their fluctuations are not explained simply by the previous return; a forecast can perform worse than predicting the test-set mean, which produces a negative R-squared.
The example reports different scores from a train/test experiment on SPY closing prices and percentage returns. However, the displayed code appears to call `lm.predict` in both evaluations instead of the separately fitted `lm_price` and `lm_returns` models, so the reported return score may not correspond to the intended return regression. The snippet also does not provide statistical diagnostics or establish a general forecasting result. A fair comparison requires checking the model variables, aligned observations, and evaluation setup.
Key ideas
- Price levels can be highly persistent, making a lagged price a strong predictor of their level variation.
- Returns often have much less serial persistence than prices.
- A negative test R-squared means the model scored worse than a baseline based on the test-set mean.
- The example’s prediction calls appear to use a different model variable than the models it fits.
- Different R-squared values do not by themselves show that one representation is a better trading forecast.
Tags
Full text
# How can the different r2 score of an AR(1) model on prices vs. returns be explained
# How can the different r2 score of an AR(1) model on prices vs. returns be explained
This is maybe a silly question, but I want to understand. As far as I understand an AR(1) model, it is basically a linear regression model with the same but lagged variable, right?
However I am wondering how the different results can be explained when we apply an OLS on the lagged price time series vs. on the return time series.
See this experiment in python:
```
from sklearn.linear_model import LinearRegression
from sklearn.metrics import r2_score
import yfinance
spydf = yfinance.download("SPY")
train_lagged_price = spydf[["Close"]][:-200].shift(1).dropna().values
train_price = spydf["Close"][:-201].values
test_lagged_price = spydf[["Close"]][-201:].shift(1).dropna().values
test_price = spydf["Close"][-200:].values
lm_price = LinearRegression().fit(train_lagged_price, train_price)
r2_price = r2_score(test_price, lm.predict(test_lagged_price))
train_lagged_returns = spydf[["Close"]].pct_change()[1:-200].shift(1).dropna().values
train_returns = spydf["Close"].pct_change()[1:-201].values
test_lagged_returns = spydf[["Close"]].pct_change().dropna()[-201:].shift(1).dropna().values
test_returns = spydf["Close"].pct_change()[-200:].values
lm_returns = LinearRegression().fit(train_lagged_returns, train_returns)
r2_returns = r2_score(test_returns, lm.predict(test_lagged_returns))
r2_price, r2_returns == 0.9333725796827039, -1.6887830366573184
```
We get 2 totally different r2 scores and I would like to explain this difference.Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.