Interpreting R-Squared for Predicting Asset Returns
Summary
The note explains why a low R-squared can still be meaningful in return prediction. In regression, R-squared measures improvement in accounting for the target’s variation relative to a baseline that always predicts its mean. A positive value may indicate smaller average prediction errors than that baseline, including in settings where the usual variance-explained interpretation does not apply. The answer cautions that this interpretation assumes the model has been checked for overfitting.
Statistical evidence that R-squared exceeds zero does not establish trading profitability. Practical value depends on the size of potential gains and costs such as fees, so predictive fit and economic usefulness are distinct questions. The note also rejects a universal threshold for a “high” score: difficulty varies by task, and the same value can be impressive in one problem but weak in another. It provides conceptual guidance, not a trading backtest or a general benchmark for acceptable performance.
Key ideas
- R-squared compares regression predictions with repeatedly predicting the target mean.
- A positive R-squared may indicate predictive improvement, subject to overfitting checks.
- Statistical significance does not show that a model can profit after trading costs.
- Whether an R-squared is high depends on the difficulty and context of the prediction task.
Tags
Full text
# R squared statistic in predictions of returns # R squared statistic in predictions of returns My question is related to an article which use predictive linear regression for the stock returns. There is told that R squared statistic of 1.6% is high. How can we measure which R squared is high? I am confused because I know that R squared statistic should be between 0 and 100 percents. Thanks a lot! ## Answer by Dave (score 2) https://quant.stackexchange.com/a/69297 The goal of regression is to account for the variance in $y$. If you are able to do that, then your predictions of the conditional mean of $y$ (conditioned on the values of your features) will be better than if you were to predict $\bar y$ every time. When $R^2$ is positive, even if just slightly, it means that you are accomplishing that goal (assuming mode validation to examine for overfitting). You have a better model than the quant team that always predicts $\bar y$. Even in the cases where $R^2$ lacks its interpretation as the proportion of variance explained, a positive value (on a model that has been shown, one way or another, not to suffer from bad overfitting) indicates that the amount by which your model misses are, in some sense, on average smaller than they would be if you always predicted $\bar y$. To determine if that is enough to profit will be a different story that will involve the money amounts at stake as well as issues external to the asset price prediction, such as trading fees. This gets into the practical significance of your model performance, which might be nil if you can’t make money, even if the model, as shown by something like an F-test in the OLS setting, has an $R^2$ that is statistically significantly greater than $0$. A danger that I see with $R^2$ is that it can get us thinking like grades in school, where we all want A-grade models with $R^2>0.9$. For difficult problems, that might be a ridiculous standard, and it might be the case that $R^2=0.016=1.6\%$ is pretty good (think of something like the Putnam competition where even one point out of the possible $120$ is pretty good). Conversely, for an easy task, $R^2=0.9$ might be rather pedestrian performance, despite looking like an A-grade in school.
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.