Skip to content
All library documents

Why R-Squared Can Be Negative on Random Binary Data

Article Quant Q&A · Author: Student

Summary

The document poses a question about a negative R-squared value obtained by comparing two random sequences whose values are limited to zero or one. The questioner expects the pairwise prediction error to match that of a constant prediction at the sample mean, reasoning that about half of binary observations should agree and half differ. The reported result challenges that intuition.

No answer or explanation is included, so the document does not resolve why the score is negative. It does identify the relevant comparison: R-squared evaluates a model against a baseline based on the observed target’s variation, and its value can fall below zero when the model performs worse than that baseline. The example supplies a random seed and sample setup, but no computed output or analysis of finite-sample variation; it is therefore a useful question prompt rather than a complete statistical explanation.

Key ideas

  • R-squared can be negative when predictions perform worse than the constant-mean baseline.
  • The document compares random binary sequences with a constant prediction at the sample mean.
  • A rough expectation of matching and differing binary values does not itself establish the resulting R-squared.
  • The document contains no answer or analysis resolving the example.

Tags

Full text
# Why is R^2 negative on random data?


# Why is R^2 negative on random data?












The introductory book says R^2 is between 0 and 1, but I have two randomly generated sequences and the R^2 is negative. So I read further and now understand the negative value is because R^2 essentially measures how well your model beats a horizontal line. OK, that makes sense. So then I made a random sequence that is always either 0 or 1. Obviously the best fitting line is a horizontal line at y=.5 and that prediction will have an error of .5 on every single point. The R^2 of that prediction is 0, so that makes sense. But then I made two random sequences that are always either 0 or 1, and the R^2 is still negative. But since each value is either 0 or 1, that means half the points will be identical (error=0) and half the points will be different (error=1), so there should be an average error of .5 exactly like the horizontal line. So why doesn't this random sequence have exactly the same explanatory power as the horizontal line? Why doesn't it also have the same R^2=0 as the horizontal line?

```
np.random.seed(10)
size=100
x = np.arange(size).reshape(-1, 1)
random_1 = np.random.randint(0,2,size=size)
random_2 = np.random.randint(0,2,size=size)

r2_vs_random = sklearn.metrics.r2_score(random_1, random_2)
print(f"R^2 vs random={r2_vs_random}")
plt.scatter(x,random_1);
plt.scatter(x, random_2);
```

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.