Skip to content
All library documents

Why Larger Samples Make Small Correlations Statistically Significant

Article Quant Q&A · Author: tweedi

Summary

The document explains why a small estimated correlation can become statistically significant as the sample size grows, even when the estimate itself stays unchanged. Significance testing addresses uncertainty in the estimate rather than whether the correlation is large enough to matter economically. With more observations, an estimate that remains near the same value is more precise, which can strengthen evidence against a zero-correlation null hypothesis.

The response presents Fisher’s z-transformation as a way to test a correlation and gives its approximate standard error, which declines as the sample grows. If the underlying correlation is truly zero, additional data would tend to move the estimate toward zero; if it is small but nonzero, more observations can make that departure easier to detect. The explanation assumes the approximation is suitable for the data. Statistical significance alone does not establish practical importance, predictive value, or stability across market regimes.

Key ideas

  • Statistical significance concerns the precision of a correlation estimate, not its economic size.
  • Fisher’s z-transformation provides an approximate test for a correlation coefficient.
  • The transformed estimate's standard error falls as the number of observations increases.
  • More data can distinguish a small nonzero correlation from zero more reliably.
  • A statistically significant correlation does not by itself imply trading usefulness or stability.

Tags

Full text
# Pearson correlation significance : Issue with $t$-statistic increasing with $N$


# Pearson correlation significance : Issue with $t$-statistic increasing with $N$












I have two assets which seem not correlated (correlation coefficient = 6.3% using monthly frequency and 48 data points).

I want to test the significance of the correlation. Null hypothesis is that correlation is nil, and if the p-value is lower than 0.05 then we can reject the null hypothesis (ie correlation is significantly different than 0, ie there is correlation).

Keeping the correlation constant (at 6.3%) I notice that as I increase $N$ (and so I increase $N-2$ degrees of freedom) the p-value reduces and will eventually be lower than 0.05.

I am confused, why having more $N$ make the correlation becoming statistically significant, if the actual correlation of returns is still low?

## Answer by kurtosis (score 1)

https://quant.stackexchange.com/a/57792

The question of significance is not about the correlation but about the precision of the estimation. If the value estimated with more data is still near the same value estimated with less data, that means you are more sure of that correlation being close to your estimate.

We can test the hypothesis that the correlation is 0 just as easily as testing that the correlation is 1 or -1. One way to do that is to use Fisher's $z$-transformation; for an estimated correlation $\hat\rho$, the test statistic $z$ is given by: $$ z = \frac{1}{2}\log\left(\frac{1+\hat\rho}{1-\hat\rho}\right) $$ and $z$ is approximately normal with standard error $\frac{1}{\sqrt{N-3}}$.

Unsurprisingly, you can note here that the standard error decreases as we have more data. If the data were just noisier, more data would give us a correlation estimate converging to 0. However, if the correlation were low but not zero, more data would reveal that.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.