Skip to content
All library documents

Why Large Samples Make Statistical Significance Less Practically Meaningful

Article Quant Q&A · Author: disha bansal

Summary

The document explains why regression studies with many observations can reject zero-effect hypotheses for very small relationships. As sample size rises, estimates become more precise and statistical tests gain power, so even a minor departure from the null can produce a low p-value. This issue applies beyond regression and cautions against treating statistical significance as proof that a finding matters in practice.

It recommends judging the size and relevance of an effect rather than making rejection harder by choosing a stricter significance threshold. A coin-weighing analogy illustrates how a sufficiently sensitive measurement can detect tiny differences that have little practical consequence. The document offers conceptual reasoning and cites statisticians' commentary, but no empirical trading example or quantitative procedure for assessing practical importance. Researchers still need to define meaningful effect sizes in the context of their question.

Key ideas

  • Large samples increase the power to detect small departures from a null hypothesis.
  • A statistically significant coefficient can have little practical importance.
  • The issue of significance in large samples extends beyond regression.
  • Assess effect size and relevance alongside p-values.
  • Changing a significance threshold does not establish whether an effect matters.

Tags

Full text
# regression analysis


# regression analysis












"A model estimated with a large no. of observations may allow one to reject null hypothesis of zero coefficients for many explanatory variables.Thus we might choose to select a somewhat lower significance to make rejection of a null hypothesis more difficult." Provide an intuitive explanation for the above statement.

## Answer by Alex C (score 2)

https://quant.stackexchange.com/a/20922

This is called the "p-value problem in large samples" and s not limited to regression. According to Cohen (1990): "a fact widely understood among statisticians: The null hypothesis, taken literally (and that’s the only way you can take it in formal hypothesis testing), is always false in the real world. If it is false, even to a tiny degree, it must be the case that a large enough sample will produce a significant result and lead to its rejection." This is a fundamental problem with p-values.

Another statistician Chatfield (1995) comments, “The question is not whether differences are ‘significant’ (they nearly always are in large samples), but whether they are interesting. Forget statistical significance, what is the practical significance of the results?” The increased power of large samples means that researchers can detect smaller, subtler, and more complex effects, but relying on p-values alone can lead to claims of support for hypotheses of little or no practical significance.

Looking at practical significance is a better solution than "select a somewhat lower significance to make rejection of a null hypothesis more difficult", although the effect is similar.

I found this article interesting : Mingfeng Lin et al.: "Too Big to Fail: Large Samples and the p-Value Problem".

An analogy might be the comparison of two coins (say 2 pennies): if you have an extremely sensitive balance, able to detect differences of a few atoms, you will find that any two coins you compare are always of different weight. A statistical estimator that uses a large sample is the equivalent of the very precise balance.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.