Skip to content
All library documents

Multiple Testing and False Discoveries in Trading Backtests

Article Quant Q&A · Author: David LE

Summary

In quantitative research, “data mining” refers to trying many hypotheses, strategies, or parameter choices and reporting only the one that appears statistically significant. Even when every tested idea has no genuine predictive value, repeated testing makes a chance success likely. A backtest can therefore look convincing because of the search process rather than durable trading skill.

The answer connects this problem to false discoveries and explains why it is especially relevant when many tests are run on limited financial data. A result selected this way may fail to replicate on a separate sample. It points to multiple-comparison controls, including family-wise error methods and Holm–Bonferroni adjustments, as ways to address the issue. It offers no worked example or specific correction procedure, so it serves as a conceptual warning rather than a complete guide to validating a strategy.

Key ideas

  • Trying many tests and publishing only a significant result raises the chance of a false discovery.
  • Chance findings can arise even when the tested hypotheses have no real effect.
  • Financial research is vulnerable because many tests may be run on limited data.
  • Failure to replicate on a separate sample can reveal selection-driven results.
  • Multiple-comparison procedures can help control false discoveries.

Tags

Full text
# What does it mean when an optimistic return forecast was due to "data mining"?


# What does it mean when an optimistic return forecast was due to "data mining"?












> The success of backtesting supports the prediction that an adaptive trading strategy fares better than using fixed rules. It also suggests that the positive results in the literature are not due to data mining

I come from a data science/engineering background, so the data mine I understand are unsupervised models building things like clustering.

But I don't think that is what it means in quantitative finance. I've read lots of paper and funds shunning the word "data-mining", as if it is some illusion a trader sees in their backtesting portfolio and not attributable to "real skill".

Explain what it means and how does one typically commit this problem.

## Answer by nbbo2 (score 4)

https://quant.stackexchange.com/a/79339

"Data mining" is also used in another negative or derogatory sense: it means to perform a large number of statistical tests hoping to find one that is statistically significant at 5% level and publishing only that one. This leads to "false discoveries". Recall that 1 out of 20 statistical test can be expected to show as positive even thogh the null hypothesis is valid. In fields like Finance or Genetics, where it very easy to run an enormous number of tests on a limited amount of data, this is a major problem, and it shows up as inability to replicate the result on another sample.

There are special techniques in Statistics for dealing with this issue, google False Discovery, FWER, Holm-Bonferroni, Multiple Comparisons Problem, and similar terms. In Finance Lopez De Prado and Cam Harvey among others have written about this issue. Anyone who performs "backtests" needs to be aware of it.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.