Reproducibility, Multiple Testing, and Overfitting in Quantitative Finance
Summary
This discussion examines why published quantitative findings and attractive backtests may fail to reproduce or translate into tradable returns. It explains how publication incentives, selective reporting, repeated testing, flexible data choices, and unrecorded experimentation can produce false positives. It recommends stating economic hypotheses in advance, tracking all tests, documenting data cleaning and outlier treatment, accounting for multiple comparisons, and validating results with genuinely independent evidence.
The article reviews competing studies on the replication of equity anomalies, reporting both skeptical findings and research that finds many factors reproducible. It emphasizes that statistical replication does not establish investability: trading costs, changing markets, limited independent market regimes, and repeated tuning can weaken results. Its examples and study summaries are presented as arguments for careful scrutiny, not as a settled verdict on every factor. Researchers are urged to interpret models, preserve a clear rationale, and treat historical sample splits cautiously when choices have been shaped by the full history.
Key ideas
- Repeated searches across variables, parameters, and samples raise the risk of false discoveries.
- Researchers should specify hypotheses and tests in advance and record unsuccessful as well as successful trials.
- Data selection, transformations, and outlier handling can materially affect reported results.
- Replication and statistical significance do not by themselves establish that a factor is profitable after trading costs.
- Repeatedly tuning a model against purportedly out-of-sample data undermines the independence of that evaluation.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.