Methods for Reducing Data Snooping in Strategy Backtests
Summary
The document distinguishes data snooping from ordinary in-sample versus out-of-sample model selection. Data snooping arises when researchers reuse the same observations across many strategy or hypothesis tests; searching enough candidates can produce an apparently successful result by chance, even when each candidate was individually validated or tuned with standard methods. It identifies multiple-testing adjustments such as Benjamini–Hochberg, White’s Reality Check, Hansen’s Superior Predictive Ability test, and Romano–Wolf procedures as ways to account for this selection process.
It also recommends practical safeguards: keep models simple, use training and test periods or cross-validation appropriately, examine performance across subperiods, and check parameter stability, execution costs, look-ahead bias, and survivorship bias. Random-trade comparisons and live trading are mentioned as additional ways to judge apparent significance. These approaches mitigate different sources of overfitting, but the discussion offers no single consensus procedure or universal guarantee of predictive power; test design and market assumptions still matter.
Key ideas
- Repeatedly testing strategies on one dataset raises the chance of finding success by luck.
- Data snooping is a multiple-testing problem distinct from ordinary train and test separation.
- Reality Check, SPA, and stepwise procedures adjust inference for searches across candidate strategies.
- Simple models, stable parameters, and subperiod checks can help reveal fragile backtest results.
- Backtests should account for execution costs, look-ahead bias, and survivorship bias.
Tags
Full text
# What are the popular methodologies to minimize data snooping?
# What are the popular methodologies to minimize data snooping?
Are there common procedures prior or posterior backtesting to ensure that a quantitative trading strategy has real predictive power and is not just one of the thing that has worked in the past by pure luck? Surely if we search long enough for working strategies we will end up finding one. Even in a walk forward approach that doesn't tell us anything about the strategy in itself.
Some people talk about white's reality check but there are no consensus in that matter.
## Answer by gappy (score 23, accepted)
https://quant.stackexchange.com/a/191
Strictly speaking, data snooping is not the same as in-sample vs out-of-sample model selection and testing, but has to deal with sequential or multiple tests of hypothesis based on the same data set. To quote Halbert White:
> Data snooping occurs when a given set of data is used more than once for purposes of inference or model selection. When such data reuse occurs, there is always the possibility that any satisfactory results obtained may simply be due to chance rather than to any merit inherent in the methody yielding the results.
Let me provide an example. Suppose that you have a time series of returns for a single asset, and that you have a large number of candidate model families. You fit each of these models, on a test data set, and then check the performance of the model prediction on a hold-out sample. If the number of models is high enough, there is a non-negligible probability that the predictions provided by one model will be considered good. This has nothing to do with bias-variance trade-offs. In fact, each model may have been fitted using cross-validation on the training set, or other in-sample criteria like AIC, BIC, Mallows etc. For examples of a typical protocol and criteria, check Ch.7 of Hastie-Friedman-Tibshirani's "The Elements of Statistical Learning". Rather the problem is that implicitly multiple tests of hypothesis are being run at the same time. Intuitively, the criterion to evaluate multiple models should be more stringent, and a naive approach would be to apply a Bonferroni correction. It turns out that this criterion is too stringent. That's where Benjamini-Hochberg, White, and Romano-Wolf kick in. They provide efficient criteria for model selection. The papers are too involved to describe here, but to get a sense of the problem, I recommend Benjamini-Hochberg first, which is both easier to read and truly seminal.
## Answer by Shane (score 19)
https://quant.stackexchange.com/a/142
Building an effective backtest is not significantly different than building any other kind of predictive model. The goal is to have similar behavior out of sample as you have in sample. As such, there are methodologies developed in statistics and machine learning that can be useful:
- Understand the bias/variance tradeoff. This is covered in many places. For a technical discussion, see lecture 9 of Andrew Ng's machine learning class at Stanford.
- You can certainly use a training and test dataset. But there are also other kinds of approaches that can be used. To list two common options: cross-validation (similar to having segmented data, but can help with parameter selection) and ensemble methods (using multiple models can outperform just one and further reduce the curve-fitting problem).
So a few general recommendations:
- Your guiding principle should be Einstein's razor: 'Everything should be kept as simple as possible, but no simpler.' In other words, less degrees of freedom in your model equates to less chance for overfitting. In the statistics world, this can involve eliminating unnecessary parameters through a selection or regularization method.
- Robustness (in every respect) is also critical. Parameters that result in sharp changes in expected prediction error will be more open to the risk of overfitting. Similarly, if the model has no fundamental basis, then it should be applicable to a wide number of assets.
- Lastly, this applies to any kind of model: understand your data, your model, your objectives, assumptions, etc. There have been countless mistakes made over time from people not understanding the meaning of their models, implications, and risks. This includes things like execution assumptions and transaction costs. Make sure that you take everything into account. Lead by being skeptical of your data, constantly asking what can go wrong, or how can the future be different. Is there any survivorship bias in your data, and if so, how can you control for it? Have you introduced any look-ahead bias?
## Answer by shabbychef (score 10)
https://quant.stackexchange.com/a/167
I have seen Hansen's SPA ('Superior Predictive Ability') test and stepwise variants used for this purpose. Hansen's test is a Studentized version of White's Reality Check. The stepwise variants allow one to accept or reject the null of no predictive ability on a subset of some tested strategies while maintaining a familywise error rate.
In his book, 'Evidence-Based Technical Analysis,' David Aronson discusses the overfit bias very well, although I believe his techniques for minimizing the bias may only apply to technical strategies, because they rely on Monte Carlo simulations.
References
- P. R. Hansen, 'A Test for Superior Predictive Ability,' Journal of Business & Economic Statistics, vol 23, no 4, 2005, http://pubs.amstat.org/doi/abs/10.1198/07350010 5000000063.
- SPA google group
- Hsu, Po-Hsuan, Hsu, Yu-Chin and Kuan, Chung-Ming, 'Testing the Predictive Ability of Technical Analysis Using a New Stepwise Test Without Data Snooping Bias,' 2008, http://ssrn.com/abstract=1087044
- Hsu, Po-Hsuan and Hsu, Yu-Chin, 'A Stepwise SPA Test for Data Snooping and its Application on Fund Performance Evaluation,' 2006, http://ssrn.com/abstract=885364
- David Aronson's Evidence-Based TA.
## Answer by Richard Herron (score 7)
https://quant.stackexchange.com/a/144
The output of your model will be a realization of your assumptions. Shane's given you a great answer. Besides doing out of sample testing (i.e., calibrating on period X then testing in period Y only using info available at the time of each trade), I would add that you should test it in sub-periods. If you have a big chunk of data, break it up and see how it works on each subset of the data.
## Answer by Patrick Burns (score 7)
https://quant.stackexchange.com/a/269
This blog post points to a presentation about backtesting and data snooping: http://www.portfolioprobe.com/2010/11/05/backtesting-almost-wordless/
I think the only non-datasnooping method there is is to trade live. But the problem of data snooping can be reduced by seeing how significant the backtest result is compared to what would have happened if the trades were random. Using this technology also makes it clear that backtesting results can easily be deceiving.
## Answer by Zarbouzou (score -2)
https://quant.stackexchange.com/a/146
Thanks for the answer as it tackles a lot of backtesting flaws, model parsimony, overfitting, survivorship bias, look ahead... But actually one can look at thousands of technical trading rules and other more sofisticated strategies, and maybe find the few ones that will answer all these problems. Nevetheless we would still be left with data snooping ie we have used our data set untill we find a satisfactory result.Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.