Detecting Backtest Data Snooping Through Selection-Induced Kurtosis
Summary
The document proposes testing whether a strong backtest was selected from many strategies by looking for excess kurtosis in its daily profit and loss observations. Under the simplified model, daily results are independent and normally distributed before selection; conditioning on total profit exceeding a threshold makes them dependent and changes their distribution. The proposed test would use observed returns alone, without access to the strategy or the search space that produced it.
The discussion explains that this is a conceptual signal rather than a validated procedure. Kurtosis estimates from a short sample are noisy, real trading returns may already have fat tails, and a strategy search could favor candidates with low measured kurtosis. It points to established alternatives, including White’s Reality Check, probability of backtest overfitting, and multiple-testing methods. It does not provide a formal derivation or empirical validation of the kurtosis proposal, and notes that the specific framing may not have been formally published.
Key ideas
- Selecting a strategy because its total backtest profit is unusually high can alter the distribution of its individual profit and loss observations.
- The proposed diagnostic is to look for excess kurtosis in the selected strategy’s returns.
- The idea assumes independent, normally distributed daily results before selection and a simplified selection rule.
- Short samples, naturally fat-tailed returns, and deliberate selection for low kurtosis make the diagnostic difficult to use reliably.
- Established data-snooping tests provide alternatives that address strategy selection more directly.
Tags
Full text
# Here is an approach for measuring Data Snooping; is it new?
# Here is an approach for measuring Data Snooping; is it new?
I came up with an approach for measuring data snooping, or overfitting. My question is whether this approach was published and expanded-on already, or is it new?
My approach relies on the observation that if we generate strategies over and over until one does well, then the distribution of PnLs of the postconditioned strategy has some extra kurtosis. This extra kurtosis arises from the sheer fact that we conditioned on the PnL being large. If we don't do any such conditioning (i.e. if our strategy was not created using snooping) then the kurtosis is as it should be, so there is no extra kurtosis. My approach has the advantage that it does not require any knowledge of the particular structure of the strategy space.
Here are the details of my approach:
Suppose our friend Alice gives us a supposedly-great strategy for investing in the stock market. Suppose we backtest this strategy (for simplicity I'll assume we have a perfect backtester) for 252 days, and the PnLs of this strategy are $P_1,\ldots,P_{252}$. We discover that the total PnL, $P = \sum_i P_i$, is very large, so this strategy looks great.
Now we want to figure out if this strategy is really great, so that it will continue being great in the future, or if this strategy is a result of data snooping (or overfitting), meaning Alice tried a bunch of mediocre strategies until she found one whose backtesting results are great.
We make a simplifying assumption, that for any particular strategy, its series of PnLs over the 252 trading days is normally distributed and independent. (So, in particular, for now we assume no kurtosis is present in the raw PnLs.) Now, suppose that Alice did not engage in data snooping. Then by this simplifying assumption, $P_1,\ldots,P_{252}$ should be independent and normally distributed. On the other hand, if Alice did indeed engage in data snooping, then $P_1,\ldots,P_{252}$ are i.i.d normally distributed (say with mean $0$), but then conditioned on the event that $\sum_i P_i \ge M$. (Here, $M$ is some large constant.) In this case, The random variables $P_1,\ldots,P_{252}$ are dependent, and perhaps we can test that. But, even better: they are not normally distributed anymore: they have some measurable kurtosis. This kurtosis is dependent on the value of $M$ and on the number of samples (in our case, $252$) but, presumably, it can be detected by just drawing a histogram of the $252$ samples.
Of course I made some simplifying assumptions that should be removed, such as the PnLs being normally distributed to start with, and our backtesting procedure being perfect. Furthermore, I assumed that Alice is snooping in a very particular prescribed way: that she tosses away a strategy if it doesn't make at least $M$ dollars, and that she gives us the first strategy she finds that makes $M$ dollars. If she tried to fool us by, say, measuring the kurtosis herself and waiting until she finds a strategy with low empirical kurtosis, then our strategy needs to be more sophisticated.
On the other hand, this approach has the advantage that it does not require any knowledge of the particular structure of the strategy space that Alice explored when she tried to find her strategy. So we don't need access to Alice's search space, only to the particular strategy she found. Better yet, we don't even need access to Alice's strategy itself: Alice can keep her strategy secret, and just truthfully report the results of backtesting her strategy.
Has this approach been explored in the literature? Also, does this approach work, or is one of the simplifying assumptions too simplistic, and is this approach bound to fail?
## Answer by nessos nessos (score 1)
https://quant.stackexchange.com/a/85547
The idea is conceptually sound, but there are already established methods that are more robust.
What already exists:
- White's Reality Check (2000) – the classic test for data snooping across multiple strategies
- Bailey et al. (2014) – "The Probability of Backtest Overfitting" (PBO) – a combinatorial approach via cross-validation that requires no distributional assumptions
- Harvey & Liu (2015) – a multiple testing framework specifically designed for financial strategies
On the kurtosis approach itself: The intuition is correct – conditioning on high total profit does introduce excess kurtosis in the individual P&L observations. However, there are practical limitations:
- Estimation noise – kurtosis estimated from 252 data points is very noisy (standard error ≈ √24/N ≈ 0.3). You need far more data to detect this reliably
- Real P&Ls already have fat tails – separating "selection-induced kurtosis" from natural market kurtosis is difficult in practice
- Gameable – a sophisticated actor could simply select strategies that happen to have low empirical kurtosis
What works in practice: Bailey's PBO framework is distribution-free and directly applicable to strategy selection problems. It measures what fraction of combinatorial train/test splits result in overfitting.
Reference: "The Probability of Backtest Overfitting" – Bailey, Borwein, Lopez de Prado, Zhu (2014)
To my knowledge, the specific kurtosis-based framing described here has not been formally published – it could be an interesting direction for further research.Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.