Skip to content
All library documents

Trade Counts and White’s Reality Check for Strategy Comparisons

Article Quant Q&A · Author: user3276418

Summary

The post questions how White’s Reality Check handles strategies that generate different numbers of trades. Its example compares one strategy with two observed returns and another with ten, then describes bootstrapping returns, taking the maximum strategy mean in each repetition, and using the resulting distribution to assess performance against a zero-return benchmark. The author worries that the strategy with fewer signals will have a more variable bootstrapped mean and may affect the p-values for both strategies.

As a possible adjustment, the post proposes comparing a t-statistic that scales the mean by the square root of the number of returns and divides by their standard deviation. It also raises a caveat: if returns are not normally distributed, that scaling may not be appropriate. The document poses the issue but provides no resolution, derivation, or empirical comparison. It therefore frames a question about test design and unequal sample sizes rather than establishing that the Reality Check is flawed or that the suggested statistic fixes the problem.

Key ideas

  • The post describes applying White’s Reality Check to multiple strategies by bootstrapping their returns and comparing the maximum mean with a zero benchmark.
  • It asks whether unequal trade counts make the bootstrapped means differ in variability.
  • The author suggests a t-statistic as a possible way to account for different sample sizes.
  • The post notes that square-root sample-size scaling may be questionable for non-normal returns.
  • No answer or empirical evidence is provided to resolve the concern.

Tags

Full text
# number of trades - flaw in White Reality Check?


# number of trades - flaw in White Reality Check?












I went through Whites paper of the reality check for multiple strategy testing. To summarize at a simple example:

I have 2 strategies, s1 and s2. s1 gives 2 signals and therefore 2 returns, s2 gives 10 signals and therefore 10 returns.

Here is the Reality Check procedure:

I compare the means of the returns to some benchmark., e.g., mean(returns) = zero. To get p-values of H0, strat returns are zero, I do a Bootstrap of both strat returns, calculate both means, take the max and repeat this prozess a high number of times. I call this array BSvals. From those values I get the p-value of H0(s1) and H0(s2) being the fraction of BSvals values, that is > mean return .

However, In this method I see a flaw. Because of the lower number of signals in s1 the variance of the bootstrapped mean of s1 will be much higher than in s2. Therefore s1 will push pvalues of H0 much higher then if s1 would have signaled more often. Therefore the mean of returns as a performance statistic might not be suitable when comparing strategies with different number of trades/signals. I think of using t statistic as a unbiased performance measure to compensate for different signal numbers, being t = mean(strat returns) * squareroot(number returns) / stddev

How is your thoughts about this? Strat returns being not normal distributed squareroot(number returns) might not be an appropriate scaling factor, though...

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.