Estimating the Sample Length Needed to Test a Sharpe Ratio
Summary
The document derives a rough out-of-sample observation length for testing whether a strategy’s mean return exceeds zero. Assuming independent, identically distributed returns, it models the sample mean’s uncertainty as the return standard deviation divided by the square root of the sample size. Combining this with the Sharpe ratio gives a minimum sample-size rule that grows as the squared inverse of the Sharpe and depends on the chosen significance threshold.
The post notes that the calculation is a simple approximation and raises use of a t distribution as a possible refinement. An answer relates the rule to statistical power, which also depends on the type I and type II error rates, and suggests squared Sharpe as a figure of merit. The derivation assumes stable return moments and independent observations; it does not account for serial dependence, non-normality, multiple testing, or the distinction between significance and power in a complete testing design.
Key ideas
- Under iid returns, the standard error of the sample mean falls with the square root of sample size.
- The rough required sample length scales inversely with the squared Sharpe ratio.
- The significance threshold affects how many observations are needed to reject a zero-mean null.
- A fuller power calculation must also account for the probability of a false negative.
- Serial correlation and other departures from the iid assumption can make the approximation unreliable.
Tags
Full text
# Significance testing of average returns from Sharpe ratio
# Significance testing of average returns from Sharpe ratio
I'm aware that one way to do significance testing on a strategy is based on the sampling distribution of its Sharpe (see, e.g., Lo, 2002 and Opdyke, 2008).
However, it appears to me that there's another very simple way to do inference directly on the average return and standard deviation implied by the Sharpe. I'm wondering (1) if this approach is valid, or whether I'm overlooking something; (2) if there's some literature on this and similar approaches.
Given a strategy with (required / targeted / IS-based) daily Sharpe ratio $SR = \frac{\mu}{\sigma}$, with returns statistics $\mu$ (mean) and $\sigma$ (standard deviation).
The question is how long of an OS period the strategy needs at minimum to establish statistically significant returns.
Assuming iid returns (big assumption), the sample mean of the daily returns in the OS period after $n$ days can be modelled with (probably more appropriate to use t distribution though) $$\hat\mu \sim \mathcal{N}\left(\mu, \frac{\sigma^2}{n}\right)$$.
We define statistical significance at level $\alpha$ as the $z_{1-\alpha}$ score (one-sided test) from zero (i.e., the null is that $\mu=0$): $$\hat\mu \geq z_{1-\alpha} \frac{\sigma}{\sqrt{n}}$$.
With the definition of the daily Sharpe $SR = \frac{\mu}{\sigma}$ and substituting $\mu=\hat\mu$ this becomes:
$$n \geq \left(\frac{z_{1-\alpha}}{SR}\right)^2$$
Intuitively this makes sense: A lower Sharpe, or higher significance level increases the minimum sampling period. For instance, at $\alpha=90\%$ significance we obtain for a strategy with an annual Sharpe of 1: $n \geq (1.28 * \sqrt{252} / 1)^2 \approx 413$ days, but at $\alpha=95\%$ already $n \geq 682$ days.
## Answer by steveo'america (score -1)
https://quant.stackexchange.com/a/37823
This looks like a power rule, as discussed in Section 3.5.6 of the Short Sharpe Course. The approximate rule is that one requires $n \ge \frac{c}{SR^2}$ for some value of $c$ that depends on the type I and type II rates (equivalently, false positives and power). There is some argument, then that Sharpe ratio squared ought to be the figure of merit quoted for strategies (the units make more sense).Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.