Testing Trading Strategy Performance Against Benchmarks
Summary
The document discusses whether a trading strategy’s returns differ meaningfully from a benchmark. It outlines comparing mean returns and variances, and introduces the Sharpe ratio test: under independent, identically distributed returns, its relationship to a t-statistic scales with the square root of the sample size. It points to research on Sharpe-ratio statistics for further treatment.
A separate example compares an HMM asset-allocation strategy with equal-weighted and buy-and-hold benchmarks using an independent-samples t-test. The reported p-values do not meet a 0.05 significance threshold, so the example finds insufficient evidence that the observed underperformance is statistically significant. The discussion also cautions that strategy comparisons depend on how the strategies differ and that return dependence or changing statistics across sampling frequencies can undermine simple tests. The example is introductory; it does not establish that the chosen test handles dependence, unequal variances, or other backtest-selection effects.
Key ideas
- A t-test can compare strategy and benchmark mean returns, but the choice of test and assumptions matter.
- The Sharpe ratio test is closely related to testing whether average excess returns differ from zero.
- The standard relationship between a Sharpe ratio and its t-statistic assumes independent, identically distributed returns.
- An HMM allocation example reports no statistically significant underperformance against two benchmarks at the stated threshold.
- Statistical comparisons should account for return dependence and the distinct construction of the strategies being compared.
Tags
Full text
# Test statistical significance of a trading strategy
# Test statistical significance of a trading strategy
I have created a trading strategy which operate every single day on the DAX 30, for the last 1700 trading sessions (some years). I have the daily returns of my strategy and also the daily returns of my index. I'm using R, therefore i can get sd, mean, var ecc...
The big issue that impresses the market watchers and financial types is the ability to consistently make above-market returns.
What are the main important instrument to test the significance of a trading strategy?Is it sufficient a simple t-test? Which type (paired, unpaires ecc..)? What are your suggestions?
## Answer by Artem Korol (score 3)
https://quant.stackexchange.com/a/29717
You can test if your mean return and variance are significantly different from benchmark statistics. Steps how to do it are described here "Chapter 9: Testing differences between two means, variances or proportions".
## Answer by lehalle (score 3)
https://quant.stackexchange.com/a/34207
"The Statistics of Sharpe Ratios" by Andy Lo is the standard reference for the widely used test for performances: the Sharpe ratio.
This ratio is close to the standard t-test: does the mean of my i.i.d. random variable of performances $R$ is significantly not zero (or better than the risk-free rate $r_f$)? $$\mbox{Sharpe Ratio}=\frac{\mbox{mean}(R) - r_f}{\mbox{std}(R)}.$$ To be compared to (where $N$ is the number of points used to compute the mean and the std) $$t\mbox{-test}=\sqrt{N}\cdot \frac{\mbox{mean}(R) - r_f}{\mbox{std}(R)}=\sqrt{N}\cdot\mbox{Sharpe Ratio}.$$
Of course this relies on the i.i.d. of the performance. It is not the case if you obtain different statistics on daily, weekly and monthly resampled timeseries of your dataset. These points are discussed in Andy Lo's paper.
If you want to go further on the nonparametric aspects of testing the Sharpe ratio, you can follow this link on the Cross-Validated stackexchange.
## Answer by ZKMathquant (score 0)
https://quant.stackexchange.com/a/85260
Before we go on with the 'how' part, let us understand what is it that we are even looking for,I would talk about an instance of an investing strategy report where instructors asked for this statistical significance part.Hope it would make sense to you why it is important:
[Cannot give you an exhaustive list of scenarios or in general,classify the problems on which one should look for stat significance,in finance. Here giving an example to have an ideea to start off]
- Suppose,we want to make two distinct trading strategy and backtest them.One based on markov chain properties and one based on Hidden Markov models. Now,I could go straight into the results/output and discuss why comparing their performance and whether they are statisitcally significant in terms of performance over other,and it would serve the purpose- however, without knowing the details of why they are 'inherently different strategies' perhaps you would overgeneralize this method everywhere.
- MC and HMM differ mainly on direct and latent(hidden) states. Direct states are easy to find,as they are observable and hence computationally efficient but latent states are hiddena and not directly observable and hence depend on heavy simulations. However for the purpose of capturing the right variables, HMM can be a better fit as direct states are not true representative of the variables relevant to the problem at hand.
- For eg,for trading only,price/volume is directly observable and a tool for Markov models,whereas bull/bear regimes or buy/sell investor sentiment- neither of them are directly observable yet truer representative of relevant characters in problems formulated accordingly.
- Hopefully, this gives an idea of inherent distinctness(which is never a rigorous mathematical/statistical signal to test for statistical significance,in any shape or form - but hope one can agree that it helps nevertheless ) of the strategies,although I have not mentioned the formulation of the strategies yet.
- Well,to be honest with you,depending on how classical this formulation is to those who are recently introduced to the concept in the context of finance,it is better not to mention the whole fomulation and rather leave it as an exercise. The gist of the main question is anyway about comparison of performance of strategies. Here is a gist howeveer:
HMM strategy permutes the available assets and change its weights depending on the detected regimes,and which one will be doing well and so on ,producing 3 states for asset allocation.
An example of the asset allocation(one just does not come up with it,you need to formulate your strategy and simulate it properly):
```
Allocation mapping per state:
State 0: {'TLT': 0.0, 'GLD': 0.4, 'SPY': 0.6}
State 1: {'TLT': 0.6, 'GLD': 0.4, 'SPY': 0.0}
State 2: {'TLT': 0.6, 'GLD': 0.4, 'SPY': 0.0}
```
Now let us directly dive into the results and performance of these strategies in my formulation,and how I dealt with the comparison part: In fact,I will go for performance of an HMM based asset-allocation strategy compared to trivial benchmarks like Buy-and-hold SPY, and Equal weighted allocation strategy (no permutation of assets and weights,rather assets equally distributed throughout ) .
Here are plots (that my simulations resulted with) to compare their performance:
- Now, there could be many claims behind this apparent under-performance of our strategy - it is poorly formulated, it leaves true bear market later than it should have and enters the true bull market late as well,hence making more loss and less profit than even the trivial counterparts (one could look at the plots of certain regime-specific window- the pre-pandemic bull-market and the '08 bear market)
- In fact, if one looks at state-conditional mean returns plot,and notice the typical nature of the assets Gold(crisis safe-heaven),SPY(good for bull market) etc the numerically obtained states 0,1 and 2 would make sense even economically,which is also one can try as an exercise.
- But now one should wonder is this apparent under-performance even statistically significant?
For that one could compare mean daily returns with that of benchmarks by operating a t-test.
For exmaple,
```
def compare_mean_returns(returns1, returns2):
returns1 = returns1.dropna()
returns2 = returns2.dropna()
t_stat, p_value = stats.ttest_ind(returns1, returns2)
return p_value
```
- And the following type of results can be interpreted as: ` vs benchmarks p_value_vs_EW p_value_vs_buy_and_hold_SPY HMM_Strategy 0.522436 0.370634 `
```
vs benchmarks p_value_vs_EW p_value_vs_buy_and_hold_SPY
HMM_Strategy 0.522436 0.370634
```
Given common significance level of α=0.05 , the above results imply the under-performance might just be off chance rather than statistical signifiance,i.e.,there is no sufficient evidence of stat significance.
Hopefully this works a beginner friendly introduction to relevance of statistical significance in investing. The question thus becomes where else one would find such requirement and how should one proceed with it in investing or trading,and what other kinds of tests might be useful in those cases.Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.