Skip to content
All library documents

Testing Whether an Automated Strategy Outperforms a Benchmark

Article Quant Q&A · Author: Studante

Summary

The discussion corrects a misunderstanding about statistical evidence for positive average returns. Rejecting a null hypothesis that mean daily return is nonpositive does not imply every day or investment period must have a positive return; it supports a claim about the average under the test’s assumptions. Individual losses remain possible.

To assess a strategy, the answer suggests comparing cumulative returns, expected returns, and volatility with a benchmark such as the S&P 500, and using simulations. The question describes a machine-learning strategy with limited backtest and test-period data, but the response does not specify a test procedure or address concerns such as dependence, model selection, transaction costs, or out-of-sample robustness. Statistical significance of an estimated mean should therefore not be read as a guarantee of future performance or of gains over a particular period.

Key ideas

  • A statistically positive average return does not rule out negative daily returns or losing periods.
  • A significance test addresses evidence about the mean under its assumptions.
  • Benchmark comparisons can include cumulative returns, expected returns, and volatility.
  • Simulation can help assess strategy behavior, but the response does not prescribe a specific method.
  • Limited historical and test data do not guarantee future performance.

Tags

Full text
# How to statistically prove that an automated trading strategy outperforms the market?


# How to statistically prove that an automated trading strategy outperforms the market?












For a university project, I have been working on a rather complex automated trading strategy, that levers machine learning techniques. I backtested the algorithm, as a result I have daily return data for various time windows (max. two years, it is a relatively new financial asset). I also simulated the performance of the algorithm on three months of test data.

Now I am writing the final report, but my daily supervisor wants more rigour i.e. statistical tests. E.g. he wants me to prove that the daily return on average is greater than zero or even the S&P500 daily return.

I believe this to be impossible. Let us assume all assumptions for whatever test are satisfied. Proving your average daily return to be higher than zero with an $\alpha = 5\%$, would imply it is statistically impossible to ever have a PERIOD (see EDIT below) with a negative return (which is definitely not the case for my strategy). Or thus, I believe my supervisor is asking me to prove that I have found the holy grail of automated trading strategies.

Two questions:

- Do you agree with my above reasoning: statistically proving an average daily return higher than zero is nonsense?

- Do you have any pointers or ideas about how I could somewhat prove or quantify that my algorithm works well, other than simulating over specific time periods and benchmarking my returns and Sharpe ratios against other assets and strategies.

EDIT: I meant period instead of day EDIT in response to @Olaf his comment: when I say statistically impossible I mean very improbable. Let us assume I would be able to reject ($\alpha=5\%$) the null hypothesis that the average daily return (e.g. calculated over 90 days) is smaller than zero. In case I would invest an equal amount every day, then given the above it is almost guaranteed that I will end up with more after 90 days. I tend to believe this is a very strong guarantee for an automated trading system.

## Answer by simmy (score 0)

https://quant.stackexchange.com/a/25846

To have an average that is statistically > 0 does NOT imply that you can NEVER have a negative daily return: it means that, on AVERAGE, you have a positive return! You can have a lot of negative days, but if the positive ones more than compensate for the losses, you end up with an AVERAGE positive return (that can or cannot be statistically significant from zero: your test will answer this question).

Check if your model outperforms S&P500 in terms of cumulative returns over the period of interest, for example. Compare the expected returns and also the volatility (this is not a negligible variable) of your model w.r.t. those of S&P. I also think that simulation methods are useful tools to test a strategy.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.