Testing a Trading Signal for Statistical Significance
Summary
The document asks how to test whether a multi-timeframe indicator signal produces alpha. The proposed entry rule requires the indicator to have a positive slope and be above zero across three timeframes for a buy, with the inverse conditions for a sell. The author suggests comparing trade returns at different holding periods with returns from a random-entry benchmark that has the same number of entries and avoids excessive clustering.
The answer proposes regression coefficient hypothesis testing: regress returns on dates and test whether the estimated slope differs significantly from zero using its standard error and a t-statistic. It mentions choosing a confidence level and degrees of freedom based on the number of observations. However, this answer does not directly construct the suggested matched random-entry benchmark or explain how to compare return distributions. A time trend in returns is also not equivalent to testing the conditional performance of a signal, and the response does not address dependence, transaction costs, or multiple testing. These omissions limit the method as a complete alpha evaluation.
Key ideas
- Compare signal outcomes with an appropriate benchmark, such as matched random entries.
- The proposed entry rule combines indicator sign and slope across multiple timeframes.
- A regression slope can be tested against zero using its standard error and a t-statistic.
- The answer does not establish that a time-trend regression tests the signal's predictive value.
Tags
Full text
# Running a simple alpha estimation test for statistical significance of a signal
# Running a simple alpha estimation test for statistical significance of a signal
I'm looking for some direction on testing whether a simple entry signal has statistical significance.
Let's say this is my simple entry signal:
Buy when some indicator has a positive slope and is above zero on 3 time frames (5M,15M,60M)
Sell when the indicator has a negative slope and is below zero on 3 time frames.
How do I go about statistically testing this signal to see if there's any alpha to be extracted?
I was thinking something along these lines:
- plot the P&L distribution for the signal should it exit 1,2,3, N bars after the entry.
- plot the P&L distribution if the entry point had been random and the exit had been again 1,2,3, N bars after the entry.
I'm not really sure where to go from there.
How do I go about creating a "random" entry signal? Perhaps taking the time difference between the first and last quote and creating a pseudo random time to enter. The number of entries would have to be equal to the number of non-random signals created from the indicator. Also, I say pseudo random because it probably wouldn't be desirable if all these "random" signals happened to cluster in one short time interval.
Then, what characteristics of these distributions should I be looking for to establish if the indicator signal has any statistical significance in terms of potential alpha that can be extracted.
Any direction or reference would be appreciated. Being that there are so many different types of statistical tests that can be done, it's hard to find anything particularly relevant.
## Answer by nsi (score 1, accepted)
https://quant.stackexchange.com/a/4254
I think the simplest way to achieve what you're looking for is through regression coefficient hypothesis testing.
- Perform linear regression on returns (y-axis) vs. dates (x-axis) over the desired time frames (do it once for 5 months, once for dataset w/15 months worth of data, and once for 60 months worth of data).
- As a result of regression, you will get coefficients for $y = mx + b$. To check if the slope is significantly different from 0.0, you would perform a t-test on m: $$ t = \frac{m - 0.0}{SE(m)} = \frac{m}{SE(m)},$$ where $SE(m)$ is standard error of $m$. Depends on how you're doing the regression, but it may be reported by the software.
- Using $n-2$ degrees of freedom (where $n$ is the number of points included in regression, look up $t$ values for the desired confidence level (typ., $\alpha=0.05$ for 95% confidence interval, CI). Let this value be $t_{critical}$.
- $m$ is significantly different from 0.0 if $t > t_{critical}$.
Slope's sign can be determined by adjusting the null hypothesis ($m > 0$, $m < 0$) & picking different $t_{critical}$ values.
...unless I misunderstood your question, that's one of the ways I see achieving what you're after.Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.