Adjusting Strategy Significance for Threshold Selection
Summary
The document considers how to evaluate a trading rule that enters an asset only when a metric exceeds a threshold. It compares the selected observations’ returns with the asset’s overall returns and illustrates the idea with correlated simulated data and a one-sided two-sample t-test. The example produces a higher subset mean, while its reported p-value leaves the strength of evidence open to judgment. The main question is how to account for strategy complexity, including the number of thresholds or metrics examined, without treating heavily overlapping threshold samples as wholly separate tests. The author also questions whether Sharpe ratios can be translated into significance measures when a strategy is invested only part of the time, and considers information criteria as a possible analogy for penalizing complexity. The document offers no solution or formal adjustment method. Its simple test setup also does not resolve dependence, selection bias, or the validity of comparing a subset with a sample that contains it.
Key ideas
- A threshold rule can be assessed by comparing returns during selected periods with overall asset returns.
- The example uses a one-sided two-sample t-test on correlated simulated observations.
- Testing many overlapping thresholds creates a multiple-comparison problem that simple test counts may not capture.
- Strategy significance adjustments should account for the number of metrics considered and the available sample size.
- The document poses these issues but does not establish a preferred correction.
Tags
Full text
# Adjusting the p-value of a strategy for number of parameters
# Adjusting the p-value of a strategy for number of parameters
Let's say I have some metric and I'm trying to evaluate whether it's predictive with respect to returns. I plan to only take trades where the value of the metric is above a certain threshold, such that I'm in the asset only part of the time.
I then want to compare the returns of the strategy to the overall returns of the asset. Below is an example where I'm taking trades if the metric `x` is above a threshold value. The red shaded region represents trades actually taken and the `y` values represent returns (scale is ignored for simplicity).
```
import numpy as np
import matplotlib.pyplot as plt
from scipy.stats import ttest_ind
np.random.seed(0)
cov = np.array([[1,0.4],
[0.4,1]])
data = np.random.multivariate_normal((0,0), cov, size=100)
x = data[:,0]
y = data[:,1]
threshold = 1.0
upper_x = x[x > threshold]
upper_y = y[x > threshold]
plt.scatter(x,y)
plt.scatter(upper_x, upper_y, color="red")
plt.axvline(threshold, color="red")
x_lim = plt.gca().get_xlim()
margins = plt.gca().margins()
plt.axvspan(threshold, x_lim[1]+margins[0], color="red", alpha=0.2)
plt.xlim(x_lim)
plt.show()
population_mean = y.mean()
subset_mean = upper_y.mean()
ttest = ttest_ind(upper_y, y, alternative="greater")
statistic = ttest.statistic
pval = ttest.pvalue
print("overall mean:", population_mean)
print("subset mean:", subset_mean)
print("p-value: ", pval)
print("statistic:", statistic)
```
```
overall mean: 0.07900435132575541
subset mean: 0.4628558666591004
p-value: 0.08187310219812942
implied z-score: 1.3925820821961334
```
In the example, the mean of the subset is higher than the overall mean, but with a p-value of 0.08, I'd have to decide whether it's worth trading the strategy.
As an aside, I've typically used Sharpe ratio and other custom metrics to evaluate strategies in backtesting, but I'm trying to use p-value more. Technically, Sharpe is basically just a Z-score, so I could convert it to a p-value, but a given strategy will not always have a position and Sharpe might be deflated because of that. I could use something like in-sample Sharpe, but I think p-value is a little more intuitive and straight to the point to tell me how likely it is I'm looking at something that represents an actual edge. Ultimately, it doesn't matter which is used as long as I can rationally adjust the value based on the number of metrics I'm using.
Regarding strategy evaluation, a couple things I'm already aware of:
- Multiple hypothesis testing: As I understand it, the idea is to multiply the p-value by the number of hypotheses tested, however, I don't think that viewing this as a multiple hypothesis test is appropriate because there will always be a degree of overlap between thresholds in that the most extreme value(s) will always be included. If one threshold leads to 20 trades and another threshold leads to 21 trades, 20 of which are from the first group, it doesn't make sense to call those different hypotheses.
- Akaike Information Criterion and Bayesian Information Criterion: If I understand these correctly, they are more for comparing models rather than looking at absolute probabilities, however, the idea of adjusting a baseline probability/likelihood for model complexity is what I'm looking for.
The example above is for a single metric/parameter, but I'm wondering if there are any other methods I can use to adjust a p-value (or Z-score) for the complexity of the metrics I use to trade on that would work for a varying number of metrics, sample size, etc.
Any suggestions are appreciated.Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.