Skip to content
All library documents

Selecting Signal Thresholds and Return Horizons Without Overfitting

Article Quant Q&A · Author: confused

Summary

The document considers how to choose both a signal threshold and the future return horizon when evaluating a trading hypothesis. It warns that searching many combinations can amount to data mining and produce overfit results, while acknowledging that there is no definitive solution. One suggested exploratory procedure is to label observations according to whether returns over a chosen horizon exceed a target, calculate strategy precision, and repeat across horizons and shifted labels.

The answer also points toward approaches intended to infer causality as a more principled direction, but does not explain or validate a specific causal method. The proposed repeated comparisons still involve many candidate horizons and can create selection bias, so they should not be mistaken for a safeguard against overfitting. The document offers framing and possible avenues for investigation rather than a complete testing protocol.

Key ideas

  • A trading hypothesis may require choosing both a signal threshold and a return horizon.
  • Searching many threshold and horizon combinations can lead to overfitting.
  • One exploratory approach compares precision across return horizons and shifted labels.
  • Causal inference methods are suggested as a more principled avenue, without a specific method being detailed.

Tags

Full text
# How do you build a model with uncertain time range?


# How do you build a model with uncertain time range?












Let's say you want to test the hypothesis that given a signal reaches some threshold, some asset will have some return over the next period.

Here we have two unknowns.

- One, the value of your threshold - that triggers the trade

- Two, the time period to measure your return after your signal occurs.

What approach do people usually take to solve the two unknowns? I've learned that data mining is bad (so I'm hesitant to test out random time periods and threshold values until I find a good combination), so what's a reasonable approach?

## Answer by scities (score 1, accepted)

https://quant.stackexchange.com/a/54460

This is a very tough problem and there are no definitive answers.

Of course you can try brute force, but you may very well overfit. One simple way to do this is to label each period in your time series with 1 if the return after X periods is above a desired threshold, 0 otherwise. Compute your strategy’s precision for a range of Xs. Then shift the labels by 1 (lag 1), do the same thing. Etc for many lags.

Or you can try to use more principled approaches, for instance methods that try to infer causality. This article is a good summary and has tons of interesting references.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.