Choosing Between Directional Labels and Return Forecasts for Trading Signals
Summary
The document raises a model-design question for trading signals: should a strategy predict return direction, such as a positive or negative label, or forecast the magnitude of the return? Directional classification can be modeled with a binary response, while return prediction retains magnitude information but asks the model to estimate noisy numerical outcomes. The author asks how either output should be used in a backtesting engine and how to interpret a near-zero coefficient of determination alongside apparently useful backtest performance.
The source presents the question rather than an answer, method comparison, or empirical evidence. It highlights that statistical fit metrics and trading utility need not align, but gives no details about the model, validation design, trading costs, position sizing, or robustness. No general recommendation can be drawn from the material alone.
Key ideas
- Directional labels discard information about the size of returns.
- Forecasting exact returns preserves magnitude but is a more demanding prediction task.
- A low coefficient of determination does not by itself establish whether a signal is useful in trading.
- The document poses the modeling tradeoff but offers no comparative evidence or recommendation.
Tags
Full text
# Trading signal strength: [-1 to 1] or [predicted return]? # Trading signal strength: [-1 to 1] or [predicted return]? In the context of a backtesting engine, is it better to have strategy generate trade signals in the range from -1 to 1 or as exact predicted returns (e.g. -12% or 26%). The difference lies in how to regress the underlying model: whether to put (the response variable) returns as 1 (for positive) or -1 (for negative) (and use a logit model) or put exact historical return values (and use a linear model). The reason for asking this is that I found a model that gives me an R^2 of almost 0 (which would seem rubbish in terms of predicting returns), but when backtesting it performs well and is actually a good proxy for relative signal strength (although the returns it gives are 0.17% or -0.09% etc). It seems that by assigning 1 for positive returns I am losing some of the information; on the other hand trying to predict exact returns seems like a tall order -- I do not like that R^2 might not correspond to backtest results at all. Which approach is better? Is there some standard literature on this (be it backtesting component design or alpha strategy design)?
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.