Log Returns as Threshold Labels for Trading Classifiers
Summary
The document raises a question about constructing buy, sell, and hold labels for a machine-learning trading classifier. In the example it describes, hourly observations are assigned classes according to whether log returns cross positive or negative percentage thresholds, with values inside the thresholds treated as hold. The author asks whether using log returns to define these classes offers predictive information beyond using ordinary percentage returns, and why it might do so.
No answer or empirical comparison is included, so the document does not establish that log-return labels improve prediction or trading performance. It also provides no details about the referenced study's validation design, transaction costs, class balance, or out-of-sample results. The useful takeaway is the modeling question itself: label construction is a consequential design choice in supervised trading research, and any claimed advantage needs to be tested against alternative return definitions using leakage-aware validation and realistic trading assumptions.
Key ideas
- The document considers classifying hourly market moves into buy, sell, and hold outcomes using return thresholds.
- It asks whether log returns produce different or more useful labels than ordinary percentage returns.
- The text supplies no evidence that log-return labels improve predictive accuracy or trading results.
- Label definitions should be compared empirically within a sound validation design.
Tags
Full text
# What benefits do using log returns for model training provide? # What benefits do using log returns for model training provide? I came across a paper that uses Support Vector Machines to classify a `buy/sell/hold` decision each hour at the $\pm$0.5% threshold. The paper can bee seen here. The paper yielded impressive predictive power as well as high returns. It was noted that during the training phase they created the `[1, 0, -1]` labels by computing the percentage rate of change on log hourly returns. I have looked at answers provided here Why should we use log returns? Log normality, however this wasn't in the context of label creation for an ML classification problem. I was wondering if this technique offers predictive insight that using normal percentage returns? And if so, why does this work?
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.