Designing Machine-Learning Labels for Stock Prediction
Summary
The document raises the problem of defining target variables for machine-learning models that predict stock behavior. Candidate labels include the next day’s net return and a return based on the maximum high over a future window. It also references a financial machine-learning text that discusses labeling methods, while noting the questioner’s concern that its justification and evaluation may be limited.
The replies offer starting points for further study rather than a specific labeling recipe: one recommends research by Gu, Kelly, and Xiu, while another directs readers to the labeling chapter’s bibliography for papers on different techniques. No empirical comparison is provided among return-based targets, future high-based targets, or classification labels. The document therefore frames label choice as a research design question, but leaves practical decisions—such as the forecast horizon, treatment of costs, evaluation criteria, and fit between labels and trading objectives—open for investigation.
Key ideas
- A prediction target can use a future return or a statistic such as the maximum high over a horizon.
- Label design is a central part of financial machine-learning research.
- The replies recommend consulting published research and bibliography sources.
- The document provides no evidence comparing the proposed target definitions.
Tags
Full text
# Any research on label/target variable design for ML training? # Any research on label/target variable design for ML training? is there any discussion or paper about how to define/design the labels for the ML training? Intuitively I can think of: - Net return of the next future day - Net return using the max candle-high value of the next N candles - There is also a process described in the book "advances in financial machine learning" (https://www.amazon.com/Advances-Financial-Machine-Learning-Marcos/dp/1119482089) but the book lacks somehow on justification and evaluation. So in general: how to define your target variables (labels if classification) for stock predictions? Thank you! ## Answer by IHonda (score 2) https://quant.stackexchange.com/a/43610 A good beggining could be the paper of Gu, Kelly and Xiu (2018). ## Answer by Jacques Joubert (score 0) https://quant.stackexchange.com/a/44875 I see there that you mentioned the text book Advances in Financial Machine Learning. At the back of every chapter are a list of good papers that provide some insight into the chapters body of knowledge. Chapter 3 is titled Labeling and it is of course a large part of the process. I suggest reading all 39 papers in the bibliography. It was a real eye opener and they are well selected so you build a very deep understanding of the various techniques.
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.