Skip to content
All library documents

Designing Machine-Learning Labels for Stock Prediction

Article Quant Q&A · Author: mojovski

Summary

The document raises the problem of defining target variables for machine-learning models that predict stock behavior. Candidate labels include the next day’s net return and a return based on the maximum high over a future window. It also references a financial machine-learning text that discusses labeling methods, while noting the questioner’s concern that its justification and evaluation may be limited.

The replies offer starting points for further study rather than a specific labeling recipe: one recommends research by Gu, Kelly, and Xiu, while another directs readers to the labeling chapter’s bibliography for papers on different techniques. No empirical comparison is provided among return-based targets, future high-based targets, or classification labels. The document therefore frames label choice as a research design question, but leaves practical decisions—such as the forecast horizon, treatment of costs, evaluation criteria, and fit between labels and trading objectives—open for investigation.

Key ideas

  • A prediction target can use a future return or a statistic such as the maximum high over a horizon.
  • Label design is a central part of financial machine-learning research.
  • The replies recommend consulting published research and bibliography sources.
  • The document provides no evidence comparing the proposed target definitions.

Tags

Full text
# Any research on label/target variable design for ML training?


# Any research on label/target variable design for ML training?












is there any discussion or paper about how to define/design the labels for the ML training? Intuitively I can think of:

- Net return of the next future day

- Net return using the max candle-high value of the next N candles

- There is also a process described in the book "advances in financial machine learning" (https://www.amazon.com/Advances-Financial-Machine-Learning-Marcos/dp/1119482089) but the book lacks somehow on justification and evaluation.

So in general: how to define your target variables (labels if classification) for stock predictions?

Thank you!

## Answer by IHonda (score 2)

https://quant.stackexchange.com/a/43610

A good beggining could be the paper of Gu, Kelly and Xiu (2018).

## Answer by Jacques Joubert (score 0)

https://quant.stackexchange.com/a/44875

I see there that you mentioned the text book Advances in Financial Machine Learning. At the back of every chapter are a list of good papers that provide some insight into the chapters body of knowledge. Chapter 3 is titled Labeling and it is of course a large part of the process.

I suggest reading all 39 papers in the bibliography. It was a real eye opener and they are well selected so you build a very deep understanding of the various techniques.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.