Skip to content
All library documents

Designing Stock Labels for Supervised Quantitative Models

Article BigQuant

Summary

This tutorial explains how to construct target labels for supervised learning in equity selection. It distinguishes classification labels, which group observations into discrete outcomes, from continuous regression targets. Label definitions should reflect the prediction objective, such as future return or volatility, while avoiding excessively fine categories that can cause a model to learn noise.

Examples include labeling stocks by forward returns, clipping extreme values at the 1st and 99th percentiles, and dividing the resulting values into 20 equal-width bins. Other examples use volatility-adjusted returns or ranks, and the platform workflow combines labels with engineered features to form training data. The examples are implementation guidance, not evidence of predictive performance. The article does not report out-of-sample results, and forward-return labels require careful separation of training and test periods to avoid leakage.

Key ideas

  • Choose labels to match the outcome the model is intended to predict.
  • Continuous outcomes can be converted into discrete classes when classification is preferred.
  • Clipping extreme returns before binning can limit the influence of outliers.
  • Labels can represent forward returns, volatility-adjusted returns, or return ranks.
  • Overly granular categories may capture noise and weaken generalization.
  • The examples describe label construction, not verified model performance.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.