Skip to content
All library documents

Choosing Regression or Classification Targets for Stock Models

Article BigQuant

Summary

The document explains how to choose between regression and classification when building machine-learning models for stock research. A regression task predicts a continuous quantity, illustrated by forecasting a stock’s five-day return. It introduces fitting a line by minimizing a loss function and mentions a random forest regressor as another example. A classification task instead predicts a discrete outcome, illustrated by whether the next day’s return exceeds a chosen threshold; binary labels represent the two outcomes, while multiclass setups are also noted.

The central lesson is to define the target according to the decision problem: regression produces a numeric estimate, whereas classification produces class probabilities. The article identifies data, algorithms, and models as components of machine learning and describes algorithms broadly enough to include the modeling process and parameter search. It gives conceptual examples and platform references, but no evaluation results, validation design, or guidance on avoiding leakage and overfitting. Predictions would need further evaluation and conversion into trading rules before they could support a strategy.

Key ideas

  • Regression is appropriate when the target is a continuous value such as a future return.
  • Classification is appropriate when the target is a discrete event such as exceeding a return threshold.
  • Binary labels encode two outcomes, while multiclass models can represent more categories.
  • Classification models can output probabilities for each class.
  • A prediction target should reflect the intended research or trading question.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.