Skip to content
All library documents

Choosing Binary or Three-Class Labels for Flat Price Moves

Article Quant Q&A · Author: Hassan Sabree

Summary

The document asks how to label minute-by-minute stock price changes when the current and future closing prices can be equal. Its proposed setup assigns two classes for rises and falls, then considers adding a third class for no change. The practical issue is whether unchanged closes are ordinary observations and whether they warrant their own target class.

No answer or empirical analysis is included, so the document does not establish how often unchanged closes occur or which labeling scheme performs better. In practice, the choice depends on the prediction objective and how equality is defined in the data, including price precision and the forecast horizon. A three-class target distinguishes flat outcomes explicitly, while a binary target needs a rule for assigning or excluding them. The document raises this modeling question but gives no evidence about resulting model metrics.

Key ideas

  • Equal current and future closes create an outcome that needs an explicit labeling rule.
  • A three-class target can distinguish unchanged prices from rises and falls.
  • The document asks whether unchanged intraday closes are normal but provides no data to answer that question.
  • Classification metrics depend on how flat observations are handled.

Tags

Full text
# Binary or Multiclass Classification?


# Binary or Multiclass Classification?












So I've been using ensemble methods to model stock price movement, using intraday per-minute data in the OHLCV format, with the prediction being a 1 if the future close goes up, and 0 if it goes down. There are rows within the data that have the same closing price, and due to my somewhat limited understanding of how to interpret intraday data, I do not know if this is a normal occurrence.

If I factor in that the stock price may not move at all, then this becomes a multiclass classification problem, as I then have 3 outcomes to consider. This obviously entails a more involved process in producing error metrics, so I wanted to know if I am mistaken in thinking this is a multiclass problem. Thanks

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.