How a Trading Label Is Discretized into Twenty Classes
Summary
This discussion clarifies how a quantitative platform’s automatic labeling feature maps a continuous score into twenty discrete classes. The questioner initially expects each class to contain an equal share of observations ranked by return, which would make adjacent class ranges non-overlapping. In the exchange, a responder explains that the classes apply to the label itself, and specifies the label as a ratio using a future close and a shifted open. The questioner later reruns the calculation and no longer observes the suspected overlap.
The exchange does not establish whether the platform uses equal-width bins, equal-frequency bins, or another precise rule. The suggestion that the result may use equal-distance bins remains speculation, while the final clarification defines discretization generally as converting a continuous variable into categorical values. For trading research, the key lesson is to verify the target’s formula and inspect actual bin boundaries rather than assume that labels represent ranked quantiles. The post offers no model comparison or evidence that one binning scheme improves predictive performance.
Key ideas
- The platform maps a continuous target into twenty categorical labels.
- The discussion distinguishes classifying the label from sorting observations directly by return.
- A responder describes the label using a shifted future close and open ratio.
- The exchange does not confirm the exact binning rule, such as equal-width or equal-frequency groups.
- Researchers should verify the target formula and inspect bin boundaries before interpreting class order.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.