How XGBoost Ranking, Classification, and Regression Differ in Stock Selection
Summary
This tutorial compares three ways to train an XGBoost model for stock selection: ranking securities by a score, classifying outcomes into categories, and predicting a numeric target through regression. It frames these choices within a broader modeling workflow involving data, features, training, and combining models. The examples use Chinese equities and features such as recent returns, trading amounts, their ranks and ratios, and price-to-earnings data. For classification, the tutorial illustrates labels based on the relationship between a future closing price and a next-day opening price, including a three-class version.
For each approach, it shows sample predictions and feature-gain tables, which illustrate how outputs and reported feature importance differ across objectives. Ranking produces scores to order stocks, classification assigns outcome classes, and regression estimates a continuous value. These examples explain setup and interpretation, but they do not establish that one objective performs better: the document reports no comparative out-of-sample returns, trading costs, or risk-adjusted results. It also identifies itself as documentation for an older platform module, so implementation details may no longer apply to current tools.
Key ideas
- Ranking produces scores that can be used to order stocks for selection.
- Classification requires outcome labels and can assign either two classes or multiple classes.
- Regression predicts a continuous target, while the tutorial uses future price-related outcomes as examples.
- Feature-gain tables and sample predictions illustrate model outputs but do not prove trading performance.
- The tutorial concerns an older version of the platform's modules.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.