Skip to content
All library documents

How XGBoost Ranking, Classification, and Regression Differ in Stock Selection

Article BigQuant

Summary

This tutorial compares three ways to train an XGBoost model for stock selection: ranking securities by a score, classifying outcomes into categories, and predicting a numeric target through regression. It frames these choices within a broader modeling workflow involving data, features, training, and combining models. The examples use Chinese equities and features such as recent returns, trading amounts, their ranks and ratios, and price-to-earnings data. For classification, the tutorial illustrates labels based on the relationship between a future closing price and a next-day opening price, including a three-class version.

For each approach, it shows sample predictions and feature-gain tables, which illustrate how outputs and reported feature importance differ across objectives. Ranking produces scores to order stocks, classification assigns outcome classes, and regression estimates a continuous value. These examples explain setup and interpretation, but they do not establish that one objective performs better: the document reports no comparative out-of-sample returns, trading costs, or risk-adjusted results. It also identifies itself as documentation for an older platform module, so implementation details may no longer apply to current tools.

Key ideas

  • Ranking produces scores that can be used to order stocks for selection.
  • Classification requires outcome labels and can assign either two classes or multiple classes.
  • Regression predicts a continuous target, while the tutorial uses future price-related outcomes as examples.
  • Feature-gain tables and sample predictions illustrate model outputs but do not prove trading performance.
  • The tutorial concerns an older version of the platform's modules.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.