Skip to content
All library documents

XGBoost for Factor-Based Stock Ranking and Portfolio Backtests

Article BigQuant

Summary

The article introduces XGBoost as a boosting ensemble and contrasts its sequential fitting of weak learners with the parallel training used in bagging. It explains residual-based improvement, complexity penalties, sample and feature subsampling, and parameters such as learning rate and estimator count. It also notes that feature calculations can be parallelized even though successive trees are trained sequentially.

Its A-share stock-selection example uses 18 factors and future five-day returns as the label, trains on 2010–2017 data, and evaluates predictions over 2017–2019. Each day it selects five top-ranked stocks, holds them for at least five days, and allocates more capital to higher ranks subject to a per-position cap. The reported comparison says XGBoost ran faster, while its strategy underperformed the random forest on return and drawdown in that test. The article is an older implementation; its results are limited to the described backtest and do not establish future performance.

Key ideas

  • Boosting trains learners sequentially, with later learners aimed at reducing residual error.
  • XGBoost adds complexity penalties and supports sample and feature subsampling to manage overfitting.
  • The stock-ranking example uses 18 factors to predict future five-day returns for A-share stocks.
  • The described backtest selects five top-ranked stocks daily and requires a minimum five-day holding period.
  • In the reported comparison, XGBoost ran faster but had lower returns and slightly higher maximum drawdown than random forest.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.