Skip to content
All library documents

Using XGBoost to Model Quantitative Stock-Selection Factors

Article BigQuant

Summary

This article introduces XGBoost as a machine-learning method for quantitative stock selection using price and volume factors. It explains boosting as a process that adds weak learners in sequence, then contrasts AdaBoost’s reweighting of misclassified observations with gradient boosting trees’ effort to reduce residual errors. It describes XGBoost as an extension of gradient-boosted trees that uses a second-order approximation of the loss, adds regularization, samples features, handles sparse inputs, and supports parallel computation. It also lists model settings such as tree depth, learning rate, objective, sampling, and L1 or L2 penalties.

The empirical section begins by proposing factors derived from market data, including open, close, high, volume, and turnover, combined through statistical aggregations such as correlations, dispersion, extrema, sums, and weighted averages. However, the supplied text ends before specifying the factor formulas, prediction target, training design, portfolio construction, or experimental results. It therefore serves mainly as an algorithm and parameter overview rather than evidence that an XGBoost stock-selection strategy works. Practical use would require a clearly defined target, time-aware validation, and realistic transaction-cost and turnover assumptions.

Key ideas

  • Boosting combines sequentially trained weak learners to improve predictive performance.
  • XGBoost uses second-order loss information and explicit regularization, with feature sampling and sparse-input handling.
  • The article surveys model controls including tree depth, learning rate, subsampling, and L1 or L2 penalties.
  • It proposes building factors from price and volume data using statistical aggregations.
  • The supplied material does not include complete factor definitions or empirical performance results.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.