Skip to content
All library documents

Using GBDT Leaf Features to Improve Logistic Regression

Article BigQuant

Summary

The document explains a two-stage model that uses gradient-boosted decision trees to create features for logistic regression. The trees split and combine the original inputs; each sample is then represented by a binary vector marking the leaf it reaches in each tree. Logistic regression learns from these leaf indicators, which can encode nonlinear relationships and feature interactions that a linear model would otherwise need researchers to specify by hand.

The article outlines preprocessing, splitting training data between the tree and logistic-regression stages, one-hot encoding the leaf assignments, and tuning with stratified folds and AUC. It gives a basic scikit-learn workflow but no measured trading results or comparative performance evidence. It warns that model quality depends on feature quality, data noise, and the application context. It also suggests alternatives such as substituting other models or cross-validating the tree stage, without evaluating those variants.

Key ideas

  • GBDT partitions inputs into leaves that can serve as automatically generated categorical features.
  • A binary indicator for each tree leaf represents whether a sample falls into that leaf.
  • Logistic regression can learn from these indicators to capture nonlinear patterns and feature combinations.
  • The example separates data for fitting the tree stage and training the logistic-regression stage to reduce overfitting risk.
  • The document offers a workflow and tuning suggestions but reports no trading performance results.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.