Combining GBDT Feature Encoding with Logistic Regression for Stock Direction
Summary
This article explains a GBDT-plus-logistic-regression method for classifying stock direction. Gradient-boosted decision trees capture nonlinear feature relationships; each tree maps observations to leaf nodes, which are encoded as indicator variables. Logistic regression then estimates a probability from those encoded features. The workflow described covers feature and label preparation, training the tree model and encoder, fitting logistic regression on a separate data subset, generating predictions, and evaluating them against labels.
The example reports overall accuracy of 0.4930 and concludes that the model does not reliably predict rises and falls, attributing the weakness mainly to input feature quality. It describes confusion-matrix and ROC analysis and notes that a trading backtest could still have positive total return if successful predictions yield more than failed ones cost. However, no detailed backtest figures or validation design are supplied. The example is therefore useful as a modeling workflow, but its reported evidence does not establish predictive value or generalizability.
Key ideas
- GBDT can model nonlinear relationships and transform observations into tree-leaf indicator features.
- Logistic regression can estimate class probabilities from the encoded leaf features.
- The described training workflow separates data used to fit the trees from data used to fit the logistic regression stage.
- The example reports 0.4930 overall accuracy and says the model did not demonstrate useful directional prediction.
- A favorable trading return despite weak classification accuracy depends on gains from correct calls outweighing losses from incorrect calls.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.