LightGBM for Predicting Next-Day NIFTY 50 Direction
Summary
The article introduces LightGBM, a tree-based gradient boosting method, and explains its leaf-wise tree growth, Gradient-based One-Side Sampling, and Exclusive Feature Bundling. It also describes histogram binning, categorical feature handling, and selected parameters such as tree depth, leaf count, and row or feature sampling. The article presents LightGBM as faster and less memory-intensive than alternatives, with a benchmark comparison to XGBoost, though the timing evidence comes from the described example rather than a broad evaluation.
Its trading example uses fifteen years of daily NIFTY 50 data and seven lagged daily-return features to classify whether the next day's price rises or falls. An 80/20 train-test split yields reported accuracies of 61% and 56%, respectively. The article also discusses confusion matrices, class-level F1 scores, and feature importance, suggesting that later lags may be removable. These classification metrics do not establish trading profitability: the document gives no transaction-cost, execution, or risk-adjusted return analysis, and treats a similar train and test accuracy as evidence against overfitting.
Key ideas
- LightGBM grows trees leaf-wise and uses gradient sampling and feature bundling to improve training efficiency.
- The example classifies next-day NIFTY 50 direction using seven lagged daily returns.
- The reported training and test accuracies are 61% and 56% after an 80/20 split.
- Confusion matrices, F1 scores, and feature importance are used to inspect classification behavior and input relevance.
- Classification accuracy alone does not show that a strategy is profitable after trading costs and execution effects.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.