Skip to content
All library documents

Boruta-Shap Feature Selection with CPU Parallelism and GPU Acceleration

Article QuantInsti blog

Summary

Boruta-Shap combines Boruta’s comparison of original features against shuffled versions with Shapley-based importance estimates. The described workflow uses a tree-based model to assess tentative features across repeated trials, counts how often features register as important, then applies probability-based thresholds to classify them as selected, tentative, or rejected. The article presents this as a way to rank features while accounting for interactions and to prepare inputs for later trading-strategy backtests.

To reduce runtime, the implementation groups trials into batches according to the available CPU worker threads and runs each batch concurrently; it also describes configuring an XGBoost classifier to use a GPU. The account gives a procedural outline but no benchmark timings, accuracy results, or trading performance evidence. It does not establish that selected features will predict future returns, and feature selection alone does not address leakage, validation design, or overfitting. Its recommendation to use the output before strategy parameter optimization should therefore be treated as a proposed workflow, not a demonstrated advantage.

Key ideas

  • Boruta-Shap compares original features with shuffled counterparts and uses Shapley values to estimate importance.
  • Repeated trials count feature hits, which are evaluated against thresholds to select or reject features.
  • CPU worker threads can process batches of trials concurrently, while the described classifier can use GPU computation.
  • Feature importance rankings do not by themselves show that features will predict future trading returns.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.