Machine Learning Feature Selection for Trading Signals
Summary
This article presents a machine learning workflow for exploring candidate predictors in a simple trading system. It discusses data mining bias, feature construction, preprocessing, removing correlated inputs, and several selection methods, including maximal information criteria, recursive elimination, embedded model selection, and Boruta. Candidate variables include returns, trend, price movement, volatility ratios, the Market Meanness Index, and the Hurst exponent. The article also tests principal components as a transformation before fitting a random forest model.
The reported analyses repeatedly favor ratios of short- and long-term volatility and a price-change oscillator, while results differ across selection methods. PCA modestly improves the model’s resampling performance in the example, but the author notes that PCA is linear and assumes future data retain the training data’s component structure. The trading results are explicitly considered upwardly biased because feature selection used the same dataset; rolling feature selection is suggested as a more informative follow-up. These findings are illustrative and do not establish a robust live edge.
Key ideas
- Feature selection can help narrow candidate predictors, but searching many patterns creates data mining bias.
- The analysis compares correlation filtering, recursive elimination, model-based selection, Boruta, and principal components.
- Volatility ratios and a price-change oscillator emerge as prominent variables across several methods.
- Recursive elimination may miss interactions, and RMSE may underweight errors in extreme outcomes.
- The example strategy’s results may be biased because feature selection was informed by the same dataset.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.