Reducing Trading Feature Dimensions with Truncated SVD and NMF
Summary
This article introduces dimensionality reduction as a way to manage large, correlated sets of trading indicators. It presents truncated singular value decomposition (SVD), which projects centered data onto a smaller number of components, and non-negative matrix factorization (NMF), which represents data using non-negative factors. The discussion covers selecting component counts, fitting transformations, and trade-offs such as interpretability, variance retention, and sensitivity to data requirements. An example uses EURUSD indicator features in a linear regression model: reducing dimensions with principal components is reported to improve the test score relative to the unreduced setup, while the model still shows a generalization gap.
These results are illustrative rather than conclusive. The article does not establish that either technique improves live trading returns, and the reported model evaluation is limited to the described dataset and procedure. Correlated indicators, overfitting, outliers, non-unique factor solutions, and hyperparameter choices can all affect results. Dimension reduction should therefore be assessed with careful out-of-sample validation, and the comparison alone does not identify a universally superior method.
Key ideas
- High-dimensional sets of correlated indicators can increase computation and make models more prone to overfitting.
- Truncated SVD reduces feature dimensions by retaining leading components of a centered data matrix.
- NMF decomposes non-negative data into lower-dimensional non-negative factors that may be easier to interpret.
- The EURUSD regression example reports improved test performance after dimension reduction but retains limitations in generalization.
- Both methods require careful choices about component count and evaluation on unseen data.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.