Preparing and Selecting Predictors for Financial Machine Learning
Summary
This article surveys preparation and selection of input variables for machine learning, with examples aimed at financial data. It recommends handling missing observations, checking near-constant predictors, reducing problematic correlations where appropriate, and removing linear dependencies among encoded factors. It also reviews scaling, standardization, transformations, imputation, dimensionality reduction, and conversion of categorical variables into dummy variables. A key procedural point is to estimate preprocessing parameters on training data and apply those same parameters to validation, test, and incoming observations.
The article also discusses target encoding and methods for assessing predictor relevance, interactions, and candidate subsets, including Random Uniform Forests and rough set analysis. It suggests that training examples themselves may be selected or optimized. Examples include classification measures and generated decision rules, but the supplied text is partial and results are not consistently reproducible across runs. Choices such as removing correlated or low-variance predictors depend on the model and data, so the recommendations are not universal rules.
Key ideas
- Fit preprocessing parameters on training data and reuse them for validation, test, and prediction data.
- Missing values, near-constant variables, correlations, and linear dependencies can affect model fitting and should be examined.
- Scaling, transformations, imputation, principal components, and dummy encoding are among the available preprocessing approaches.
- Predictor selection can consider relevance to the target, interactions among predictors, and the quality of predictor subsets.
- The article's example results vary between runs, and cleanup choices depend on the model.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.