Quantizing CatBoost Inputs and Selecting Predictor Preprocessing Methods
Summary
The article presents an MQL5 workflow for preprocessing training data and selecting quantization tables for CatBoost predictors. A script loads a sample, tests candidate tables by measuring reconstruction error for each predictor, and records the best options under a chosen criterion. It also describes handling rare predictor values as spikes: identify tails using rank frequencies, then either move those values near a split or replace them with randomized values drawn from the normal range. Related options save transformed samples, annotate spike counts, or remove rows containing outliers.
Other preprocessing steps include excluding highly correlated predictors with alternative selection rules and filtering predictors whose mean shifts across segments of the sample. The article also compares uniform and randomized quantization and describes repeatable random searches. Its experiment suggests table selection can affect model results, but the author cautions that a single experiment cannot establish one method as best. The reported selection relies on approximation error, and other evaluation criteria and broader tests may be useful; the described preprocessing methods are not individually validated for effectiveness.
Key ideas
- Candidate quantization tables are compared by their reconstruction error for each predictor.
- Rare predictor values can be transformed, annotated, or removed, with different choices affecting the training data.
- Correlated predictors and predictors with shifting segment means can be identified for possible exclusion.
- Uniform and randomized quantization offer different splits, but the article does not establish one as universally superior.
- The experiment is limited, and approximation error is only one possible criterion for choosing tables.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.