Quantizing Features for CatBoost Tree Models
Summary
This article explains quantization as mapping continuous or finely measured feature values into a smaller set of discrete levels. It describes how a histogram can help inspect a feature’s distribution and how cutoffs define bins, with values then encoded by a bin index or a representative level. It also discusses reconstructing approximate values from those bins and notes that choices of boundaries depend on the data and intended method. The practical focus is quantizing predictors for tree models, including a uniform quantization example and CatBoost’s use of saved borders. The article compares several boundary-selection approaches across example distributions and explains how CatBoost can load a quantization table when training. It frames financial data as often non-representative, so observed distributions may not cover future market behavior. The article offers conceptual and implementation guidance, but does not establish that a particular quantization method improves trading performance; the promised follow-up covers selecting tables for specific predictors.
Key ideas
- Quantization compresses feature values into discrete groups while accepting some loss of measurement precision.
- A histogram can reveal a feature’s distribution and help inform the placement of quantization boundaries.
- Values can be encoded by their bin index or by a representative boundary value.
- CatBoost can use stored border tables to quantize predictors during model training.
- Financial samples may not represent future market conditions, limiting conclusions drawn from their observed feature distributions.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.