Information Value Can Exceed One in Credit Scoring
Summary
This note answers whether a credit-scoring variable's Information Value (IV) can exceed one and what happens when a feature perfectly separates good and bad outcomes. IV sums the differences between the good and bad population shares in each category, weighted by that category's Weight of Evidence. With two categories that approach perfect separation, the response reduces IV to a difference of log-ratios. As the cross-category shares approach zero, the log terms can grow without bound, so IV has no finite theoretical maximum under this setup.
The response says values above one are possible and describes values below ten as common in practice. Those practical observations are not supported here by a dataset or further evidence, and the note does not detail smoothing or handling empty categories. In a perfect-separation case, the formulas involve logarithms of zero, so the unsmoothed calculation is undefined; the limiting argument explains why IV can become arbitrarily large rather than yielding a finite perfect-fit value.
Key ideas
- Information Value is computed by weighting category-level differences in good and bad shares by Weight of Evidence.
- Information Value can exceed one and has no finite theoretical upper bound in the limiting perfect-separation case.
- Perfect separation causes zero category shares, making the unsmoothed logarithmic formulas undefined.
- The response reports that practical IV values are usually below ten but gives no supporting sample or methodology.
- The document does not specify how to smooth or otherwise handle empty categories.
Tags
Full text
# Credit Scoring model IV max value
# Credit Scoring model IV max value
I'm making a credit scoring model and I get that one variable has Information Value (IV) more than 1, is that possible?
Formulas are pretty simple for weight of evidence(WoE) and information value(IV)
$$WoE_i = \log \left( \dfrac{\dfrac{g_i}{g}}{\dfrac{b_i}{b}} \right)$$ where $g_i$ represents the number of goods (no default) in category $i$ of variable $x_i$, $b_i$ represents the number of bads (default) in category $i$ of variable $x_i$, $g$ represents the number of goods (no default) in the entire dataset, $b$ represents the number of bads (default) in the entire dataset, $N(x)$ is the number of levels in the variable $x$, that is number of categories $$IV = \sum_{i=1}^{N(x)}\left( \dfrac{g_i}{g} - \dfrac{b_i}{b} \right) \cdot WoE_i$$
Also, what is a perfect fit in the model?
By the perfect fit I understand that there is just two categories in x: the first includes all goods and the second includes all bads. In that case when computing $WoE_1$ I get 0 in $log$ denominator, because $b_1 = 0$. When computing $WoE_2$ I get 0 in $log$ numerator, because $g_2 = 0$. Does that make sense?
## Answer by Magic is in the chain (score 3, accepted)
https://quant.stackexchange.com/a/45974
IV of greater than 1 is possible, and very common. Assume you are using the natural logarithm? You can deduce the limit by using two categories. Say the model/feature can separate the good and bad perfectly, and let $G_1, B_1, G_2, B_2$ be the proportion of goods and bads falling in the two categories. Then let $G_1 \to 1, B_1 \to 0, G_2 \to 0, B_2 \to 1$. So we are assuming a large pool and the Model is perfectly separating the goods/bads. Then easy to check that that the IV equals:
$\ln \left( \frac{G_1}{B_1}\right) -\ln \left( \frac{G_2}{B_2}\right) $
And minus sign in front of the second means you can invert the second ratio, and hence it is 2 times the log of a quantity that goes to infinity, so you can get infinity. But in practice you should be getting IV of less than 10 most of the time.Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.