Skip to content
All library documents

Encoding Credit Ratings for Machine Learning Models

Article Quant Q&A · Author: wanna_be_quant

Summary

The document considers how to encode credit ratings as categorical features when the distance between rating notches may not be linear. It notes ordinal encoding as a common approach and mentions probit, logistic, and exponential transformations as alternatives raised by the questioner.

A response suggests James–Stein encoding, a target-based method, and points to its use in a credit-scoring study. It also identifies weight of evidence encoding as a commonly used target-based technique in credit scoring. These suggestions offer candidate methods for representing ratings in predictive models, but the document does not compare their performance or explain how to implement them. The best encoding depends on the target, available data, and validation design; the brief exchange does not establish that any method captures rating-notch nonlinearity better than the others.

Key ideas

  • Ordinal encoding is presented as a standard way to represent credit ratings.
  • The question raises probit, logistic, and exponential transformations as possible alternatives.
  • James–Stein encoding is suggested as a target-based approach for credit-scoring data.
  • Weight of evidence encoding is identified as another target-based option used in credit scoring.
  • The document provides suggestions rather than a comparative evaluation of encoding performance.

Tags

Full text
# Efficient encoding technique for credit ratings


# Efficient encoding technique for credit ratings












Is there any categorical encoding technique for credit ratings that take into account the kind of non linear nature of the notches of the credit ratings?

The literature standard is the ordinal one unless I have missed something. Some other papers have tried probit and logistic/exponential encoding (?)

Any ideas are really welcome

## Answer by William Wu (score 1)

https://quant.stackexchange.com/a/71685

Other than the ones you have mentioned, James-Stein encoding (a target-based encoder) could be a possible idea.

James-Stein encoding was used in this paper titled "Machine Learning approach for Credit Scoring".

Also, this Kaggle website lists 11 categorical encoders. Notably, "Weight Of Evidence" is quoted as a commonly used target-based encoder in credit scoring.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.