Skip to content
All library documents

Choosing Logistic or Poisson Models for Default Probability

Article Quant Q&A · Author: user9406

Summary

The document compares logistic regression with Poisson-based generalized linear models for retail credit default probability. It favors logistic regression as the more common choice for this purpose because default is a categorical outcome, while Poisson models are typically used for counts. The logistic function also produces probabilities bounded between zero and one, which makes its output directly interpretable as a probability.

The answers cite predictive performance, implementation simplicity, and potential overdispersion in Poisson models as reasons logistic regression may be preferred. They also describe limits of Poisson intensity models, including difficulty representing clustered defaults and aggregating results across time horizons when future risk factors are unknown. These are broad observations rather than a reported head-to-head experiment: model performance can vary with the sample, economic cycle, and application. The document recommends comparing models on the specific data, using measures such as accuracy ratio and ROC analysis, rather than treating one model as universally superior.

Key ideas

  • Logistic regression is commonly used when the target is a binary default outcome.
  • Its output is bounded between zero and one, so it can represent a probability directly.
  • Poisson models are associated with count data and can face overdispersion and default-clustering limitations.
  • Compare candidate models on the relevant data because performance may depend on the sample and economic conditions.
  • The document describes general industry preferences, not a universal rule backed by a reported experiment.

Tags

Full text
# Answer by BCLC (score 3)


# For Probability of Default in retail credit what is more popular logistic regression or GLM with Poisson distribution and why?












Trying to understand which regression model is more popular in retail credit card industry Logistic regression or GLM with Poisson distribution and why?

## Answer by BCLC (score 3)

https://quant.stackexchange.com/a/16653

From http://rmi.nus.edu.sg/gcr/files/04%20GCR%20vol%201.pdf

- "One of the attractive features of the logistic function is the fact that it is bounded between 0 and 1, making it suitable to represent probabilities. "

- "The Poisson intensity model introduced in this article still has serious shortcomings despite the major advancement offered by its dynamic features. First, it is known to be unable to properly capture the clustered default phenomenon such as is documented in Das et al. (2007). Another limitation is that the time aggregation to different horizons is easy in principle but difficult in reality. The Poisson intensity is a known function of common risk factors and individual firm attributes. For time aggregation to get to a longer horizon of interest, one must prescribe the dynamic processes for all these variables whose future values are unknown. The dimension of the dyna"

- etc

## Answer by Priyanka Gupta (score 1)

https://quant.stackexchange.com/a/33040

http://jgscott.github.io/SDS325H_Spring2015/files/logit_poisson_cox.pdf

Based on my understanding from reading the above document, I think it could be because Poisson is used for count data and Logistic is used for categorical data and we have a categorical data while doing Probability of Default (PD) modelling.

## Answer by Quantopik (score 0)

https://quant.stackexchange.com/a/17085

In the credit modelling industry is more popular the use of the logistic regression with respect to the Poisson one.

This is for several reasons. Here I listed the main ones:

1) The Logistic regression is empirically shown to be better in describing that kind of phaenomenona in terms of forecasting performances and predictive capacity (try to compare the performance ratio for both of them: Accuracy ratio, ROC,...).

2) The Logistic regression suffers less the overdispersion problem that is a features of the Poisson regression models and only sometimes can be solved by using a Bivariate regression model, as, for instance, in the health care industry analysis case.

3) The Logistic regression is simpler to be implemented with respect to both the programming point of view and a theoretical point of view.

This is as regards the estimation of the probability of default. In other cases, this could be not completely true.

I suggest you to read Categorical Data Analysis by Agresti, to have a more deepen knowledge of this topic (from an econometric point of view) and, moreover, to try to test which model is better; generally, it is as above, but it depends on the economic cycle, data sample,... etc.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.