Skip to content
All library documents

Sampling Data for Logistic Credit-Risk Models

Article Quant Q&A · Author: sergio.azevedo

Summary

The document considers how to assemble observations for a logistic regression that estimates insolvency probabilities and produces credit scores. The question arises when a database contains repeated monthly records for each company, raising concerns about dependence between observations and whether firms have had enough time to default. The proposed sampling advice includes selecting one record per company, choosing observations with an adequate performance history, and using financial ratios as predictors.

The answer emphasizes random sampling, sufficient sample size, and including relevant predictors such as credit history and industry. It also recommends balancing solvent and insolvent cases, although balancing outcomes can affect how predicted probabilities represent the population and may require adjustment. The discussion gives general suggestions rather than a validated sampling design; it does not resolve how to define the prediction horizon, account for repeated observations, or evaluate out-of-sample calibration.

Key ideas

  • Repeated monthly records from the same company can create dependence that matters for model design.
  • Sampling should represent the population and provide enough cases for estimation.
  • Predictors may include financial ratios, credit history, and industry controls.
  • Outcome balancing may help model fitting, but its effect on population probability estimates needs consideration.
  • Sampling should align the observation period with the horizon over which insolvency is predicted.

Tags

Full text
# Sample selection for a credit scoring model


# Sample selection for a credit scoring model












Suppose that our goal is to fit a logistic regression in order to obtain insolvency probabilities for possible credit takers and produce a credit risk score.

If I have a database available that contains monthly data for every customer and a binary variable that says whether said company is solvent or not, what are the sampling best practices?

A friend gave me the following advice:

- Select one row of data for each company in order to avoid auto-correlation and other types of association that violate the GLM assumptions;

- Select a large enough month-of-book in order to match the contract lengths and obtain a good amount of insolvents if possible. For example, if the average operation length is 2 years, we should select only observations that are already 6 months or 1 year old, otherwise they had little time to "become insolvent";

- Use financial and economical ratios as covariates, such as $EBITDA/Loans$ and $Sales_t/ Sale_{t-1}.$

What are the best practices when assembling a sample to fit my statistical model?

## Answer by Sane (score 1, accepted)

https://quant.stackexchange.com/a/79227

A couple of remarks:

- $\textbf{Ensure random sampling}$: It is important to randomly select your sample from the dataset to avoid bias and ensure that the sample is representative of the population.

- $\textbf{Balance between solvent and insolvent cases}$: Ensure that your sample includes a balanced number of solvent and insolvent cases to avoid any class imbalance issues in your logistic regression model.

- $\textbf{Adequate sample size}$: Make sure that your sample size is large enough to provide sufficient power for your logistic regression model. A larger sample size will also help in generalizing the results to the population.

- $\textbf{Controls}$: Among financial and economic ratios, include other relevant controls as well, such as credit history, sector/industry.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.