Selecting Observation Windows for Behavioral Credit Default Models
Summary
The document raises a model-development question about choosing observation dates for behavioral credit scoring. Such models use account activity, arrears, advances, and other customer information to estimate whether an existing borrower will default during the following year. The author asks how to construct training observations when customers can contribute multiple non-default periods, while defaulted customers may have several possible pre-default observation dates.
It highlights choices that can affect the development sample: how far before default to measure predictors, whether to use multiple dates for a non-default customer, and how to handle accounts that default soon after opening. The example development period is 2014–2018, and the stated target is default within the next 12 months. The document provides no recommended sampling scheme, references, or empirical comparison. It is therefore useful as a framing of the observation-selection problem, but does not establish best practice or address related issues such as dependence among repeated observations or representativeness of the sample.
Key ideas
- Behavioral credit models use account information to predict default over a future horizon.
- Non-default customers may contribute multiple eligible observation periods.
- Defaulted customers can have several candidate observation dates before the event.
- The choice of observation date and window structure needs explicit justification.
- The document asks for guidance but does not provide a sampling method or references.
Tags
Full text
# Choosing observations/sample selection in behaviour credit scoring models # Choosing observations/sample selection in behaviour credit scoring models In retail banking the credit risk of a creditor after the credit had been granted is often modeled using behavioral credit scoring. In this setting the customer already has an account (or a few) and the bank can observe inflow, outflows, arrears and advances and use such information to assess the credibility of the customer. We can think of loans but the setting also applies to current accounts with the possibility to use overdraft. For the capital requirement we have to consider default within the forthcoming 12 months. To this aim we can think of risk factors observed at some point before the reference date and the outcome default/non-default. But there are many possible choices for when to observe the risk factors. Thus, what I was wondering is whether there are references or best practice examples on how to choose the observations. Say we develop a model on the years 2014-2018. - Then there are customers that never defaulted and I can use several one year periods for observations of non-defaults. - On the other hand I have defaulted customers. In such cases I can use observations before this very default. But which ones? 12 months before, 11 or 10? What if the customer defaults just (e.g.) 6 months after opening the account. The target to model "default during the forthcoming year" leaves open many degrees of freedom. Could anyone please point me to references where the structure of observations on the data set where the model is developed is clearly justified?
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.