Engineering Bank Transaction Features for Loan Default Prediction
Summary
The document describes a feature-engineering problem: turning a bank’s customer transaction history into inputs for a model that predicts loan default. It gives examples based on account balances, gambling expenditure, declined payments, ATM withdrawals relative to spending, and fines. Several proposed features aggregate transaction amounts or ratios by month, while transaction descriptions help categorize spending.
The material is a question seeking additional candidate variables, so it does not provide a tested feature set, model, or evidence that any listed signal predicts default. It also does not address validation, leakage, changing customer behavior, or whether labels and transaction histories are aligned in time. The examples illustrate how raw dated transactions can be summarized at customer level, but the document offers no basis for treating these spending categories as proof of irresponsibility or as reliable predictors without empirical testing.
Key ideas
- Transaction records can be aggregated into customer-level predictors for default modeling.
- Transaction descriptions can help categorize spending and payment events.
- Monthly amounts, balances, and ratios are examples of possible features.
- Candidate features require empirical validation before being treated as predictive.
Tags
Full text
# Using transaction data to predict default of the customer # Using transaction data to predict default of the customer I am trying to build a prediction model that utilize the huge transaction database of all the customers of a bank. My dataset currently looks like this: ``` transaction_date customer_id transaction_key transaction_amt cleaned_transaction_description 1 2017-03-22 137 AA_15 -50.00 ATM withdrawal 2 2017-03-30 129 AA_15 -50.00 ATM withdrawal 3 2017-03-14 142 AA_07 -100.00 grocery store Y 4 2017-03-20 120 AA_07 -30.00 clothing store X 5 2017-03-03 129 AA_07 -200.00 Pharmacy Z 6 2017-03-16 140 AA_20 -78.31 SEPA transaction ``` So, from a dataset like this, I am supposed to create variables, that could be used in a model that predicts the probability of default on a customer loan. So far I created variables like: - balance - the average net sum of all credits and debits for the customer ID - gaming - the average monthly sum of money spent on betting/gambling - stornos - the average monthly sum of declined payments (due to the lack of funds) - ATM to all - the average monthly ratio of ATM withdrawals to all the money spent - fines - the average monthly spending on fines As can be seen, most of the variables were created from the transaction description. I cannot think of more variables that I could create, and it seems like a waste of potential of the dataset. Can you think of more variables, that would indicate a "bad/irresponsible" customer? Thanks!
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.