Finding Data and Resources for Machine Learning Credit Scoring
Summary
The document outlines ways to begin building a machine learning credit scoring model for a thesis, focusing on public datasets, example competitions, and learning materials. It points to credit default prediction tasks as sources of data and published solutions, and mentions R and Python’s scikit-learn as common tools for modeling.
It also suggests exploring loan data released by peer-to-peer lenders and coursework projects on credit scoring. The main caveat is that useful predictive fields may be proprietary, so publicly available loan data may limit model performance. The document offers starting points rather than a detailed modeling workflow, validation method, or guidance on score calibration and fairness.
Key ideas
- Credit scoring can be studied through public loan default prediction competitions and their released solutions.
- Peer-to-peer lending platforms may provide data for developing and evaluating models.
- R and Python’s scikit-learn are identified as practical modeling tools.
- Public datasets may omit highly predictive proprietary fields, limiting achievable performance.
Tags
Full text
# How do I use machine learning to build a credit scoring model? # How do I use machine learning to build a credit scoring model? There are currently a lot of ways for credit scoring. The most popular one is the FICO score, and its variants. For my masters thesis, I would like to work on making my own credit scoring system using machine learning. The idea would be to obtain some real life data, and evaluate the credit scores, not necessarily in the 300-850 range as in the FICO score. What are some good resources to understand how to go about doing the same? Any new ideas are appreciated! Also, what are some places I could get free data (or not so expensive data) to build my model? ## Answer by Dom (score 10) https://quant.stackexchange.com/a/29928 One excellent resource is to try Kaggle and to examine some of the competitions, some of which are specifically on the application of machine learning to credit scoring. https://www.kaggle.com/c/GiveMeSomeCredit You wil see that the winning solution is made public, including source code and output. https://github.com/IdoZehori/Credit-Score/blob/master/Credit%20score.ipynb There are other problems in addition to this one so you should spend some time looking around. Here is another. https://www.kaggle.com/c/loan-default-prediction The preferred ML libraries are either in R or increasingly it seems that Python's Scikit learn is becoming very popular. Also note that there are a number of p2p loan platforms in the US (and now in the UK) that provide some loan data for such analysis. Google Prosper and Lending Club. One final point is that if there is a data field with high predictivity, the p2p providers may prefer to keep it proprietary. As a result it may be hard to find models for these loans with good AUC statistics. Stanford University also runs an ML course that covers credit scoring in the student projects submitted. Look here. http://cs229.stanford.edu/ I hope that is enough to get you started.
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.