Skip to content
All library documents

Federated Learning for Borrower Income Estimation Under Data Constraints

Article arXiv papers · Author: Sultan Amed et al.

Summary

The document presents a federated learning approach for estimating borrower income when lenders cannot combine raw applicant records. Institutions train a shared model while keeping data local. Its evaluation simulates a consortium of state-based clients using historical lending records and compares federated estimates with pooled and local-only models on later observations.

Federated performance is close to the pooled benchmark overall, and small-data clients outperform that benchmark on average in this setting. The study also finds that federation beats local-only training across client size groups, with the largest improvements for the least data-rich. It combines income estimates with state- and income-specific debt-to-income limits in a retrospective approval analysis, reporting higher simulated approvals with modest changes in observed defaults. Results depend on this dataset, client partition, and evaluation design; the analysis does not establish that the approval changes would hold in live lending or other populations.

Key ideas

  • Federated learning lets institutions train a shared income estimator while retaining borrower records locally.
  • The evaluation compares federated models with both pooled and institution-local training.
  • Small-data clients see the clearest benefits over local-only estimation.
  • Federated estimates are paired with tailored debt-to-income limits in a retrospective approval analysis.
  • The reported results are specific to the historical data and simulated consortium.

Tags

Full text
# FedIncome: Federated Learning for Income Estimation in Digital Lending Under Data Sovereignty Constraints


# FedIncome: Federated Learning for Income Estimation in Digital Lending Under Data Sovereignty Constraints









Verified income is often unavailable in digital loan applications, forcing lenders to rely on reported income and potentially leading to over-lending, overly conservative offers, or rejection of creditworthy applicants. Cross-institutional data-sharing constraints make this problem especially difficult for smaller lenders with limited training data. We introduce FedIncome, a federated learning framework for income estimation that enables institutions to train a shared model without pooling raw borrower records. Using more than one million LendingClub loans partitioned into $50$ state-level clients, we simulate a heterogeneous lending consortium. The best federated model achieves out-of-time $R^2=0.608$, compared with $0.619$ for a pooled centralised benchmark. Small-sample clients obtain an average out-of-time $R^2$ improvement of $3.8$ percentage points relative to the pooled centralised benchmark, while the fitted client-level relationship places the empirical crossover at approximately $4,790$ training observations in this setting. When pooling is infeasible and the relevant alternative is local-only training, federation improves out-of-time performance across all sample-size groups, with the largest gains for data-scarce clients. We also combine federated income estimates with state- and income-specific debt-to-income thresholds. In a retrospective decision analysis, replacing reported income with the federated estimate increases simulated approval rates with only modest changes in observed default rates. FedIncome supports collaborative learning under data-locality constraints with little aggregate loss relative to pooled training and larger gains relative to local-only estimation.

Shown in full with attribution under the source's licence. Licence: abstract CC0

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.