跳至正文
返回文库全部文档

数据受限条件下借款人收入估算的联邦学习

文章 arXiv papers · 作者: Sultan Amed et al.

总结

该文提出一种联邦学习方法,在贷款机构无法合并申请人原始记录时估算借款人收入。各机构在数据保留于本地的同时训练共享模型。研究通过模拟由各州客户组成的联合体,使用历史贷款记录进行评估,并在后续观测数据上比较联邦估算、汇总数据模型和仅使用本地数据的模型。

总体而言,联邦学习的表现接近汇总数据基准;在这一设置中,数据量较少的客户平均表现优于该基准。研究还发现,在不同客户规模组中,联邦学习均优于仅使用本地数据训练的模型,数据最少的客户提升最大。研究将收入估算与按州和收入划分的债务收入比上限结合,进行回顾性审批分析;模拟审批数量有所增加,观察到的违约变化幅度较小。结果取决于所用数据集、客户划分方式和评估设计;该分析无法证明这些审批变化会在实际贷款业务或其他人群中重现。

核心观点

  • 联邦学习让机构在借款人记录保留于本地的同时,共同训练收入估算模型。
  • 评估将联邦模型与汇总数据模型以及机构本地模型进行比较。
  • 与仅使用本地数据进行估算相比,数据量较少的客户获益最明显。
  • 研究将联邦估算与定制的债务收入比上限结合,进行回顾性审批分析。
  • 报告结果仅适用于这组历史数据和模拟联合体。

标签

全文
# FedIncome: Federated Learning for Income Estimation in Digital Lending Under Data Sovereignty Constraints


# FedIncome: Federated Learning for Income Estimation in Digital Lending Under Data Sovereignty Constraints









Verified income is often unavailable in digital loan applications, forcing lenders to rely on reported income and potentially leading to over-lending, overly conservative offers, or rejection of creditworthy applicants. Cross-institutional data-sharing constraints make this problem especially difficult for smaller lenders with limited training data. We introduce FedIncome, a federated learning framework for income estimation that enables institutions to train a shared model without pooling raw borrower records. Using more than one million LendingClub loans partitioned into $50$ state-level clients, we simulate a heterogeneous lending consortium. The best federated model achieves out-of-time $R^2=0.608$, compared with $0.619$ for a pooled centralised benchmark. Small-sample clients obtain an average out-of-time $R^2$ improvement of $3.8$ percentage points relative to the pooled centralised benchmark, while the fitted client-level relationship places the empirical crossover at approximately $4,790$ training observations in this setting. When pooling is infeasible and the relevant alternative is local-only training, federation improves out-of-time performance across all sample-size groups, with the largest gains for data-scarce clients. We also combine federated income estimates with state- and income-specific debt-to-income thresholds. In a retrospective decision analysis, replacing reported income with the federated estimate increases simulated approval rates with only modest changes in observed default rates. FedIncome supports collaborative learning under data-locality constraints with little aggregate loss relative to pooled training and larger gains relative to local-only estimation.

在遵守原作品许可的前提下,附作者信息全文展示。 许可协议: abstract CC0

此摘要由 Stratmill 研究智能体根据原文撰写,并非原文副本。