본문으로 건너뛰기
라이브러리 문서 전체

데이터 제약 환경의 차주 소득 추정을 위한 연합학습

기사 arXiv papers · 저자: Sultan Amed et al.

요약

이 문서는 대출 기관들이 신청자 원자료를 통합할 수 없을 때 차주의 소득을 추정하는 연합학습 접근법을 제시합니다. 각 기관은 데이터를 로컬에 보관하면서 공동 모델을 학습합니다. 평가는 과거 대출 기록을 이용해 주별 클라이언트로 구성된 컨소시엄을 시뮬레이션하고, 이후 관측치에서 연합 추정치와 통합 모델 및 로컬 전용 모델을 비교합니다.

이 설정에서 연합학습의 전반적인 성능은 통합 모델의 기준 성능에 가깝고, 데이터가 적은 클라이언트는 평균적으로 그 기준을 웃돕니다. 또한 클라이언트 규모별로 연합학습이 로컬 전용 학습보다 우수하며, 데이터가 가장 부족한 곳에서 개선 폭이 가장 큽니다. 소득 추정치와 주별·소득별 총부채상환비율 한도를 결합한 사후 대출 승인 분석에서는 관측된 부도율이 소폭 변하는 가운데 시뮬레이션상 승인 건수가 증가했다고 보고합니다. 결과는 이 데이터셋, 클라이언트 분할, 평가 설계에 좌우되며, 승인 변화가 실제 대출이나 다른 모집단에서도 나타날 것임을 입증하지 않습니다.

핵심 아이디어

  • 연합학습을 통해 기관들은 차주 기록을 로컬에 보관하면서 공동 소득 추정 모델을 학습할 수 있습니다.
  • 평가에서 연합 모델을 통합 학습 및 기관별 로컬 학습과 비교합니다.
  • 데이터가 적은 클라이언트에서 로컬 전용 추정보다 뚜렷한 이점이 나타납니다.
  • 사후 승인 분석에서 연합 추정치와 맞춤형 총부채상환비율 한도를 함께 적용합니다.
  • 보고된 결과는 과거 데이터와 시뮬레이션 컨소시엄에 한정됩니다.

태그

전문
# FedIncome: Federated Learning for Income Estimation in Digital Lending Under Data Sovereignty Constraints


# FedIncome: Federated Learning for Income Estimation in Digital Lending Under Data Sovereignty Constraints









Verified income is often unavailable in digital loan applications, forcing lenders to rely on reported income and potentially leading to over-lending, overly conservative offers, or rejection of creditworthy applicants. Cross-institutional data-sharing constraints make this problem especially difficult for smaller lenders with limited training data. We introduce FedIncome, a federated learning framework for income estimation that enables institutions to train a shared model without pooling raw borrower records. Using more than one million LendingClub loans partitioned into $50$ state-level clients, we simulate a heterogeneous lending consortium. The best federated model achieves out-of-time $R^2=0.608$, compared with $0.619$ for a pooled centralised benchmark. Small-sample clients obtain an average out-of-time $R^2$ improvement of $3.8$ percentage points relative to the pooled centralised benchmark, while the fitted client-level relationship places the empirical crossover at approximately $4,790$ training observations in this setting. When pooling is infeasible and the relevant alternative is local-only training, federation improves out-of-time performance across all sample-size groups, with the largest gains for data-scarce clients. We also combine federated income estimates with state- and income-specific debt-to-income thresholds. In a retrospective decision analysis, replacing reported income with the federated estimate increases simulated approval rates with only modest changes in observed default rates. FedIncome supports collaborative learning under data-locality constraints with little aggregate loss relative to pooled training and larger gains relative to local-only estimation.

출처의 라이선스에 따라 출처를 표시하고 전문을 공개합니다. 라이선스: abstract CC0

이 요약은 원문을 바탕으로 Stratmill의 리서치 에이전트가 작성했으며, 원문을 복사한 것이 아닙니다.