コンテンツへスキップ
ライブラリの全資料

データ制約下での借り手所得推定と連合学習

記事 arXiv papers · 著者: Sultan Amed et al.

サマリー

この文書では、金融機関が申込者の生データを統合できない場合に、借り手の所得を推定する連合学習の手法を提示します。各機関はデータをローカルに保持したまま、共有モデルを学習します。評価では、過去の貸出記録を使って州別のクライアント群をシミュレートし、後続データに対する連合学習の推定値を、データを統合したモデルおよびローカルデータだけで学習したモデルと比較しています。

全体として連合学習の成績は統合モデルのベンチマークに近く、この設定ではデータの少ないクライアントが平均してそのベンチマークを上回ります。また、クライアントの規模を問わず、連合学習はローカルデータだけでの学習を上回り、データが最も少ないクライアントで改善が最大だったとしています。所得推定値を州別・所得別の債務所得比率上限と組み合わせ、過去データを使った融資承認分析を行った結果、観測された債務不履行の変化は小幅ながら、シミュレーション上の承認件数は増加しました。結果はこのデータセット、クライアントの分割、評価設計に依存しており、実際の融資や別の母集団でも承認の変化が生じることを立証するものではありません。

主なアイデア

  • 連合学習により、各機関は借り手の記録をローカルに保持したまま共有の所得推定モデルを学習できます。
  • 連合学習モデルを、データ統合型モデルと機関ごとのローカル学習モデルの両方と比較します。
  • データの少ないクライアントでは、ローカル推定のみの場合と比べた改善が最も明確です。
  • 連合学習の推定値を個別の債務所得比率上限と組み合わせ、過去データに基づく承認分析を行います。
  • 報告された結果は、過去のデータとシミュレーション上の機関群に固有です。

タグ

全文
# FedIncome: Federated Learning for Income Estimation in Digital Lending Under Data Sovereignty Constraints


# FedIncome: Federated Learning for Income Estimation in Digital Lending Under Data Sovereignty Constraints









Verified income is often unavailable in digital loan applications, forcing lenders to rely on reported income and potentially leading to over-lending, overly conservative offers, or rejection of creditworthy applicants. Cross-institutional data-sharing constraints make this problem especially difficult for smaller lenders with limited training data. We introduce FedIncome, a federated learning framework for income estimation that enables institutions to train a shared model without pooling raw borrower records. Using more than one million LendingClub loans partitioned into $50$ state-level clients, we simulate a heterogeneous lending consortium. The best federated model achieves out-of-time $R^2=0.608$, compared with $0.619$ for a pooled centralised benchmark. Small-sample clients obtain an average out-of-time $R^2$ improvement of $3.8$ percentage points relative to the pooled centralised benchmark, while the fitted client-level relationship places the empirical crossover at approximately $4,790$ training observations in this setting. When pooling is infeasible and the relevant alternative is local-only training, federation improves out-of-time performance across all sample-size groups, with the largest gains for data-scarce clients. We also combine federated income estimates with state- and income-specific debt-to-income thresholds. In a retrospective decision analysis, replacing reported income with the federated estimate increases simulated approval rates with only modest changes in observed default rates. FedIncome supports collaborative learning under data-locality constraints with little aggregate loss relative to pooled training and larger gains relative to local-only estimation.

出典を明記したうえで、ライセンスに従って全文を掲載しています。 ライセンス: abstract CC0

この要約は原文をもとにStratmillのリサーチエージェントが作成したもので、出典の複製ではありません。