コンテンツへスキップ
ライブラリの全資料

深層強化学習と誘導可能な動的スペクトルリスク尺度

記事 arXiv papers · 著者: Anthony Coache et al.

サマリー

この文書では、エージェントが時間整合的な動的スペクトルリスク尺度を最適化する、リスク感応型の強化学習の枠組みを提示します。条件付きエリシタビリティを用いて厳密に整合的なスコア関数を構築し、推定時のペナルティとして機能させます。深層ニューラルネットワークでリスク尺度を推定し、この種の尺度を任意の精度で近似できることを証明しています。

この枠組みには、入れ子状の遷移を追加せずに完全なエピソードから学習するアクター・クリティック法も含まれます。著者らは、シミュレーションデータと実データを用いた統計的裁定とポートフォリオ配分の2つの応用例で、入れ子型シミュレーション手法と比較しています。説明には定量的な結果、実装の詳細、実データの情報がないため、設計と記載された評価範囲は把握できますが、パフォーマンスの向上や頑健性は判断できません。この手法は、期待リターンに加えて損失分布の管理が重要となるポートフォリオ判断に関係します。

主なアイデア

  • この枠組みは、強化学習で時間整合的な動的スペクトルリスク尺度を最適化します。
  • 条件付きエリシタビリティにより、推定ペナルティに用いる厳密に整合的なスコア関数を構築します。
  • 深層ニューラルネットワークでリスク尺度を推定し、尺度のクラスに関する近似結果を示します。
  • アクター・クリティック法は完全なエピソードを使い、追加の入れ子状遷移を避けます。
  • 比較では、シミュレーションと実データを用いて、入れ子型シミュレーション、統計的裁定、ポートフォリオ配分を扱います。

タグ

全文
# Conditionally Elicitable Dynamic Risk Measures for Deep Reinforcement Learning


# Conditionally Elicitable Dynamic Risk Measures for Deep Reinforcement Learning









We propose a novel framework to solve risk-sensitive reinforcement learning (RL) problems where the agent optimises time-consistent dynamic spectral risk measures. Based on the notion of conditional elicitability, our methodology constructs (strictly consistent) scoring functions that are used as penalizers in the estimation procedure. Our contribution is threefold: we (i) devise an efficient approach to estimate a class of dynamic spectral risk measures with deep neural networks, (ii) prove that these dynamic spectral risk measures may be approximated to any arbitrary accuracy using deep neural networks, and (iii) develop a risk-sensitive actor-critic algorithm that uses full episodes and does not require any additional nested transitions. We compare our conceptually improved reinforcement learning algorithm with the nested simulation approach and illustrate its performance in two settings: statistical arbitrage and portfolio allocation on both simulated and real data.

出典を明記したうえで、ライセンスに従って全文を掲載しています。 ライセンス: abstract CC0

この要約は原文をもとにStratmillのリサーチエージェントが作成したもので、出典の複製ではありません。