動的凸リスク尺度を用いた方策勾配強化学習
記事 arXiv papers · 著者: Anthony Coache et al.
サマリー
本稿は、動的凸リスク尺度で成果を評価する逐次最適化のため、モデルフリーの強化学習手法を提示します。時間整合的な動的計画法の原理で方策を評価し、より良い方策を見つけるための方策勾配更新を導きます。ニューラルネットワークを用いたアクター・クリティック手法で方策を最適化します。
この手法は、統計的裁定取引、金融ヘッジ、ロボットの障害物回避制御という3つの課題で実演されています。金融分野と非金融分野の双方への適用を示していますが、説明には成績数値、比較手法、市場の仮定、実装の詳細がありません。そのため、最適化の構造と適用範囲は示される一方、取引の有効性や実務上の限界を示す証拠は明らかにされていません。
主なアイデア
- 動的凸リスク尺度により、不確実な結果の系列を評価できます。
- 候補方策の評価には、時間整合的な動的計画法を用います。
- 方策勾配ルールとニューラルネットワークのアクター・クリティック手法で方策を最適化します。
- 統計的裁定取引、金融ヘッジ、ロボットの障害物回避を実演課題とします。
- 提供された説明には比較結果や市場実装の前提がありません。
タグ
全文
# Reinforcement Learning with Dynamic Convex Risk Measures # Reinforcement Learning with Dynamic Convex Risk Measures We develop an approach for solving time-consistent risk-sensitive stochastic optimization problems using model-free reinforcement learning (RL). Specifically, we assume agents assess the risk of a sequence of random variables using dynamic convex risk measures. We employ a time-consistent dynamic programming principle to determine the value of a particular policy, and develop policy gradient update rules that aid in obtaining optimal policies. We further develop an actor-critic style algorithm using neural networks to optimize over policies. Finally, we demonstrate the performance and flexibility of our approach by applying it to three optimization problems: statistical arbitrage trading strategies, financial hedging, and obstacle avoidance robot control.
出典を明記したうえで、ライセンスに従って全文を掲載しています。 ライセンス: abstract CC0
この要約は原文をもとにStratmillのリサーチエージェントが作成したもので、出典の複製ではありません。