基于动态凸风险度量的策略梯度强化学习
文章 arXiv papers · 作者: Anthony Coache et al.
总结
本文提出一种无模型强化学习方法,用于在结果通过动态凸风险度量评估时进行序贯优化。该方法采用时间一致的动态规划原理评估策略,并推导策略梯度更新规则以寻找更优策略。研究还提出使用神经网络的演员评论家方法来优化这些策略。
该方法在三项任务中进行演示:统计套利交易、金融对冲和机器人避障控制。这表明该框架应用于金融与非金融场景,但说明未提供表现数据、比较方法、市场假设或实施细节。因此,文本呈现了优化结构和应用范围,但未说明其交易效果和实际局限性的相关证据。
核心观点
- 动态凸风险度量可用于评估一系列不确定结果。
- 研究采用时间一致的动态规划评估候选策略。
- 策略梯度规则和基于神经网络的演员评论家方法可用于优化策略。
- 演示任务包括统计套利、金融对冲和机器人避障。
- 现有说明未提供比较结果或市场实施假设。
标签
全文
# Reinforcement Learning with Dynamic Convex Risk Measures # Reinforcement Learning with Dynamic Convex Risk Measures We develop an approach for solving time-consistent risk-sensitive stochastic optimization problems using model-free reinforcement learning (RL). Specifically, we assume agents assess the risk of a sequence of random variables using dynamic convex risk measures. We employ a time-consistent dynamic programming principle to determine the value of a particular policy, and develop policy gradient update rules that aid in obtaining optimal policies. We further develop an actor-critic style algorithm using neural networks to optimize over policies. Finally, we demonstrate the performance and flexibility of our approach by applying it to three optimization problems: statistical arbitrage trading strategies, financial hedging, and obstacle avoidance robot control.
在遵守原作品许可的前提下,附作者信息全文展示。 许可协议: abstract CC0
此摘要由 Stratmill 研究智能体根据原文撰写,并非原文副本。