基于可诱导动态谱风险度量的深度强化学习 | Stratmill
文章 arXiv papers · 作者: Anthony Coache et al.
总结
本文提出一种风险敏感型强化学习框架,智能体在其中优化时间一致的动态谱风险度量。该框架利用条件可诱导性构造严格一致的评分函数,并将其作为估计过程中的惩罚项。深度神经网络用于估计风险度量,作者还证明,这类度量可以由此类网络以任意精度逼近。
该框架还包含一种演员—评论家算法,可利用完整回合进行学习,而无需添加嵌套转移。作者在统计套利和投资组合配置两个应用中,使用模拟数据和真实数据,将该方法与嵌套模拟方法进行比较。说明中未提供数值结果、实现细节或真实数据集的信息,因此可据此了解设计和所述评估范围,但无法判断性能提升或稳健性。该方法适用于投资组合决策,其中控制损失分布与预期收益同样重要。
核心观点
- 该框架在强化学习中优化时间一致的动态谱风险度量。
- 条件可诱导性提供严格一致的评分函数,并将其用作估计惩罚项。
- 深度神经网络用于估计风险度量,文中还给出了关于该度量类别的逼近结果。
- 演员—评论家方法利用完整回合,避免增加嵌套转移。
- 比较涵盖嵌套模拟、统计套利和投资组合配置,并使用模拟数据和真实数据。
标签
全文
# Conditionally Elicitable Dynamic Risk Measures for Deep Reinforcement Learning # Conditionally Elicitable Dynamic Risk Measures for Deep Reinforcement Learning We propose a novel framework to solve risk-sensitive reinforcement learning (RL) problems where the agent optimises time-consistent dynamic spectral risk measures. Based on the notion of conditional elicitability, our methodology constructs (strictly consistent) scoring functions that are used as penalizers in the estimation procedure. Our contribution is threefold: we (i) devise an efficient approach to estimate a class of dynamic spectral risk measures with deep neural networks, (ii) prove that these dynamic spectral risk measures may be approximated to any arbitrary accuracy using deep neural networks, and (iii) develop a risk-sensitive actor-critic algorithm that uses full episodes and does not require any additional nested transitions. We compare our conceptually improved reinforcement learning algorithm with the nested simulation approach and illustrate its performance in two settings: statistical arbitrage and portfolio allocation on both simulated and real data.
在遵守原作品许可的前提下,附作者信息全文展示。 许可协议: abstract CC0
此摘要由 Stratmill 研究智能体根据原文撰写,并非原文副本。