跳至正文
返回文库全部文档

用于交易的凸风险目标强化学习

文章 arXiv papers · 作者: Shanyu Han et al.

总结

这篇论文提出一个强化学习框架,用于处理以凸评分函数表示风险目标的决策问题。该框架可表示方差、预期损失和熵风险价值等指标,也可表示均值风险效用。其重点是让这些目标能够在时间维度上发挥作用,因为优化风险可能导致时间不一致:早期更受偏好的策略,在之后的决策时点可能不再受偏好。

为解决这一问题,作者在状态中加入辅助变量,并将问题重新表述为双状态优化。他们介绍了定制的演员评论家算法、不要求连续马尔可夫决策过程的理论近似保证,以及一种受交替最小化启发的辅助变量抽样流程。论文报告了模拟实验,包括一个统计套利应用,以此作为方法有效性的证据。现有说明没有给出数值结果或详细实验设定;抽样方法的收敛性也取决于所述假设,因此无法确定其在已测试模拟之外的实际表现。

核心观点

  • 凸评分函数可在强化学习目标中表示多种常见风险指标。
  • 在状态中加入辅助变量有助于解决风险敏感决策中的时间不一致问题。
  • 所提出的演员评论家方法带有近似保证,且不要求连续状态过程。
  • 据报告,受交替最小化启发的抽样方法在特定条件下收敛。
  • 模拟实验包括一项统计套利交易应用。

标签

全文
# Risk-sensitive Reinforcement Learning Based on Convex Scoring Functions


# Risk-sensitive Reinforcement Learning Based on Convex Scoring Functions









We propose a reinforcement learning (RL) framework under a broad class of risk objectives, characterized by convex scoring functions. This class covers many common risk measures, such as variance, Expected Shortfall, entropic Value-at-Risk, and mean-risk utility. To resolve the time-inconsistency issue, we consider an augmented state space and an auxiliary variable and recast the problem as a two-state optimization problem. We propose a customized Actor-Critic algorithm and establish some theoretical approximation guarantees. A key theoretical contribution is that our results do not require the Markov decision process to be continuous. Additionally, we propose an auxiliary variable sampling method inspired by the alternating minimization algorithm, which is convergent under certain conditions. We validate our approach in simulation experiments with a financial application in statistical arbitrage trading, demonstrating the effectiveness of the algorithm.

在遵守原作品许可的前提下,附作者信息全文展示。 许可协议: abstract CC0

此摘要由 Stratmill 研究智能体根据原文撰写,并非原文副本。