跳至正文
返回文库全部文档

面向投资组合策略的稳健风险感知强化学习

文章 arXiv papers · 作者: Sebastian Jaimungal et al.

总结

这项研究提出一种强化学习方法,依据风险感知的绩效标准优化策略。研究使用秩依赖期望效用评估策略,使智能体能够表达对收益和下行结果的不同偏好。为应对模型不确定性,研究在围绕模型分布的 Wasserstein 距离球内,采用最坏情形收益分布评估策略。这形成了一个嵌套决策问题:智能体选择策略,随后对手选择一个邻近分布来降低该策略的表现。

作者推导了策略选择问题和对抗分布问题各自的策略梯度公式,并在稳健投资组合配置、基准优化和统计套利中展示了该方法。这些应用体现了框架的目标范围,但文档没有提供比较绩效数据、市场数据细节或实际实施指引。因此,该方法是否有用,取决于能否在具体交易环境中进一步验证模型、不确定性集合和风险偏好。

核心观点

  • 秩依赖期望效用使策略能够体现收益与下行风险之间的不同权衡。
  • 该框架在 Wasserstein 邻域内用最坏情形分布评估每项策略。
  • 对抗性的内层问题会寻找降低所选策略表现的分布。
  • 作者分别为策略优化和对抗问题推导了策略梯度。
  • 示例涵盖投资组合配置、基准优化和统计套利。

标签

全文
# Robust Risk-Aware Reinforcement Learning


# Robust Risk-Aware Reinforcement Learning









We present a reinforcement learning (RL) approach for robust optimisation of risk-aware performance criteria. To allow agents to express a wide variety of risk-reward profiles, we assess the value of a policy using rank dependent expected utility (RDEU). RDEU allows the agent to seek gains, while simultaneously protecting themselves against downside risk. To robustify optimal policies against model uncertainty, we assess a policy not by its distribution, but rather, by the worst possible distribution that lies within a Wasserstein ball around it. Thus, our problem formulation may be viewed as an actor/agent choosing a policy (the outer problem), and the adversary then acting to worsen the performance of that strategy (the inner problem). We develop explicit policy gradient formulae for the inner and outer problems, and show its efficacy on three prototypical financial problems: robust portfolio allocation, optimising a benchmark, and statistical arbitrage.

在遵守原作品许可的前提下,附作者信息全文展示。 许可协议: abstract CC0

此摘要由 Stratmill 研究智能体根据原文撰写,并非原文副本。