用于投资组合优化的深度确定性策略梯度
文章 arXiv papers · 作者: Ayman Chaouki et al.
总结
本研究检验深度强化学习能否在易于描述但数学上具有挑战性的环境中求得最优交易策略。作者重点研究深度确定性策略梯度,这是一种通过与环境交互来学习行动选择策略的方法。他们选择了最优策略已知或近似已知的场景,因此可以将学习到的行为和奖励与参考解进行比较。
报告结果显示,智能体能够找回已知策略的关键特征,并获得接近最优的奖励。这为该算法能够在所测试环境中充当求解器提供了证据。摘录没有指出具体市场模型、投资组合约束、训练流程或评估指标。因此,结论仅适用于所选的受控场景,不能证明该方法在嘈杂的实盘市场、存在交易成本的情况下,或最优策略未知时的表现。
核心观点
- 本研究评估深度强化学习能否作为交易策略的求解方法。
- 研究在概念简单但数学上具有挑战性的环境中使用深度确定性策略梯度。
- 测试环境中的最优策略已知或近似已知,可用于比较。
- 据报告,智能体找回了策略的关键特征,并取得接近最优的奖励。
- 摘录不能证明该方法在实盘市场摩擦或最优策略未知时的表现。
标签
全文
# Deep Deterministic Portfolio Optimization # Deep Deterministic Portfolio Optimization Can deep reinforcement learning algorithms be exploited as solvers for optimal trading strategies? The aim of this work is to test reinforcement learning algorithms on conceptually simple, but mathematically non-trivial, trading environments. The environments are chosen such that an optimal or close-to-optimal trading strategy is known. We study the deep deterministic policy gradient algorithm and show that such a reinforcement learning agent can successfully recover the essential features of the optimal trading strategies and achieve close-to-optimal rewards.
在遵守原作品许可的前提下,附作者信息全文展示。 许可协议: abstract CC0
此摘要由 Stratmill 研究智能体根据原文撰写,并非原文副本。