比较 DQN 与 PPO 在外汇强化学习交易中的表现
文章 arXiv papers · 作者: Yun-Cheng Tsai et al.
总结
这项研究将深度强化学习用于外汇交易,把交易选择建模为一系列决策,而非依赖直接的价格预测。研究将 Sure-Fire 统计套利策略调整为三种可能的操作,并把连续价格历史转换为格拉米角场图像。作者使用 EUR/USD、GBP/USD 和 AUD/USD 的四小时级数据,对比深度 Q 学习与近端策略优化。
训练使用 1 August 至 30 November 2018 的数据,测试期为 December 2018。作者报告称,能够表示复杂随机市场行为、且状态能充分描述交易环境的模型取得了较好的投资表现。所提供的描述没有给出收益数据、风险指标、费用假设或详细的模型设置,因此无法在此评估结果的强弱和可比性。证据来自三个货币对在较短历史区间上的可行性测试,并不能证明任一算法在其他时期或实盘执行条件下都能可靠运行。
核心观点
- 该方法通过强化学习将外汇交易视为序贯决策问题。
- 调整后的统计套利策略中定义了三种交易操作。
- 价格窗口被编码为格拉米角场图像,作为模型输入。
- 研究在三个货币对上比较了深度 Q 学习与近端策略优化。
- 较短的历史评估报告了较好的表现,但未提供详细指标和交易成本假设。
标签
全文
# Deep Reinforcement Learning for Foreign Exchange Trading # Deep Reinforcement Learning for Foreign Exchange Trading Reinforcement learning can interact with the environment and is suitable for applications in decision control systems. Therefore, we used the reinforcement learning method to establish a foreign exchange transaction, avoiding the long-standing problem of unstable trends in deep learning predictions. In the system design, we optimized the Sure-Fire statistical arbitrage policy, set three different actions, encoded the continuous price over a period of time into a heat-map view of the Gramian Angular Field (GAF) and compared the Deep Q Learning (DQN) and Proximal Policy Optimization (PPO) algorithms. To test feasibility, we analyzed three currency pairs, namely EUR/USD, GBP/USD, and AUD/USD. We trained the data in units of four hours from 1 August 2018 to 30 November 2018 and tested model performance using data between 1 December 2018 and 31 December 2018. The test results of the various models indicated that favorable investment performance was achieved as long as the model was able to handle complex and random processes and the state was able to describe the environment, validating the feasibility of reinforcement learning in the development of trading strategies.
在遵守原作品许可的前提下,附作者信息全文展示。 许可协议: abstract CC0
此摘要由 Stratmill 研究智能体根据原文撰写,并非原文副本。