跳至正文
返回文库全部文档

用于配对交易的风险感知循环强化学习

文章 arXiv papers · 作者: Weiguang Han et al.

总结

本文介绍CREDIT,一种用于配对交易的强化学习方法,采用双向门控循环单元和时间注意力机制。该架构旨在捕捉两种资产价格变动随时间变化的关系,包括在单独考察市场状态时可能被忽略的模式。该方法还使用同时考虑交易利润和风险的奖励函数,旨在抑制潜在收益和损失都较高的交易。

作者将该方法定位为对早期强化学习方法中频繁交易、交易成本和风险承担问题的回应。他们报告称,在使用五年美国股票数据的实验中,CREDIT优于现有强化学习方法并获得显著利润。摘录没有提供基准详情、成本假设、风险度量或数值结果,因此无法评估所报告优势的大小或稳健性。这里也未证明其在其他市场或时期的表现。

核心观点

  • CREDIT将循环学习和时间注意力应用于双资产配对交易的价格历史。
  • 其奖励函数同时考虑利润和交易风险。
  • 该设计旨在学习较长期的价格模式并减少过度交易。
  • 据报告,在五年美国股票数据上的实验优于其他强化学习方法。
  • 摘录没有提供足够细节来评估稳健性或现实交易成本。

标签

全文
# Mastering Pair Trading with Risk-Aware Recurrent Reinforcement Learning


# Mastering Pair Trading with Risk-Aware Recurrent Reinforcement Learning









Although pair trading is the simplest hedging strategy for an investor to eliminate market risk, it is still a great challenge for reinforcement learning (RL) methods to perform pair trading as human expertise. It requires RL methods to make thousands of correct actions that nevertheless have no obvious relations to the overall trading profit, and to reason over infinite states of the time-varying market most of which have never appeared in history. However, existing RL methods ignore the temporal connections between asset price movements and the risk of the performed trading. These lead to frequent tradings with high transaction costs and potential losses, which barely reach the human expertise level of trading. Therefore, we introduce CREDIT, a risk-aware agent capable of learning to exploit long-term trading opportunities in pair trading similar to a human expert. CREDIT is the first to apply bidirectional GRU along with the temporal attention mechanism to fully consider the temporal correlations embedded in the states, which allows CREDIT to capture long-term patterns of the price movements of two assets to earn higher profit. We also design the risk-aware reward inspired by the economic theory, that models both the profit and risk of the tradings during the trading period. It helps our agent to master pair trading with a robust trading preference that avoids risky trading with possible high returns and losses. Experiments show that it outperforms existing reinforcement learning methods in pair trading and achieves a significant profit over five years of U.S. stock data.

在遵守原作品许可的前提下,附作者信息全文展示。 许可协议: abstract CC0

此摘要由 Stratmill 研究智能体根据原文撰写,并非原文副本。