连续动作TD3用于股票与比特币交易
文章 arXiv papers · 作者: Naseh Majidi et al.
总结
这项研究将双延迟深度确定性策略梯度(TD3)这一连续动作强化学习方法应用于亚马逊股票和比特币的每日收盘价。交易动作同时表示持仓方向和交易股数,扩展了使用离散动作的方法。论文将由此产生的策略与技术分析方法、其他强化学习方法以及随机和确定性策略进行比较。
研究使用收益率和夏普比率评估表现。报告结果表明,同时选择持仓方向和交易规模可以改善这些指标。文档没有提供评估细节、具体结果,也没有证据说明该方法在其他资产或市场时期的稳健性。因此,其结论仅描述所研究的市场和比较方法,并未证明连续动作TD3在其他情境下会表现更好。
核心观点
- 研究使用TD3根据亚马逊股票和比特币的每日收盘价生成交易动作。
- 动作空间同时表示交易持仓和交易股数。
- 研究将TD3与技术分析、强化学习、随机及确定性策略进行比较。
- 收益率和夏普比率是文中列出的绩效指标。
- 在研究测试的情境中,报告的比较结果更支持将持仓选择与交易规模结合起来。
标签
全文
# Algorithmic Trading Using Continuous Action Space Deep Reinforcement Learning # Algorithmic Trading Using Continuous Action Space Deep Reinforcement Learning Price movement prediction has always been one of the traders' concerns in financial market trading. In order to increase their profit, they can analyze the historical data and predict the price movement. The large size of the data and complex relations between them lead us to use algorithmic trading and artificial intelligence. This paper aims to offer an approach using Twin-Delayed DDPG (TD3) and the daily close price in order to achieve a trading strategy in the stock and cryptocurrency markets. Unlike previous studies using a discrete action space reinforcement learning algorithm, the TD3 is continuous, offering both position and the number of trading shares. Both the stock (Amazon) and cryptocurrency (Bitcoin) markets are addressed in this research to evaluate the performance of the proposed algorithm. The achieved strategy using the TD3 is compared with some algorithms using technical analysis, reinforcement learning, stochastic, and deterministic strategies through two standard metrics, Return and Sharpe ratio. The results indicate that employing both position and the number of trading shares can improve the performance of a trading system based on the mentioned metrics.
在遵守原作品许可的前提下,附作者信息全文展示。 许可协议: abstract CC0
此摘要由 Stratmill 研究智能体根据原文撰写,并非原文副本。