跳至正文
返回文库全部文档

集成深度强化学习智能体进行股票交易

文章 arXiv papers · 作者: Hongyang Yang et al.

总结

论文介绍一种自动化股票交易方法,结合三种演员-评论家强化学习算法:近端策略优化、优势演员-评论家和深度确定性策略梯度。智能体经过训练来学习交易决策,目标是最大化投资回报;随后将各算法的优势整合进一个集成模型,旨在适应不同市场环境。该方法还使用按需加载技术处理大规模数据,以减少连续动作训练期间的内存占用。

该方法在道琼斯指数成分股中流动性较高的股票上进行评估,并与各个单独的学习算法、道琼斯工业平均指数和最小方差投资组合进行比较。据报告,集成模型的夏普比率高于这些单独方法和基准。摘录没有提供评估时段、详细投资组合约束、交易成本假设或稳健性检验。因此,其绩效主张应理解为仅适用于该研究所述的测试设置,而非未来收益的证据。

核心观点

  • 该策略将PPO、A2C和DDPG演员-评论家智能体集成为一个模型。
  • 智能体学习股票交易决策,目标是获得投资回报。
  • 训练时使用按需加载的数据处理方式,以减少内存需求。
  • 评估将集成模型与组成它的各个方法、市场指数和最小方差投资组合进行比较。
  • 据报告,在这些比较中,集成模型的夏普比率最高;结论受研究测试设置限制。

标签

全文
# Deep Reinforcement Learning for Automated Stock Trading: An Ensemble Strategy


# Deep Reinforcement Learning for Automated Stock Trading: An Ensemble Strategy









Stock trading strategies play a critical role in investment. However, it is challenging to design a profitable strategy in a complex and dynamic stock market. In this paper, we propose an ensemble strategy that employs deep reinforcement schemes to learn a stock trading strategy by maximizing investment return. We train a deep reinforcement learning agent and obtain an ensemble trading strategy using three actor-critic based algorithms: Proximal Policy Optimization (PPO), Advantage Actor Critic (A2C), and Deep Deterministic Policy Gradient (DDPG). The ensemble strategy inherits and integrates the best features of the three algorithms, thereby robustly adjusting to different market situations. In order to avoid the large memory consumption in training networks with continuous action space, we employ a load-on-demand technique for processing very large data. We test our algorithms on the 30 Dow Jones stocks that have adequate liquidity. The performance of the trading agent with different reinforcement learning algorithms is evaluated and compared with both the Dow Jones Industrial Average index and the traditional min-variance portfolio allocation strategy. The proposed deep ensemble strategy is shown to outperform the three individual algorithms and two baselines in terms of the risk-adjusted return measured by the Sharpe ratio. This work is fully open-sourced at \href{https://github.com/AI4Finance-Foundation/Deep-Reinforcement-Learning-for-Automated-Stock-Trading-Ensemble-Strategy-ICAIF-2020}{GitHub}.

在遵守原作品许可的前提下,附作者信息全文展示。 许可协议: abstract CC0

此摘要由 Stratmill 研究智能体根据原文撰写,并非原文副本。