Skip to content
All library documents

Ensembling Deep Reinforcement Learning Agents for Stock Trading

Article arXiv papers · Author: Hongyang Yang et al.

Summary

The paper describes an automated stock-trading approach that combines three actor-critic reinforcement learning algorithms: Proximal Policy Optimization, Advantage Actor Critic, and Deep Deterministic Policy Gradient. The agents are trained to learn trading decisions with the aim of maximizing investment returns, and their strengths are integrated into an ensemble intended to adapt across market conditions. A load-on-demand technique is also used to process large data while limiting memory use during training with continuous actions.

The approach is evaluated on liquid stocks from the Dow Jones universe and compared with the individual learning algorithms, the Dow Jones Industrial Average, and a minimum-variance portfolio. The reported result is that the ensemble achieves a higher Sharpe ratio than those individual methods and baselines. The excerpt does not give the evaluation period, detailed portfolio constraints, transaction-cost assumptions, or robustness checks. Its performance claim should therefore be understood within the study's stated test setup rather than as evidence of future returns.

Key ideas

  • The strategy combines PPO, A2C, and DDPG actor-critic agents into an ensemble.
  • The agents learn stock-trading decisions with investment return as the objective.
  • Load-on-demand data processing is used to reduce memory demands during training.
  • The evaluation compares the ensemble with its constituent methods, a market index, and a minimum-variance portfolio.
  • The ensemble is reported to have the strongest Sharpe ratio in those comparisons, subject to the study's test setup.

Tags

Full text
# Deep Reinforcement Learning for Automated Stock Trading: An Ensemble Strategy


# Deep Reinforcement Learning for Automated Stock Trading: An Ensemble Strategy









Stock trading strategies play a critical role in investment. However, it is challenging to design a profitable strategy in a complex and dynamic stock market. In this paper, we propose an ensemble strategy that employs deep reinforcement schemes to learn a stock trading strategy by maximizing investment return. We train a deep reinforcement learning agent and obtain an ensemble trading strategy using three actor-critic based algorithms: Proximal Policy Optimization (PPO), Advantage Actor Critic (A2C), and Deep Deterministic Policy Gradient (DDPG). The ensemble strategy inherits and integrates the best features of the three algorithms, thereby robustly adjusting to different market situations. In order to avoid the large memory consumption in training networks with continuous action space, we employ a load-on-demand technique for processing very large data. We test our algorithms on the 30 Dow Jones stocks that have adequate liquidity. The performance of the trading agent with different reinforcement learning algorithms is evaluated and compared with both the Dow Jones Industrial Average index and the traditional min-variance portfolio allocation strategy. The proposed deep ensemble strategy is shown to outperform the three individual algorithms and two baselines in terms of the risk-adjusted return measured by the Sharpe ratio. This work is fully open-sourced at \href{https://github.com/AI4Finance-Foundation/Deep-Reinforcement-Learning-for-Automated-Stock-Trading-Ensemble-Strategy-ICAIF-2020}{GitHub}.

Shown in full with attribution under the source's licence. Licence: abstract CC0

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.