株式売買における深層強化学習エージェントのアンサンブル
記事 arXiv papers · 著者: Hongyang Yang et al.
サマリー
この論文は、近接方策最適化、Advantage Actor Critic、Deep Deterministic Policy Gradientという3つのアクター・クリティック型強化学習アルゴリズムを組み合わせた自動株式売買手法を説明します。エージェントは投資リターンの最大化を目指して売買判断を学習し、それぞれの強みを市場環境の変化に適応することを意図したアンサンブルに統合します。また、連続行動の訓練時にメモリ使用量を抑えながら大規模データを処理するため、オンデマンド読み込み手法も用います。
この手法はダウ・ジョーンズ構成銘柄のうち流動性の高い株式で評価され、各学習アルゴリズム単体、ダウ・ジョーンズ工業株平均、最小分散ポートフォリオと比較されています。報告された結果では、アンサンブルは各手法とベースラインを上回るシャープレシオを達成しています。抜粋には評価期間、詳細なポートフォリオ制約、取引コストの仮定、頑健性検証が記載されていません。したがって、この成績は将来のリターンの証拠ではなく、研究で示された検証条件の範囲内で理解する必要があります。
主なアイデア
- この戦略は、PPO、A2C、DDPGのアクター・クリティック型エージェントをアンサンブルに統合します。
- エージェントは投資リターンを目的として、株式の売買判断を学習します。
- 訓練時のメモリ負荷を抑えるため、オンデマンドのデータ処理を使用します。
- アンサンブルを構成手法、株価指数、最小分散ポートフォリオと比較しています。
- 研究の検証条件の範囲で、アンサンブルは比較対象の中で最も高いシャープレシオを示すと報告されています。
タグ
全文
# Deep Reinforcement Learning for Automated Stock Trading: An Ensemble Strategy
# Deep Reinforcement Learning for Automated Stock Trading: An Ensemble Strategy
Stock trading strategies play a critical role in investment. However, it is challenging to design a profitable strategy in a complex and dynamic stock market. In this paper, we propose an ensemble strategy that employs deep reinforcement schemes to learn a stock trading strategy by maximizing investment return. We train a deep reinforcement learning agent and obtain an ensemble trading strategy using three actor-critic based algorithms: Proximal Policy Optimization (PPO), Advantage Actor Critic (A2C), and Deep Deterministic Policy Gradient (DDPG). The ensemble strategy inherits and integrates the best features of the three algorithms, thereby robustly adjusting to different market situations. In order to avoid the large memory consumption in training networks with continuous action space, we employ a load-on-demand technique for processing very large data. We test our algorithms on the 30 Dow Jones stocks that have adequate liquidity. The performance of the trading agent with different reinforcement learning algorithms is evaluated and compared with both the Dow Jones Industrial Average index and the traditional min-variance portfolio allocation strategy. The proposed deep ensemble strategy is shown to outperform the three individual algorithms and two baselines in terms of the risk-adjusted return measured by the Sharpe ratio. This work is fully open-sourced at \href{https://github.com/AI4Finance-Foundation/Deep-Reinforcement-Learning-for-Automated-Stock-Trading-Ensemble-Strategy-ICAIF-2020}{GitHub}.出典を明記したうえで、ライセンスに従って全文を掲載しています。 ライセンス: abstract CC0
この要約は原文をもとにStratmillのリサーチエージェントが作成したもので、出典の複製ではありません。