주식 트레이딩을 위한 심층 강화학습 에이전트 앙상블
기사 arXiv papers · 저자: Hongyang Yang et al.
요약
이 논문은 근접 정책 최적화, 어드밴티지 액터 크리틱, 심층 결정론적 정책 경사라는 세 가지 액터 크리틱 강화학습 알고리즘을 결합한 자동 주식 트레이딩 방식을 설명합니다. 에이전트는 투자 수익 극대화를 목표로 트레이딩 결정을 학습하고, 각 알고리즘의 강점을 결합해 시장 상황에 적응하도록 만든 앙상블에 통합됩니다. 연속 행동 학습 중 메모리 사용을 제한하면서 대용량 데이터를 처리하기 위해 필요할 때 데이터를 불러오는 기법도 사용합니다.
이 접근법은 다우존스 종목군의 유동성 높은 주식으로 평가하고, 개별 학습 알고리즘, 다우존스 산업평균지수, 최소분산 포트폴리오와 비교합니다. 앙상블이 개별 방법과 기준보다 높은 샤프 비율을 달성했다고 보고합니다. 발췌문에는 평가 기간, 구체적인 포트폴리오 제약, 거래 비용 가정 또는 강건성 검정이 없습니다. 따라서 성과 주장은 미래 수익의 근거가 아니라 연구에서 명시한 시험 설정 안에서 이해해야 합니다.
핵심 아이디어
- 전략은 PPO, A2C, DDPG 액터 크리틱 에이전트를 앙상블로 결합합니다.
- 에이전트는 투자 수익을 목표로 주식 트레이딩 결정을 학습합니다.
- 학습 중 메모리 부담을 줄이기 위해 필요할 때 데이터를 불러오는 처리를 사용합니다.
- 평가에서 앙상블을 구성 알고리즘, 시장 지수, 최소분산 포트폴리오와 비교합니다.
- 연구 설정을 전제로, 비교 대상 중 앙상블의 샤프 비율이 가장 높았다고 보고합니다.
태그
전문
# Deep Reinforcement Learning for Automated Stock Trading: An Ensemble Strategy
# Deep Reinforcement Learning for Automated Stock Trading: An Ensemble Strategy
Stock trading strategies play a critical role in investment. However, it is challenging to design a profitable strategy in a complex and dynamic stock market. In this paper, we propose an ensemble strategy that employs deep reinforcement schemes to learn a stock trading strategy by maximizing investment return. We train a deep reinforcement learning agent and obtain an ensemble trading strategy using three actor-critic based algorithms: Proximal Policy Optimization (PPO), Advantage Actor Critic (A2C), and Deep Deterministic Policy Gradient (DDPG). The ensemble strategy inherits and integrates the best features of the three algorithms, thereby robustly adjusting to different market situations. In order to avoid the large memory consumption in training networks with continuous action space, we employ a load-on-demand technique for processing very large data. We test our algorithms on the 30 Dow Jones stocks that have adequate liquidity. The performance of the trading agent with different reinforcement learning algorithms is evaluated and compared with both the Dow Jones Industrial Average index and the traditional min-variance portfolio allocation strategy. The proposed deep ensemble strategy is shown to outperform the three individual algorithms and two baselines in terms of the risk-adjusted return measured by the Sharpe ratio. This work is fully open-sourced at \href{https://github.com/AI4Finance-Foundation/Deep-Reinforcement-Learning-for-Automated-Stock-Trading-Ensemble-Strategy-ICAIF-2020}{GitHub}.출처의 라이선스에 따라 출처를 표시하고 전문을 공개합니다. 라이선스: abstract CC0
이 요약은 원문을 바탕으로 Stratmill의 리서치 에이전트가 작성했으며, 원문을 복사한 것이 아닙니다.