외환 강화학습에서 DQN와 PPO 비교
기사 arXiv papers · 저자: Yun-Cheng Tsai et al.
요약
이 연구는 외환 트레이딩에 심층 강화학습을 적용하며, 직접 가격을 예측하는 대신 매매 선택을 연속된 의사결정으로 다룹니다. Sure-Fire 통계적 차익거래 정책을 세 가지 가능한 행동으로 조정하고, 연속적인 가격 이력을 Gramian Angular Field 이미지로 변환합니다. 저자들은 EUR/USD, GBP/USD, AUD/USD의 4시간 간격 데이터를 사용해 Deep Q Learning과 Proximal Policy Optimization을 비교합니다.
학습에는 1 8월부터 30 11월 2018까지의 데이터를 사용하고, 테스트 기간은 2018 12월입니다. 저자들은 복잡하고 무작위적인 시장 행동과 트레이딩 환경을 적절히 설명하는 상태를 표현할 수 있는 모델이 양호한 투자 성과를 보였다고 보고합니다. 제공된 설명에는 수익률 수치, 위험 지표, 수수료 가정, 상세한 모델 설정이 없어 결과의 강도와 비교 가능성을 평가할 수 없습니다. 근거는 짧은 과거 데이터 분할에서 세 통화 페어를 대상으로 한 가능성 검증이며, 어느 알고리즘이든 다른 기간이나 실거래 환경에서 안정적으로 성과를 낼 것임을 입증하지는 않습니다.
핵심 아이디어
- 외환 트레이딩을 강화학습을 활용한 순차적 의사결정 문제로 다룹니다.
- 조정된 통계적 차익거래 정책 안에서 세 가지 매매 행동을 정의합니다.
- 모델 입력을 위해 가격 구간을 Gramian Angular Field 이미지로 인코딩합니다.
- 세 통화 페어에서 Deep Q Learning과 Proximal Policy Optimization을 비교합니다.
- 짧은 과거 데이터 평가에서 양호한 성과를 보고하지만, 상세 지표와 거래 비용 가정은 제시하지 않습니다.
태그
전문
# Deep Reinforcement Learning for Foreign Exchange Trading # Deep Reinforcement Learning for Foreign Exchange Trading Reinforcement learning can interact with the environment and is suitable for applications in decision control systems. Therefore, we used the reinforcement learning method to establish a foreign exchange transaction, avoiding the long-standing problem of unstable trends in deep learning predictions. In the system design, we optimized the Sure-Fire statistical arbitrage policy, set three different actions, encoded the continuous price over a period of time into a heat-map view of the Gramian Angular Field (GAF) and compared the Deep Q Learning (DQN) and Proximal Policy Optimization (PPO) algorithms. To test feasibility, we analyzed three currency pairs, namely EUR/USD, GBP/USD, and AUD/USD. We trained the data in units of four hours from 1 August 2018 to 30 November 2018 and tested model performance using data between 1 December 2018 and 31 December 2018. The test results of the various models indicated that favorable investment performance was achieved as long as the model was able to handle complex and random processes and the state was able to describe the environment, validating the feasibility of reinforcement learning in the development of trading strategies.
출처의 라이선스에 따라 출처를 표시하고 전문을 공개합니다. 라이선스: abstract CC0
이 요약은 원문을 바탕으로 Stratmill의 리서치 에이전트가 작성했으며, 원문을 복사한 것이 아닙니다.