본문으로 건너뛰기
라이브러리 문서 전체

페어 트레이딩을 위한 위험 인식 순환형 강화학습

기사 arXiv papers · 저자: Weiguang Han et al.

요약

이 문서는 CREDIT을 설명합니다. 양방향 게이트 순환 유닛과 시간 어텐션을 사용하는 페어 트레이딩 강화학습 접근법입니다. 이 구조는 두 자산 가격 움직임의 시간에 따른 관계를 포착하도록 설계됐으며, 시장 상태를 개별적으로 볼 때 놓칠 수 있는 패턴도 다룹니다. 또한 거래 수익과 위험을 모두 반영하는 보상을 사용해 잠재적인 이익과 손실이 큰 거래를 억제하려 합니다.

저자들은 이 방법이 기존 강화학습 접근법의 잦은 거래, 거래 비용, 위험 감수 문제에 대응한다고 설명합니다. 미국 주식 데이터 5년을 사용한 실험에서 CREDIT이 기존 강화학습 방법보다 뛰어난 성능을 보이고 상당한 수익을 냈다고 보고합니다. 발췌문에는 벤치마크 세부 정보, 비용 가정, 위험 측정치 또는 수치 결과가 없어 보고된 우위의 크기나 견고성을 평가하기 어렵습니다. 다른 시장이나 기간에서의 성과도 여기서는 입증되지 않았습니다.

핵심 아이디어

  • CREDIT은 두 자산 페어 트레이딩의 가격 이력에 순환형 학습과 시간 어텐션을 적용합니다.
  • 보상은 수익과 거래 위험을 모두 반영합니다.
  • 이 설계는 장기 가격 패턴을 학습하고 과도한 거래를 줄이는 것을 목표로 합니다.
  • 미국 주식 데이터 5년을 사용한 실험에서 다른 강화학습 방법보다 뛰어났다고 보고합니다.
  • 발췌문만으로는 견고성이나 실거래 비용을 평가할 만큼 충분한 세부 정보를 확인할 수 없습니다.

태그

전문
# Mastering Pair Trading with Risk-Aware Recurrent Reinforcement Learning


# Mastering Pair Trading with Risk-Aware Recurrent Reinforcement Learning









Although pair trading is the simplest hedging strategy for an investor to eliminate market risk, it is still a great challenge for reinforcement learning (RL) methods to perform pair trading as human expertise. It requires RL methods to make thousands of correct actions that nevertheless have no obvious relations to the overall trading profit, and to reason over infinite states of the time-varying market most of which have never appeared in history. However, existing RL methods ignore the temporal connections between asset price movements and the risk of the performed trading. These lead to frequent tradings with high transaction costs and potential losses, which barely reach the human expertise level of trading. Therefore, we introduce CREDIT, a risk-aware agent capable of learning to exploit long-term trading opportunities in pair trading similar to a human expert. CREDIT is the first to apply bidirectional GRU along with the temporal attention mechanism to fully consider the temporal correlations embedded in the states, which allows CREDIT to capture long-term patterns of the price movements of two assets to earn higher profit. We also design the risk-aware reward inspired by the economic theory, that models both the profit and risk of the tradings during the trading period. It helps our agent to master pair trading with a robust trading preference that avoids risky trading with possible high returns and losses. Experiments show that it outperforms existing reinforcement learning methods in pair trading and achieves a significant profit over five years of U.S. stock data.

출처의 라이선스에 따라 출처를 표시하고 전문을 공개합니다. 라이선스: abstract CC0

이 요약은 원문을 바탕으로 Stratmill의 리서치 에이전트가 작성했으며, 원문을 복사한 것이 아닙니다.