본문으로 건너뛰기
라이브러리 문서 전체

잠재 레짐을 고려한 순환 강화학습 트레이딩

기사 arXiv papers · 저자: Andrea Macrì et al.

요약

이 논문은 신호가 평균회귀하는 오른슈타인-울렌벡 과정에 따르고, 그 모수가 잠재 레짐에 따라 달라질 때의 최적 트레이딩을 연구합니다. 순환 신경망과 강화학습을 결합해 관측값으로부터 신호의 잠재 평균, 회귀 속도, 변동성에 관한 정보를 추론한 뒤, 그 정보를 매매 지침으로 활용합니다.

저자들은 DDPG 기반 접근법 세 가지를 제안합니다. 하나는 순환 신경망의 은닉 상태를 트레이딩 에이전트에 전달하고, 나머지 둘은 추정된 레짐 확률이나 다음 신호값 예측을 제공합니다. 레짐 역학이 점점 복잡해지는 시뮬레이션과 주식 페어 트레이딩 실증 적용에서는 확률 기반 접근법이 누적 보상과 해석 가능성 면에서 더 나은 결과를 보였습니다. 예측 기반 방법의 추가 이점은 제한적이었고, 은닉 상태 방법은 두 방식의 중간 성과를 보였습니다. 이 결과는 추론한 정보를 표현하는 방식이 중요함을 시사하지만, 근거는 명시된 시뮬레이션과 적용 사례에 한정되며 다른 시장이나 실제 운용 환경에서의 성과를 입증하지는 않습니다.

핵심 아이디어

  • 트레이딩 신호는 숨겨진 레짐에 따라 모수가 달라지는 평균회귀 과정으로 모델링됩니다.
  • 순환 신경망은 신호 관측값에서 시간적 정보를 추출합니다.
  • 세 가지 DDPG 설계는 에이전트에 은닉 상태, 레짐 확률 또는 신호 예측값을 제공합니다.
  • 보고된 시뮬레이션과 페어 트레이딩 적용 사례에서는 레짐 확률 접근법이 가장 좋은 성과를 보입니다.
  • 에이전트에 제공하는 정보에 따라 보고된 보상과 전략의 해석 가능성이 달라집니다.

태그

전문
# Deep reinforcement learning for optimal trading with partial information


# Deep reinforcement learning for optimal trading with partial information









Reinforcement Learning (RL) applied to financial problems has been the subject of a lively area of research. The use of RL for optimal trading strategies that exploit latent information in the market is, to the best of our knowledge, not widely tackled. In this paper we study an optimal trading problem, where a trading signal follows an Ornstein-Uhlenbeck process with regime-switching dynamics. We employ a blend of RL and Recurrent Neural Networks (RNN) in order to make the most at extracting underlying information from the trading signal with latent parameters. The latent parameters driving mean reversion, speed, and volatility are filtered from observations of the signal, and trading strategies are derived via RL. To address this problem, we propose three Deep Deterministic Policy Gradient (DDPG)-based algorithms that integrate Gated Recurrent Unit (GRU) networks to capture temporal dependencies in the signal. The first, a one -step approach (hid-DDPG), directly encodes hidden states from the GRU into the RL trader. The second and third are two-step methods: one (prob-DDPG) makes use of posterior regime probability estimates, while the other (reg-DDPG) relies on forecasts of the next signal value. Through extensive simulations with increasingly complex Markovian regime dynamics for the trading signal's parameters, as well as an empirical application to equity pair trading, we find that prob-DDPG achieves superior cumulative rewards and exhibits more interpretable strategies. By contrast, reg-DDPG provides limited benefits, while hid-DDPG offers intermediate performance with less interpretable strategies. Our results show that the quality and structure of the information supplied to the agent are crucial: embedding probabilistic insights into latent regimes substantially improves both profitability and robustness of reinforcement learning-based trading strategies.

출처의 라이선스에 따라 출처를 표시하고 전문을 공개합니다. 라이선스: abstract CC0

이 요약은 원문을 바탕으로 Stratmill의 리서치 에이전트가 작성했으며, 원문을 복사한 것이 아닙니다.