투기 매매와 페어 트레이딩을 위한 탐색적 강화학습
기사 arXiv papers · 저자: Yun Zhao et al.
요약
이 연구는 일반적인 효용 함수와 가격 과정 아래 진입 시점과 청산 시점을 선택하는 투기 매매를 순차적 최적 정지 문제로 구성합니다. 먼저 진입과 청산 사건을 콕스 과정의 점프로 나타내 정지 문제를 완화하며, 에이전트는 유계 강도 제어값을 선택합니다.
탐색적 강화학습 구성에서는 무작위 정책이 이러한 강도에 확률 분포를 부여하고, 섀넌 미분 엔트로피로 목적 함수를 정규화합니다. 그 결과 도출된 탐색적 해밀턴-야코비-벨만 방정식은 깁스 형태의 최적 정책을 냅니다. 저자들은 오차 추정과 학습 목적의 원래 문제 가치 함수에 대한 수렴을 입증한 뒤 페어 트레이딩 응용에서 알고리즘을 시연합니다. 제공된 설명에는 해당 응용의 데이터나 성과 결과가 없으며, 명시된 수렴은 모델링된 목적에 관한 것이지 입증된 실거래 결과에 관한 것이 아닙니다.
핵심 아이디어
- 투기 매매를 진입 및 청산 시점에 관한 최적 정지 문제로 모델링합니다.
- 완화된 구성에서는 유계 강도를 통해 제어되는 콕스 과정 점프로 정지 사건을 나타냅니다.
- 섀넌 미분 엔트로피로 무작위 강도 정책을 정규화합니다.
- 이 틀은 탐색적 HJB 방정식과 깁스 형태의 최적 정책을 도출합니다.
- 오차 및 수렴 결과를 입증하고 페어 트레이딩 알고리즘을 시연합니다.
태그
전문
# Reinforcement Learning for Speculative Trading under Exploratory Framework # Reinforcement Learning for Speculative Trading under Exploratory Framework We study a speculative trading problem within the exploratory reinforcement learning (RL) framework of Wang et al. [2020]. The problem is formulated as a sequential optimal stopping problem over entry and exit times under general utility function and price process. We first consider a relaxed version of the problem in which the stopping times are modeled by the jump times of Cox processes driven by bounded, non-randomized intensity controls. Under the exploratory formulation, the agent's randomized control is characterized via the probability measure over the jump intensities, and their objective function is regularized by Shannon's differential entropy. This yields a system of the exploratory HJB equations and Gibbs distributions in closed-form as the optimal policy. Error estimates and convergence of the RL objective to the value function of the original problem are established. Finally, an RL algorithm is designed, and its implementation is showcased in a pairs-trading application.
출처의 라이선스에 따라 출처를 표시하고 전문을 공개합니다. 라이선스: abstract CC0
이 요약은 원문을 바탕으로 Stratmill의 리서치 에이전트가 작성했으며, 원문을 복사한 것이 아닙니다.