암호화폐 페어 트레이딩을 위한 심층 강화학습
기사 arXiv papers · 저자: Damian Lebiedź et al.
요약
이 연구는 변동성이 큰 암호화폐 시장에서 페어 트레이딩의 체결 보조 기법으로 심층 강화학습을 테스트합니다. 프레임워크는 먼저 후보 페어를 추리고 순위를 매긴 뒤, Proximal Policy Optimization 에이전트와 Long Short-Term Memory 계층을 사용해 체결 결정을 내립니다. 고정 위험과 적응형 평균 설계, 결정론적 위험 통제로 에이전트를 제한해 괴리 위험을 줄이고자 합니다. 평가는 시간별 Binance USD-M 선물 데이터를 사용하고 학습된 정책을 휴리스틱 기준선과 비교합니다.
보고된 표본 외 결과는 강화학습 정책에 유리하며, 정상 원형 블록 부트스트랩에서는 위험조정 초과성과가 10퍼센트 수준에서 통계적으로 유의한 것으로 나타났습니다. 이 결과는 더 엄격한 5퍼센트 유의수준에는 미치지 않으며, 연구는 이를 디지털 자산의 높은 개별 요인 분산과 연관 짓습니다. 증거는 테스트한 데이터와 설정에 한정됩니다. 설명에는 성과 규모, 거래 비용 가정, 다른 거래소나 기간에서의 테스트가 제시되지 않습니다. 하이브리드 설계는 학습된 체결 결정을 제한하는 방법을 제시하지만, 더 폭넓은 강건성은 여전히 확인되지 않았습니다.
핵심 아이디어
- 통계적 페어 선정과 학습 기반 체결 보조 기법을 결합한 시스템을 제안합니다.
- PPO 에이전트와 LSTM를 사용해 결정론적 위험 한도 안에서 체결 결정을 내립니다.
- 시간별 Binance USD-M 선물 데이터로 휴리스틱 기준선과 접근법을 비교합니다.
- 보고된 위험조정 초과성과는 10퍼센트 수준에서는 유의하지만 5퍼센트 수준에서는 유의하지 않습니다.
- 결과는 테스트한 시장, 데이터 및 전략 설정에 한정됩니다.
태그
전문
# Dynamic Multi-Pair Trading Strategy in Cryptocurrency Markets with Deep Reinforcement Learning # Dynamic Multi-Pair Trading Strategy in Cryptocurrency Markets with Deep Reinforcement Learning This study aims to determine whether the application of Deep Reinforcement Learning (DRL) as a specialized execution overlay can enhance pair trading in highly volatile cryptocurrency markets. Although classical implementations of the strategy have proven successful in traditional equities, they frequently exhibit rigidity and suffer from severe divergence risks when applied to high-variance environments. To address this need, this research introduces novel concepts. To construct a robust system, we developed a hierarchical "Filter-then-Rank" pair selection methodology and a proprietary "Fixed Risk, Adaptive Mean" execution model. The system employs a Proximal Policy Optimization (PPO) agent with a Long Short-Term Memory (LSTM) layer to govern execution decisions within strict deterministic risk management boundaries. Evaluated on 1-hour interval data from the Binance USD-M Futures market, the optimized RL policy achieved an out-of-sample performance that substantially outperformed the heuristic baseline. A stationary circular block bootstrap robustness check confirms that the agent's risk-adjusted outperformance is statistically significant at the 10 percent level. Although falling marginally short of the stricter 5 percent threshold, this result highlights the extreme idiosyncratic variance characteristic of digital assets. Ultimately, this thesis contributes to the quantitative finance literature by introducing a hybrid architecture that combines statistical arbitrage with DRL execution policies. Furthermore, it delivers a novel framework for safe reinforcement learning via deterministic shielding, proving that anchoring a neural policy to statistically robust boundaries successfully mitigates severe divergence risks.
출처의 라이선스에 따라 출처를 표시하고 전문을 공개합니다. 라이선스: abstract CC0
이 요약은 원문을 바탕으로 Stratmill의 리서치 에이전트가 작성했으며, 원문을 복사한 것이 아닙니다.