드로다운 관리를 위한 시장 국면 인식형 노출 사전분포를 적용한 PPO 트레이딩
기사 arXiv papers · 저자: Duong Hien Chi Kien et al.
요약
PPO-HRAP는 근접 정책 최적화와 시장 국면에 따라 정한 목표 노출을 결합합니다. 정책은 시장 특성과 포트폴리오 상태를 관찰한 뒤 학습된 행동을 국면별 목표와 혼합합니다. 보상은 포트폴리오 로그 수익률, VIX에 따라 조정한 드로다운 증가분, 목표 노출과의 편차, 회전율 비용을 반영합니다. 수익에 집중하는 정책이 과도한 투자를 유지하고, 강한 위험 페널티가 정책을 지나치게 보수적으로 만들 수 있는 문제를 고려해 수익 참여와 드로다운 관리를 균형 있게 추구합니다.
SPY의 별도 검증 구간에서 2020년부터 2022년까지 연구를 진행한 결과, 매수 후 보유보다 최대 낙폭이 낮았으며 수익률 및 위험조정 성과 지표도 함께 보고했습니다. SPY개 시드에 걸친 결과는 안정적이라고 설명하며, QQQ와 DIA의 단일 실행 테스트에서도 비교 대상 방법 중 총수익률과 샤프 비율이 가장 높았습니다. 저자들은 높은 회전율과 제한적인 자산 간 증거를 인정하므로, 더 폭넓은 견고성은 여전히 불확실합니다.
핵심 아이디어
- PPO-HRAP는 학습된 정책 행동을 시장 국면에서 도출한 목표 노출과 혼합합니다.
- 보상은 수익률, VIX 조건부 드로다운 증가분, 목표 노출 편차, 회전율을 반영합니다.
- 홀드아웃 SPY 평가에서 매수 후 보유보다 최대 드로다운이 낮았다고 보고합니다.
- SPY 시드 5개에서 결과는 안정적이지만, QQQ와 DIA 근거는 단일 실행에서 나왔습니다.
- 이 방법은 여전히 회전율이 높고 자산 간 견고성 근거가 제한적입니다.
태그
전문
# PPO-HRAP: Proximal Policy Optimization with a Hybrid Regime-Aware Policy for Risk-Controlled Trading # PPO-HRAP: Proximal Policy Optimization with a Hybrid Regime-Aware Policy for Risk-Controlled Trading Reinforcement learning for trading often struggles to balance upside participation with drawdown control. Profit-only policies can collapse toward passive long exposure on upward-drifting assets, while aggressively risk-penalized rewards can become too defensive during volatile periods. This paper proposes PPO-HRAP, a hybrid regime-aware policy that combines Proximal Policy Optimization with an interpretable regime prior. The agent observes both market features and portfolio-state variables, receives a reward combining portfolio log return, VIX-conditioned drawdown-increase penalty, target-exposure deviation, and turnover cost, and executes a blended action between the PPO actor output and a regime-derived target exposure. On the held-out 2020-2022 SPY test window, PPO-HRAP achieves 27.62% total return, 8.48% annualized return, 0.6447 Sharpe ratio, 0.8588 Sortino ratio, and 0.4592 Calmar ratio, while reducing maximum drawdown from 34.10% for Buy and Hold to 18.47%. Across five SPY seeds, PPO-HRAP remains stable with mean total return $0.2725 \pm 0.0109$ and mean Sharpe ratio $0.6219 \pm 0.0565$. Single-run cross-asset tests on QQQ and DIA further show that the proposed method ranks first on total return and Sharpe ratio for all three reported assets. These results suggest that blending learned actions with a volatility-aware regime prior is a practical way to improve risk-adjusted trading behavior, although the current policy still incurs high turnover and cross-asset robustness beyond SPY remains limited to single-run evidence.
출처의 라이선스에 따라 출처를 표시하고 전문을 공개합니다. 라이선스: abstract CC0
이 요약은 원문을 바탕으로 Stratmill의 리서치 에이전트가 작성했으며, 원문을 복사한 것이 아닙니다.