コンテンツへスキップ
ライブラリの全資料

PPO取引と市場局面に応じたドローダウン管理

記事 arXiv papers · 著者: Duong Hien Chi Kien et al.

サマリー

PPO-HRAPは、近接方策最適化と市場局面から導いた目標エクスポージャーを組み合わせます。方策は市場の特徴とポートフォリオの状態を観測し、学習した行動を市場局面の目標と組み合わせます。報酬にはポートフォリオの対数リターン、VIXに応じたドローダウンの増加、目標エクスポージャーからの乖離、売買回転コストを反映します。この設計は、利益重視の方策が投資比率を高く保ち続ける傾向と、強いリスクペナルティが方策を過度に慎重にする傾向に対処し、利益への参加とドローダウン抑制の両立を目指します。

2020から2022までの未使用SPY期間では、バイ・アンド・ホールドより最大ドローダウンが低かったことに加え、リターンとリスク調整後パフォーマンス指標も報告されています。SPYの5つのシードで結果は安定しているとされ、QQQとDIAの単一実行テストでも、比較対象の中で総リターンとシャープレシオが最上位となっています。著者らは売買回転率の高さと、異なる資産での証拠の少なさを認めており、より広い頑健性は不確かです。

主なアイデア

  • PPO-HRAPは、学習した方策の行動を市場局面から導いた目標エクスポージャーと組み合わせます。
  • 報酬にはリターン、VIXに応じたドローダウンの増加、目標エクスポージャーからの乖離、売買回転を反映します。
  • 未使用のSPY評価では、バイ・アンド・ホールドより最大ドローダウンが低いと報告されています。
  • SPYの5つのシードで結果は安定していますが、QQQとDIAの証拠は単一実行によるものです。
  • この手法は依然として売買回転率が高く、異なる資産での頑健性を示す証拠も限られています。

タグ

全文
# PPO-HRAP: Proximal Policy Optimization with a Hybrid Regime-Aware Policy for Risk-Controlled Trading


# PPO-HRAP: Proximal Policy Optimization with a Hybrid Regime-Aware Policy for Risk-Controlled Trading









Reinforcement learning for trading often struggles to balance upside participation with drawdown control. Profit-only policies can collapse toward passive long exposure on upward-drifting assets, while aggressively risk-penalized rewards can become too defensive during volatile periods. This paper proposes PPO-HRAP, a hybrid regime-aware policy that combines Proximal Policy Optimization with an interpretable regime prior. The agent observes both market features and portfolio-state variables, receives a reward combining portfolio log return, VIX-conditioned drawdown-increase penalty, target-exposure deviation, and turnover cost, and executes a blended action between the PPO actor output and a regime-derived target exposure. On the held-out 2020-2022 SPY test window, PPO-HRAP achieves 27.62% total return, 8.48% annualized return, 0.6447 Sharpe ratio, 0.8588 Sortino ratio, and 0.4592 Calmar ratio, while reducing maximum drawdown from 34.10% for Buy and Hold to 18.47%. Across five SPY seeds, PPO-HRAP remains stable with mean total return $0.2725 \pm 0.0109$ and mean Sharpe ratio $0.6219 \pm 0.0565$. Single-run cross-asset tests on QQQ and DIA further show that the proposed method ranks first on total return and Sharpe ratio for all three reported assets. These results suggest that blending learned actions with a volatility-aware regime prior is a practical way to improve risk-adjusted trading behavior, although the current policy still incurs high turnover and cross-asset robustness beyond SPY remains limited to single-run evidence.

出典を明記したうえで、ライセンスに従って全文を掲載しています。 ライセンス: abstract CC0

この要約は原文をもとにStratmillのリサーチエージェントが作成したもので、出典の複製ではありません。