Skip to content
All library documents

PPO Trading with a Regime-Aware Exposure Prior for Drawdown Control

Article arXiv papers · Author: Duong Hien Chi Kien et al.

Summary

PPO-HRAP combines Proximal Policy Optimization with a regime-derived target exposure. Its policy observes market features and portfolio state, then blends the learned action with the regime target. The reward accounts for portfolio log returns, increases in drawdown conditioned on VIX, deviation from target exposure, and turnover costs. The design aims to balance participation in gains with drawdown control, addressing the tendency of profit-focused policies to stay heavily invested and of strong risk penalties to make policies overly cautious.

On a held-out SPY window from 2020 to 2022, the study reports lower maximum drawdown than Buy and Hold alongside its return and risk-adjusted performance measures. Results across five SPY seeds are described as stable; single-run tests on QQQ and DIA also rank the method first on total return and Sharpe ratio among the compared methods. The authors acknowledge high turnover and limited cross-asset evidence, so broader robustness remains uncertain.

Key ideas

  • PPO-HRAP blends the learned policy action with a target exposure derived from market regimes.
  • Its reward accounts for returns, VIX-conditioned drawdown increases, target-exposure deviation, and turnover.
  • The held-out SPY evaluation reports lower maximum drawdown than Buy and Hold.
  • Results across five SPY seeds are stable, while QQQ and DIA evidence comes from single runs.
  • The method still has high turnover and limited cross-asset robustness evidence.

Tags

Full text
# PPO-HRAP: Proximal Policy Optimization with a Hybrid Regime-Aware Policy for Risk-Controlled Trading


# PPO-HRAP: Proximal Policy Optimization with a Hybrid Regime-Aware Policy for Risk-Controlled Trading









Reinforcement learning for trading often struggles to balance upside participation with drawdown control. Profit-only policies can collapse toward passive long exposure on upward-drifting assets, while aggressively risk-penalized rewards can become too defensive during volatile periods. This paper proposes PPO-HRAP, a hybrid regime-aware policy that combines Proximal Policy Optimization with an interpretable regime prior. The agent observes both market features and portfolio-state variables, receives a reward combining portfolio log return, VIX-conditioned drawdown-increase penalty, target-exposure deviation, and turnover cost, and executes a blended action between the PPO actor output and a regime-derived target exposure. On the held-out 2020-2022 SPY test window, PPO-HRAP achieves 27.62% total return, 8.48% annualized return, 0.6447 Sharpe ratio, 0.8588 Sortino ratio, and 0.4592 Calmar ratio, while reducing maximum drawdown from 34.10% for Buy and Hold to 18.47%. Across five SPY seeds, PPO-HRAP remains stable with mean total return $0.2725 \pm 0.0109$ and mean Sharpe ratio $0.6219 \pm 0.0565$. Single-run cross-asset tests on QQQ and DIA further show that the proposed method ranks first on total return and Sharpe ratio for all three reported assets. These results suggest that blending learned actions with a volatility-aware regime prior is a practical way to improve risk-adjusted trading behavior, although the current policy still incurs high turnover and cross-asset robustness beyond SPY remains limited to single-run evidence.

Shown in full with attribution under the source's licence. Licence: abstract CC0

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.