PPO 交易:利用市场状态感知敞口先验控制回撤
文章 arXiv papers · 作者: Duong Hien Chi Kien et al.
总结
PPO-HRAP 将近端策略优化与根据市场状态得出的目标敞口结合。其策略观察市场特征和投资组合状态,然后将学习所得动作与市场状态目标相结合。奖励函数考虑投资组合对数收益、以 VIX 为条件的回撤增加、相对目标敞口的偏离以及换手成本。该设计旨在兼顾参与收益和控制回撤,解决以盈利为导向的策略倾向于保持高仓位,以及过重风险惩罚可能使策略过于谨慎的问题。
在留出测试的SPY窗口(2020至2022)中,研究报告称,与买入并持有相比,该方法的最大回撤较低,同时给出了收益和风险调整后绩效指标。研究称,五个SPY随机种子的结果稳定;在QQQ和DIA上的单次运行测试中,该方法的总收益和夏普比率也在所比较的方法中排名第一。作者承认换手率较高,跨资产证据有限,因此其更广泛的稳健性仍不确定。
核心观点
- PPO-HRAP 将学习策略动作与根据市场状态得出的目标敞口相结合。
- 奖励函数考虑收益、以 VIX 为条件的回撤增加、目标敞口偏离和换手率。
- 留出的 SPY 评估报告称,最大回撤低于买入并持有策略。
- 五个 SPY 随机种子下的结果稳定,而 QQQ 和 DIA 的证据来自单次运行。
- 该方法仍存在换手率较高、跨资产稳健性证据有限的问题。
标签
全文
# PPO-HRAP: Proximal Policy Optimization with a Hybrid Regime-Aware Policy for Risk-Controlled Trading # PPO-HRAP: Proximal Policy Optimization with a Hybrid Regime-Aware Policy for Risk-Controlled Trading Reinforcement learning for trading often struggles to balance upside participation with drawdown control. Profit-only policies can collapse toward passive long exposure on upward-drifting assets, while aggressively risk-penalized rewards can become too defensive during volatile periods. This paper proposes PPO-HRAP, a hybrid regime-aware policy that combines Proximal Policy Optimization with an interpretable regime prior. The agent observes both market features and portfolio-state variables, receives a reward combining portfolio log return, VIX-conditioned drawdown-increase penalty, target-exposure deviation, and turnover cost, and executes a blended action between the PPO actor output and a regime-derived target exposure. On the held-out 2020-2022 SPY test window, PPO-HRAP achieves 27.62% total return, 8.48% annualized return, 0.6447 Sharpe ratio, 0.8588 Sortino ratio, and 0.4592 Calmar ratio, while reducing maximum drawdown from 34.10% for Buy and Hold to 18.47%. Across five SPY seeds, PPO-HRAP remains stable with mean total return $0.2725 \pm 0.0109$ and mean Sharpe ratio $0.6219 \pm 0.0565$. Single-run cross-asset tests on QQQ and DIA further show that the proposed method ranks first on total return and Sharpe ratio for all three reported assets. These results suggest that blending learned actions with a volatility-aware regime prior is a practical way to improve risk-adjusted trading behavior, although the current policy still incurs high turnover and cross-asset robustness beyond SPY remains limited to single-run evidence.
在遵守原作品许可的前提下,附作者信息全文展示。 许可协议: abstract CC0
此摘要由 Stratmill 研究智能体根据原文撰写,并非原文副本。