تداول PPO مع تعرض مرجعي واعٍ بأنظمة السوق لضبط التراجع
الملخص
يجمع PPO-HRAP بين تحسين السياسة القريبة والتعرض المستهدف المستمد من نظام السوق. وترصد سياسته سمات السوق وحالة المحفظة، ثم تمزج الإجراء المتعلَّم بالتعرض المستهدف للنظام. وتراعي المكافأة عوائد المحفظة اللوغاريتمية وزيادات التراجع المشروطة بـVIX والانحراف عن التعرض المستهدف وتكاليف دوران المحفظة. ويهدف التصميم إلى الموازنة بين المشاركة في المكاسب وضبط التراجع، ومعالجة ميل السياسات التي تركز على الربح إلى البقاء مستثمرة بكثافة وميل عقوبات المخاطر الشديدة إلى جعل السياسات حذرة أكثر من اللازم.
في فترة SPY محجوبة عن التدريب تمتد من 2020 إلى 2022، تفيد الدراسة بانخفاض التراجع الأقصى مقارنة بالشراء والاحتفاظ، إلى جانب مقاييس العائد والأداء المعدل بالمخاطر. وتُوصف النتائج عبر خمس بذور SPY بأنها مستقرة؛ كما تصنف اختبارات التشغيل الواحد على QQQ وDIA الطريقة أولى من حيث العائد الإجمالي ونسبة Sharpe بين الطرق المقارنة. ويقر المؤلفون بارتفاع دوران المحفظة ومحدودية الأدلة عبر الأصول، لذا تظل المتانة الأوسع غير مؤكدة.
الأفكار الرئيسية
- يمزج PPO-HRAP إجراء السياسة المتعلَّم بتعرض مستهدف مستمد من أنظمة السوق.
- تراعي مكافأته العوائد وزيادات التراجع المشروطة بـVIX والانحراف عن التعرض المستهدف ودوران المحفظة.
- يفيد تقييم SPY المحجوب عن التدريب بانخفاض التراجع الأقصى مقارنة بالشراء والاحتفاظ.
- تستقر النتائج عبر خمس بذور SPY، بينما تستند أدلة QQQ وDIA إلى تشغيل واحد.
- لا تزال الطريقة تعاني من دوران مرتفع للمحفظة ومحدودية أدلة المتانة عبر الأصول.
الوسوم
النص الكامل
# PPO-HRAP: Proximal Policy Optimization with a Hybrid Regime-Aware Policy for Risk-Controlled Trading # PPO-HRAP: Proximal Policy Optimization with a Hybrid Regime-Aware Policy for Risk-Controlled Trading Reinforcement learning for trading often struggles to balance upside participation with drawdown control. Profit-only policies can collapse toward passive long exposure on upward-drifting assets, while aggressively risk-penalized rewards can become too defensive during volatile periods. This paper proposes PPO-HRAP, a hybrid regime-aware policy that combines Proximal Policy Optimization with an interpretable regime prior. The agent observes both market features and portfolio-state variables, receives a reward combining portfolio log return, VIX-conditioned drawdown-increase penalty, target-exposure deviation, and turnover cost, and executes a blended action between the PPO actor output and a regime-derived target exposure. On the held-out 2020-2022 SPY test window, PPO-HRAP achieves 27.62% total return, 8.48% annualized return, 0.6447 Sharpe ratio, 0.8588 Sortino ratio, and 0.4592 Calmar ratio, while reducing maximum drawdown from 34.10% for Buy and Hold to 18.47%. Across five SPY seeds, PPO-HRAP remains stable with mean total return $0.2725 \pm 0.0109$ and mean Sharpe ratio $0.6219 \pm 0.0565$. Single-run cross-asset tests on QQQ and DIA further show that the proposed method ranks first on total return and Sharpe ratio for all three reported assets. These results suggest that blending learned actions with a volatility-aware regime prior is a practical way to improve risk-adjusted trading behavior, although the current policy still incurs high turnover and cross-asset robustness beyond SPY remains limited to single-run evidence.
يُعرض النص كاملًا مع نسبه إلى مصدره وفقًا لترخيصه. الترخيص: abstract CC0
أعدّ وكيل الأبحاث في Stratmill هذا الملخص استنادًا إلى المصدر الأصلي؛ وهو ليس نسخة منه.