دمج وكلاء التعلم المعزز العميق لتداول الأسهم
الملخص
تصف الورقة نهجًا آليًا لتداول الأسهم يجمع ثلاثة خوارزميات للتعلم المعزز من نوع الفاعل والناقد: تحسين السياسة القريب، والفاعل الناقد ذي الميزة، وتدرج السياسة الحتمية العميق. ويُدرَّب الوكلاء على تعلم قرارات التداول بهدف تعظيم عوائد الاستثمار، وتُدمج نقاط قوتهم في مجموعة تهدف إلى التكيف مع ظروف السوق. كما تُستخدم تقنية التحميل عند الطلب لمعالجة كميات كبيرة من البيانات مع الحد من استخدام الذاكرة أثناء التدريب على الأفعال المستمرة.
يُقيَّم النهج على أسهم سائلة من مجموعة داو جونز ويُقارن بخوارزميات التعلم المنفردة ومؤشر داو جونز الصناعي ومحفظة تباينها أدنى. والنتيجة المُبلغ عنها هي أن المجموعة تحقق نسبة شارب أعلى من تلك الطرق المنفردة والمعايير المرجعية. ولا يحدد المقتطف فترة التقييم أو قيود المحفظة التفصيلية أو افتراضات تكاليف المعاملات أو اختبارات المتانة. لذلك ينبغي فهم ادعاء الأداء ضمن إعداد الاختبار المحدد في الدراسة، لا بوصفه دليلًا على عوائد مستقبلية.
الأفكار الرئيسية
- تجمع الاستراتيجية وكلاء الفاعل والناقد PPO وA2C وDDPG في مجموعة.
- يتعلم الوكلاء قرارات تداول الأسهم، ويكون عائد الاستثمار هو الهدف.
- تُستخدم معالجة البيانات عند الطلب لتقليل متطلبات الذاكرة أثناء التدريب.
- يقارن التقييم المجموعة بأساليبها المكونة لها ومؤشر سوق ومحفظة ذات تباين أدنى.
- يُذكر أن المجموعة حققت أعلى نسبة شارب في تلك المقارنات، وفقًا لإعداد الاختبار في الدراسة.
الوسوم
النص الكامل
# Deep Reinforcement Learning for Automated Stock Trading: An Ensemble Strategy
# Deep Reinforcement Learning for Automated Stock Trading: An Ensemble Strategy
Stock trading strategies play a critical role in investment. However, it is challenging to design a profitable strategy in a complex and dynamic stock market. In this paper, we propose an ensemble strategy that employs deep reinforcement schemes to learn a stock trading strategy by maximizing investment return. We train a deep reinforcement learning agent and obtain an ensemble trading strategy using three actor-critic based algorithms: Proximal Policy Optimization (PPO), Advantage Actor Critic (A2C), and Deep Deterministic Policy Gradient (DDPG). The ensemble strategy inherits and integrates the best features of the three algorithms, thereby robustly adjusting to different market situations. In order to avoid the large memory consumption in training networks with continuous action space, we employ a load-on-demand technique for processing very large data. We test our algorithms on the 30 Dow Jones stocks that have adequate liquidity. The performance of the trading agent with different reinforcement learning algorithms is evaluated and compared with both the Dow Jones Industrial Average index and the traditional min-variance portfolio allocation strategy. The proposed deep ensemble strategy is shown to outperform the three individual algorithms and two baselines in terms of the risk-adjusted return measured by the Sharpe ratio. This work is fully open-sourced at \href{https://github.com/AI4Finance-Foundation/Deep-Reinforcement-Learning-for-Automated-Stock-Trading-Ensemble-Strategy-ICAIF-2020}{GitHub}.يُعرض النص كاملًا مع نسبه إلى مصدره وفقًا لترخيصه. الترخيص: abstract CC0
أعدّ وكيل الأبحاث في Stratmill هذا الملخص استنادًا إلى المصدر الأصلي؛ وهو ليس نسخة منه.