שילוב סוכני למידת חיזוק עמוקה למסחר במניות
סיכום
המאמר מתאר גישה למסחר אוטומטי במניות המשלבת שלושה אלגוריתמי למידת חיזוק מסוג שחקן-מבקר: אופטימיזציה של מדיניות פרוקסימלית, שחקן-מבקר יתרון וגרדיאנט מדיניות דטרמיניסטי עמוק. הסוכנים מאומנים ללמוד החלטות מסחר במטרה למקסם תשואות השקעה, ויתרונותיהם משולבים באנסמבל שנועד להסתגל לתנאי שוק משתנים. כמו כן, נעשה שימוש בטכניקת טעינה לפי דרישה לעיבוד נתונים בהיקף גדול תוך הגבלת השימוש בזיכרון במהלך אימון עם פעולות רציפות.
הגישה מוערכת על מניות נזילות מתוך יקום דאו ג'ונס ומושווית לאלגוריתמי הלמידה הבודדים, למדד Dow Jones Industrial Average ולתיק שונות מינימלית. התוצאה המדווחת היא שהאנסמבל משיג יחס שארפ גבוה יותר מהשיטות הבודדות וממדדי הבסיס הללו. הקטע אינו מציין את תקופת ההערכה, אילוצי תיק מפורטים, הנחות על עלויות עסקה או בדיקות עמידות. לכן, יש להבין את טענת הביצועים במסגרת הגדרת המבחן של המחקר, ולא כראיה לתשואות עתידיות.
רעיונות מרכזיים
- האסטרטגיה משלבת סוכני שחקן-מבקר של PPO, A2C ו-DDPG באנסמבל.
- הסוכנים לומדים החלטות מסחר במניות כשהתשואה על ההשקעה היא היעד.
- עיבוד נתונים בטעינה לפי דרישה משמש להפחתת דרישות הזיכרון במהלך האימון.
- ההערכה משווה את האנסמבל לשיטות המרכיבות אותו, למדד שוק ולתיק שונות מינימלית.
- לפי הדיווח, לאנסמבל יחס השארפ הגבוה ביותר בהשוואות האלה, בכפוף להגדרת המבחן של המחקר.
תגיות
הטקסט המלא
# Deep Reinforcement Learning for Automated Stock Trading: An Ensemble Strategy
# Deep Reinforcement Learning for Automated Stock Trading: An Ensemble Strategy
Stock trading strategies play a critical role in investment. However, it is challenging to design a profitable strategy in a complex and dynamic stock market. In this paper, we propose an ensemble strategy that employs deep reinforcement schemes to learn a stock trading strategy by maximizing investment return. We train a deep reinforcement learning agent and obtain an ensemble trading strategy using three actor-critic based algorithms: Proximal Policy Optimization (PPO), Advantage Actor Critic (A2C), and Deep Deterministic Policy Gradient (DDPG). The ensemble strategy inherits and integrates the best features of the three algorithms, thereby robustly adjusting to different market situations. In order to avoid the large memory consumption in training networks with continuous action space, we employ a load-on-demand technique for processing very large data. We test our algorithms on the 30 Dow Jones stocks that have adequate liquidity. The performance of the trading agent with different reinforcement learning algorithms is evaluated and compared with both the Dow Jones Industrial Average index and the traditional min-variance portfolio allocation strategy. The proposed deep ensemble strategy is shown to outperform the three individual algorithms and two baselines in terms of the risk-adjusted return measured by the Sharpe ratio. This work is fully open-sourced at \href{https://github.com/AI4Finance-Foundation/Deep-Reinforcement-Learning-for-Automated-Stock-Trading-Ensemble-Strategy-ICAIF-2020}{GitHub}.מוצג במלואו בציון המקור ובהתאם לרישיון שלו. רישיון: abstract CC0
הסיכום נכתב בידי סוכן המחקר של Stratmill על סמך המקור; הוא אינו העתק של המקור.