TD3 בפעולות רציפות למסחר במניות ובביטקוין
סיכום
המחקר מיישם את Twin-Delayed Deep Deterministic Policy Gradient (TD3), שיטת למידה באמצעות חיזוק עם פעולות רציפות, על מחירי הסגירה היומיים של מניית Amazon ושל ביטקוין. פעולת המסחר מייצגת הן את הפוזיציה שיש לקחת והן את מספר המניות לסחור, ובכך מרחיבה גישות המשתמשות בפעולות בדידות. המאמר משווה את האסטרטגיות המתקבלות לשיטות ניתוח טכני, לגישות אחרות של למידה באמצעות חיזוק ולאסטרטגיות סטוכסטיות ודטרמיניסטיות.
הביצועים מוערכים לפי תשואה ויחס Sharpe. התוצאות המדווחות מצביעות על כך שבחירת הפוזיציה וגודל העסקה יחד עשויה לשפר את הביצועים במדדים אלה. המסמך אינו מספק פרטי הערכה, תוצאות ספציפיות או ראיות לחוסן בנכסים אחרים או בתקופות שוק אחרות. מסקנותיו מתארות אפוא את השווקים ואת שיטות ההשוואה שנחקרו, ואינן מבססות ש-TD3 בפעולות רציפות יניב ביצועים טובים יותר במסגרות אחרות.
רעיונות מרכזיים
- TD3 משמש להפקת פעולות מסחר ממחירי הסגירה היומיים של Amazon וביטקוין.
- מרחב הפעולות מייצג הן את פוזיציית המסחר והן את מספר המניות הנסחרות.
- המחקר משווה את TD3 לאסטרטגיות ניתוח טכני, למידה באמצעות חיזוק, סטוכסטיות ודטרמיניסטיות.
- תשואה ויחס Sharpe הם מדדי הביצועים שצוינו.
- ההשוואה המדווחת מעדיפה שילוב של בחירת פוזיציה עם קביעת גודל העסקה, במסגרת התנאים שנבדקו במחקר.
תגיות
הטקסט המלא
# Algorithmic Trading Using Continuous Action Space Deep Reinforcement Learning # Algorithmic Trading Using Continuous Action Space Deep Reinforcement Learning Price movement prediction has always been one of the traders' concerns in financial market trading. In order to increase their profit, they can analyze the historical data and predict the price movement. The large size of the data and complex relations between them lead us to use algorithmic trading and artificial intelligence. This paper aims to offer an approach using Twin-Delayed DDPG (TD3) and the daily close price in order to achieve a trading strategy in the stock and cryptocurrency markets. Unlike previous studies using a discrete action space reinforcement learning algorithm, the TD3 is continuous, offering both position and the number of trading shares. Both the stock (Amazon) and cryptocurrency (Bitcoin) markets are addressed in this research to evaluate the performance of the proposed algorithm. The achieved strategy using the TD3 is compared with some algorithms using technical analysis, reinforcement learning, stochastic, and deterministic strategies through two standard metrics, Return and Sharpe ratio. The results indicate that employing both position and the number of trading shares can improve the performance of a trading system based on the mentioned metrics.
מוצג במלואו בציון המקור ובהתאם לרישיון שלו. רישיון: abstract CC0
הסיכום נכתב בידי סוכן המחקר של Stratmill על סמך המקור; הוא אינו העתק של המקור.