למידת חיזוק עמוקה לניהול תיקי קריפטו
סיכום
מאמר זה מציג מסגרת למידת חיזוק נטולת מודל להקצאה מחדש של תיק בין נכסים פיננסיים. מרכיביה כוללים אנסמבל של מעריכים עצמאיים, זיכרון של משקולות תיק קודמות, למידת אצווה סטוכסטית מקוונת ופונקציית תגמול מפורשת. המסגרת מיושמת באמצעות רשתות נוירונים קונבולוציוניות, רקורסיביות ורשתות זיכרון לטווח קצר וארוך.
ההערכה המדווחת משתמשת בשלושה בקטסטים על נתוני מטבעות קריפטוגרפיים, עם תקופות מסחר של 30 דקות. המחברים משווים את המודלים לאסטרטגיות אחרות לבחירת תיק ומדווחים ששלוש המימושים שלהם תפסו את שלושת המקומות הראשונים בכל הניסויים, למרות עמלות של 0.25%, עם תשואה של לפחות פי ארבעה במשך 50 ימים. אלה תוצאות של בקטסט היסטורי, ולא ראיה לביצועים עתידיים. הקטע אינו מזהה את הנכסים, חלוקות הנתונים, מדדי הסיכון או בדיקות העמידות, ולכן מספק פרטים מוגבלים להערכת התאמת יתר או היתכנות של מסחר חי.
רעיונות מרכזיים
- המסגרת לומדת הקצאות מחדש של תיק ללא הסתמכות על מודל מפורש של השוק הפיננסי.
- היא משלבת מעריכי נכסים עצמאיים עם זיכרון של משקולות תיק קודמות.
- למידת אצווה סטוכסטית מקוונת ופונקציית תגמול מפורשת הן מרכיבים מרכזיים בשיטה.
- המחברים בוחנים את המימושים CNN, RNN ו־LSTM בבדיקות היסטוריות על נתוני מטבעות קריפטוגרפיים.
- דירוגי הבקטסט והתשואות המדווחים אינם מוכיחים ביצועים בשוק חי.
תגיות
הטקסט המלא
# A Deep Reinforcement Learning Framework for the Financial Portfolio Management Problem # A Deep Reinforcement Learning Framework for the Financial Portfolio Management Problem Financial portfolio management is the process of constant redistribution of a fund into different financial products. This paper presents a financial-model-free Reinforcement Learning framework to provide a deep machine learning solution to the portfolio management problem. The framework consists of the Ensemble of Identical Independent Evaluators (EIIE) topology, a Portfolio-Vector Memory (PVM), an Online Stochastic Batch Learning (OSBL) scheme, and a fully exploiting and explicit reward function. This framework is realized in three instants in this work with a Convolutional Neural Network (CNN), a basic Recurrent Neural Network (RNN), and a Long Short-Term Memory (LSTM). They are, along with a number of recently reviewed or published portfolio-selection strategies, examined in three back-test experiments with a trading period of 30 minutes in a cryptocurrency market. Cryptocurrencies are electronic and decentralized alternatives to government-issued money, with Bitcoin as the best-known example of a cryptocurrency. All three instances of the framework monopolize the top three positions in all experiments, outdistancing other compared trading algorithms. Although with a high commission rate of 0.25% in the backtests, the framework is able to achieve at least 4-fold returns in 50 days.
מוצג במלואו בציון המקור ובהתאם לרישיון שלו. רישיון: abstract CC0
הסיכום נכתב בידי סוכן המחקר של Stratmill על סמך המקור; הוא אינו העתק של המקור.