למידת Q עמוקה למסחר בתיק רב-נכסי
סיכום
מחקר זה מנסח מסחר בתיק השקעות כתהליך החלטה מרקובי ומאמן סוכן באמצעות למידת Q עמוקה לבחירת הקצאות בין נכסים. מרחב הפעולות שלו בדיד וקומבינטורי: עבור כל נכס, הסוכן בוחר כיוון מסחר בגודל מוגדר מראש. כדי להתמודד עם פעולות המפרות אילוצים, השיטה ממפה הצעה לא ישימה לחלופה הישימה הקרובה ביותר. היא מתארת גם תכנון של סוכן ורשת Q שנועד לנהל את מרחב הפעולות הרב-נכסי ולדמות פעולות ישימות בכל מצב.
הגישה נבחנת באמצעות בקטסטים על שני תיקי השקעות מייצגים, והתוצאות מדווחות כעדיפות על פני אסטרטגיות ייחוס. המסמך אינו מזהה את התיקים, מדדי הייחוס, תקופת ההערכה או הנחות על עלויות עסקה, ולכן אי אפשר לשפוט על סמך התיאור בלבד עד כמה ההשוואה רחבה או רלוונטית למסחר חי.
רעיונות מרכזיים
- תהליך ההחלטה בתיק ממודל כתהליך החלטה מרקובי ומאומן באמצעות למידת Q עמוקה.
- הסוכן בוחר לכל נכס כיוונים בדידים וגדלי מסחר מוגדרים מראש.
- שלב מיפוי ממיר פעולות מוצעות שאינן ישימות לפעולות ישימות קרובות.
- לפי הדיווח, הגישה השיגה תוצאות טובות יותר מאסטרטגיות הייחוס בבקטסטים על שני תיקים, אך פרטי ההערכה אינם נמסרים.
תגיות
הטקסט המלא
# An intelligent financial portfolio trading strategy using deep Q-learning # An intelligent financial portfolio trading strategy using deep Q-learning Portfolio traders strive to identify dynamic portfolio allocation schemes so that their total budgets are efficiently allocated through the investment horizon. This study proposes a novel portfolio trading strategy in which an intelligent agent is trained to identify an optimal trading action by using deep Q-learning. We formulate a Markov decision process model for the portfolio trading process, and the model adopts a discrete combinatorial action space, determining the trading direction at prespecified trading size for each asset, to ensure practical applicability. Our novel portfolio trading strategy takes advantage of three features to outperform in real-world trading. First, a mapping function is devised to handle and transform an initially found but infeasible action into a feasible action closest to the originally proposed ideal action. Second, by overcoming the dimensionality problem, this study establishes models of agent and Q-network for deriving a multi-asset trading strategy in the predefined action space. Last, this study introduces a technique that has the advantage of deriving a well-fitted multi-asset trading strategy by designing an agent to simulate all feasible actions in each state. To validate our approach, we conduct backtests for two representative portfolios and demonstrate superior results over the benchmark strategies.
מוצג במלואו בציון המקור ובהתאם לרישיון שלו. רישיון: abstract CC0
הסיכום נכתב בידי סוכן המחקר של Stratmill על סמך המקור; הוא אינו העתק של המקור.