עבור לתוכן
כל מסמכי הספרייה

למידת חיזוק חוזרת לניהול סיכונים במסחר זוגות

מאמר arXiv papers · מחבר: Weiguang Han et al.

סיכום

המסמך מתאר את CREDIT, גישת למידת חיזוק למסחר זוגות המשתמשת ביחידה חוזרת מגודרת דו־כיוונית ובתשומת לב זמנית. הארכיטקטורה נועדה ללכוד קשרים לאורך זמן בתנועות המחיר של שני נכסים, לרבות דפוסים שעלולים לחמוק כאשר מצבי השוק נבחנים בנפרד. הגישה משתמשת גם בתגמול המביא בחשבון הן רווח ממסחר והן סיכון, במטרה להרתיע מפני עסקאות עם פוטנציאל גבוה לרווח ולהפסד.

המחברים מציגים את השיטה כמענה למסחר תכוף, לעלויות עסקה ולנטילת סיכונים בגישות למידת חיזוק קודמות. הם מדווחים כי CREDIT עולה בביצועיה על שיטות למידת חיזוק קיימות ומניבה רווח משמעותי בניסויים שהתבססו על נתוני מניות בארצות הברית במשך חמש שנים. הקטע אינו מפרט מדדי ייחוס, הנחות על עלויות, מדידות סיכון או תוצאות מספריות, ולכן אי אפשר להעריך את גודל היתרון המדווח או את חוסנו. לא הוכחו כאן ביצועים בשווקים או בתקופות אחרות.

רעיונות מרכזיים

  • CREDIT מיישמת למידה חוזרת ותשומת לב זמנית על היסטוריית מחירים לצורך מסחר זוגות בשני נכסים.
  • פונקציית התגמול שלה מביאה בחשבון גם רווח וגם סיכון מסחר.
  • התכנון נועד ללמוד דפוסי מחיר ארוכי טווח יותר ולהפחית מסחר מוגזם.
  • לפי הדיווח, השיטות שנבדקו בניסויים על נתוני מניות מארצות הברית לאורך חמש שנים השיגו ביצועים טובים יותר משיטות למידת חיזוק אחרות.
  • הקטע אינו מספק די פרטים להערכת החוסן או עלויות המסחר בעולם האמיתי.

תגיות

הטקסט המלא
# Mastering Pair Trading with Risk-Aware Recurrent Reinforcement Learning


# Mastering Pair Trading with Risk-Aware Recurrent Reinforcement Learning









Although pair trading is the simplest hedging strategy for an investor to eliminate market risk, it is still a great challenge for reinforcement learning (RL) methods to perform pair trading as human expertise. It requires RL methods to make thousands of correct actions that nevertheless have no obvious relations to the overall trading profit, and to reason over infinite states of the time-varying market most of which have never appeared in history. However, existing RL methods ignore the temporal connections between asset price movements and the risk of the performed trading. These lead to frequent tradings with high transaction costs and potential losses, which barely reach the human expertise level of trading. Therefore, we introduce CREDIT, a risk-aware agent capable of learning to exploit long-term trading opportunities in pair trading similar to a human expert. CREDIT is the first to apply bidirectional GRU along with the temporal attention mechanism to fully consider the temporal correlations embedded in the states, which allows CREDIT to capture long-term patterns of the price movements of two assets to earn higher profit. We also design the risk-aware reward inspired by the economic theory, that models both the profit and risk of the tradings during the trading period. It helps our agent to master pair trading with a robust trading preference that avoids risky trading with possible high returns and losses. Experiments show that it outperforms existing reinforcement learning methods in pair trading and achieves a significant profit over five years of U.S. stock data.

מוצג במלואו בציון המקור ובהתאם לרישיון שלו. רישיון: abstract CC0

הסיכום נכתב בידי סוכן המחקר של Stratmill על סמך המקור; הוא אינו העתק של המקור.