עבור לתוכן
כל מסמכי הספרייה

למידת חיזוק עמוקה למסחר דינמי בזוגות מטבעות קריפטו

מאמר arXiv papers · מחבר: Damian Lebiedź et al.

סיכום

המחקר בוחן שימוש בלמידת חיזוק עמוקה כשכבת ביצוע למסחר בזוגות בשווקי מטבעות קריפטוגרפיים תנודתיים. תחילה המסגרת מסננת ומדרגת זוגות מועמדים, ואז סוכן Proximal Policy Optimization עם שכבת Long Short-Term Memory מקבל החלטות ביצוע. תכנון עם סיכון קבוע וממוצע אדפטיבי, לצד בקרות סיכון דטרמיניסטיות, מגביל את הסוכן ונועד לצמצם סיכון לסטייה. ההערכה משתמשת בנתוני חוזים עתידיים שעתיים של Binance USD-M ומשווה את המדיניות שנלמדה לקו בסיס היוריסטי.

התוצאות המדווחות מחוץ למדגם נוטות לטובת מדיניות למידת החיזוק, ומבחן בוטסטרפ של בלוקים מעגליים סטציונריים מצביע על ביצועי יתר מובהקים סטטיסטית לאחר התאמה לסיכון, ברמת מובהקות של 10 אחוז. התוצאה אינה עומדת ברף המחמיר יותר של 5 אחוז מובהקות, והמחקר קושר זאת לשונות אידיוסינקרטית גבוהה בנכסים דיגיטליים. הראיות ייחודיות לנתונים ולהגדרות שנבדקו; התיאור אינו מספק את היקף הביצועים, הנחות לגבי עלויות עסקה או בדיקות בזירות ובתקופות אחרות. התכנון ההיברידי מציע דרך להגביל החלטות ביצוע שנלמדו, אך העמידות הרחבה יותר עדיין אינה מוכרעת.

רעיונות מרכזיים

  • המערכת המוצעת משלבת בחירה סטטיסטית של זוגות עם שכבת ביצוע נלמדת.
  • סוכן PPO עם LSTM מקבל החלטות ביצוע בתוך גבולות סיכון דטרמיניסטיים.
  • המחקר מעריך את הגישה מול קו בסיס היוריסטי, על נתוני חוזים עתידיים של Binance USD-M בתדירות שעתית.
  • ביצועי היתר המדווחים לאחר התאמה לסיכון מובהקים ברמת 10 אחוז, אך לא ברמת 5 אחוז.
  • התוצאות ייחודיות לשוק, לנתונים ולהגדרות האסטרטגיה שנבדקו.

תגיות

הטקסט המלא
# Dynamic Multi-Pair Trading Strategy in Cryptocurrency Markets with Deep Reinforcement Learning


# Dynamic Multi-Pair Trading Strategy in Cryptocurrency Markets with Deep Reinforcement Learning









This study aims to determine whether the application of Deep Reinforcement Learning (DRL) as a specialized execution overlay can enhance pair trading in highly volatile cryptocurrency markets. Although classical implementations of the strategy have proven successful in traditional equities, they frequently exhibit rigidity and suffer from severe divergence risks when applied to high-variance environments. To address this need, this research introduces novel concepts. To construct a robust system, we developed a hierarchical "Filter-then-Rank" pair selection methodology and a proprietary "Fixed Risk, Adaptive Mean" execution model. The system employs a Proximal Policy Optimization (PPO) agent with a Long Short-Term Memory (LSTM) layer to govern execution decisions within strict deterministic risk management boundaries. Evaluated on 1-hour interval data from the Binance USD-M Futures market, the optimized RL policy achieved an out-of-sample performance that substantially outperformed the heuristic baseline. A stationary circular block bootstrap robustness check confirms that the agent's risk-adjusted outperformance is statistically significant at the 10 percent level. Although falling marginally short of the stricter 5 percent threshold, this result highlights the extreme idiosyncratic variance characteristic of digital assets. Ultimately, this thesis contributes to the quantitative finance literature by introducing a hybrid architecture that combines statistical arbitrage with DRL execution policies. Furthermore, it delivers a novel framework for safe reinforcement learning via deterministic shielding, proving that anchoring a neural policy to statistically robust boundaries successfully mitigates severe divergence risks.

מוצג במלואו בציון המקור ובהתאם לרישיון שלו. רישיון: abstract CC0

הסיכום נכתב בידי סוכן המחקר של Stratmill על סמך המקור; הוא אינו העתק של המקור.