עבור לתוכן
כל מסמכי הספרייה

למידת חיזוק לניהול תיקי מטבעות קריפטוגרפיים

מאמר arXiv papers · מחבר: Kamal Paykan

סיכום

המאמר מציע להשתמש ב-Soft Actor-Critic (SAC) וב-Deep Deterministic Policy Gradient (DDPG) לניהול תיקי מטבעות קריפטוגרפיים. הסוכנים לומדים פעולות מסחר רציפות בסביבה מדומה באמצעות נתוני שוק היסטוריים, ומתאימים את משקלי התיק כדי לשאוף לתשואות מצטברות תוך התחשבות בסיכון לירידות ובעלויות עסקה. SAC משתמש בפונקציית מטרה עם רגולריזציה של אנטרופיה, שהמאמר מקשר ליציבות רבה יותר בתנאים רועשים.

ניסויים בכמה מטבעות קריפטוגרפיים מדווחים ששני הסוכנים עולים בביצועיהם על אמות מידה של משקל שווה ושל תיק ממוצע-שונות, כאשר SAC יציב וחסין יותר מ-DDPG. התקציר אינו מפרט את הנכסים, תקופת המדגם, הנחות השוק, מודל עלויות העסקה או תוצאות מספריות. מכיוון שהראיות מבוססות על הערכה ניסויית עם נתונים היסטוריים וסביבה מדומה, הן אינן מראות כשלעצמן שהשיטות יניבו תוצאות דומות במסחר חי או בתנאי שוק אחרים.

רעיונות מרכזיים

  • סוכני SAC ו-DDPG לומדים פעולות תיק רציפות מנתוני שוק היסטוריים בסביבה מדומה.
  • הסוכנים מתאימים את משקלי התיק תוך התחשבות בתשואות, בסיכון לירידות ובעלויות עסקה.
  • המאמר מדווח על ביצועים טובים מאמות מידה של משקל שווה ושל ממוצע-שונות בכמה מטבעות קריפטוגרפיים.
  • לפי הדיווח, SAC יציב יותר מ-DDPG בתנאי שוק רועשים.
  • התקציר אינו מספק את פרטי הניסוי או נתוני ביצועים מספריים.

תגיות

הטקסט המלא
# 2511.20678


# Cryptocurrency Portfolio Management with Reinforcement Learning: Soft Actor--Critic and Deep Deterministic Policy Gradient Algorithms









This paper proposes a reinforcement learning--based framework for cryptocurrency portfolio management using the Soft Actor--Critic (SAC) and Deep Deterministic Policy Gradient (DDPG) algorithms. Traditional portfolio optimization methods often struggle to adapt to the highly volatile and nonlinear dynamics of cryptocurrency markets. To address this, we design an agent that learns continuous trading actions directly from historical market data through interaction with a simulated trading environment. The agent optimizes portfolio weights to maximize cumulative returns while minimizing downside risk and transaction costs. Experimental evaluations on multiple cryptocurrencies demonstrate that the SAC and DDPG agents outperform baseline strategies such as equal-weighted and mean--variance portfolios. The SAC algorithm, with its entropy-regularized objective, shows greater stability and robustness in noisy market conditions compared to DDPG. These results highlight the potential of deep reinforcement learning for adaptive and data-driven portfolio management in cryptocurrency markets.

מוצג במלואו בציון המקור ובהתאם לרישיון שלו. רישיון: abstract CC0

הסיכום נכתב בידי סוכן המחקר של Stratmill על סמך המקור; הוא אינו העתק של המקור.