رفتن به محتوا
همه اسناد کتابخانه

یادگیری تقویتی عمیق برای مدیریت سبد رمزارز با آگاهی از ریسک

مقاله arXiv papers · نویسنده: Wonsup Shin et al.

خلاصه

این مقاله عامل یادگیری تقویتی عمیقی را برای مدیریت سبد معرفی می‌کند که جست‌وجوی بازده را با مهار ریسک متوازن می‌سازد. سیاست هدف آن تنظیم می‌کند که عامل تا چه اندازه از اقدام بهینه پیروی کند و با استفاده از پارامتر قابل‌تنظیم طمع، انتخاب‌های کم‌ریسک‌تر را تشویق می‌کند. نویسندگان این رویکرد را با داده‌های بازار رمزارز ارزیابی می‌کنند؛ این داده‌ها به‌دلیل فراوانی مشاهدات دقیقه‌ای و نوسان‌پذیری بالا انتخاب شده‌اند.

در دوره آزمون گزارش‌شده، عامل بازده 1800% داشت و در میان روش‌های مقایسه‌شده کمترین ریسک را ثبت کرد. آزمایش‌های بیشتر نشان‌دهنده عملکردی باثبات در نوسان‌پذیری بالای بازار و دوره‌های آموزشی کوتاه‌اند. متن سنجه ریسک، روش‌های مقایسه، دارایی‌ها، هزینه‌های معاملاتی یا طراحی ارزیابی را مشخص نمی‌کند؛ بنابراین ادعاهای عملکردی را نمی‌توان با جزئیات ارزیابی کرد یا تعمیم آن‌ها را فراتر از آزمایش‌های بیان‌شده فرض گرفت.

ایده‌های کلیدی

  • عامل پیشنهادی، مدیریت سبد را با توجه هم‌زمان به سود و مهار ریسک بهینه می‌کند.
  • سیاست هدف قابل‌تنظیم، ترجیح برای اقدام بهینه را کنترل می‌کند و هدف آن گرایش به اقدامات کم‌ریسک‌تر است.
  • این رویکرد با داده‌های بازار رمزارز و مشاهدات دقیقه‌ای ارزیابی می‌شود.
  • نویسندگان بازده 1800% در دوره آزمون و کمترین ریسک را در میان روش‌های مقایسه‌شده گزارش می‌کنند.
  • آزمایش‌های بیشتر از پایداری در برابر نوسان بالا و دوره آموزشی کوتاه حکایت دارند، هرچند متن جزئیات ارزیابی را ارائه نمی‌کند.

برچسب‌ها

متن کامل
# Automatic Financial Trading Agent for Low-risk Portfolio Management using Deep Reinforcement Learning


# Automatic Financial Trading Agent for Low-risk Portfolio Management using Deep Reinforcement Learning









The autonomous trading agent is one of the most actively studied areas of artificial intelligence to solve the capital market portfolio management problem. The two primary goals of the portfolio management problem are maximizing profit and restrainting risk. However, most approaches to this problem solely take account of maximizing returns. Therefore, this paper proposes a deep reinforcement learning based trading agent that can manage the portfolio considering not only profit maximization but also risk restraint. We also propose a new target policy to allow the trading agent to learn to prefer low-risk actions. The new target policy can be reflected in the update by adjusting the greediness for the optimal action through the hyper parameter. The proposed trading agent verifies the performance through the data of the cryptocurrency market. The Cryptocurrency market is the best test-ground for testing our trading agents because of the huge amount of data accumulated every minute and the market volatility is extremely large. As a experimental result, during the test period, our agents achieved a return of 1800% and provided the least risky investment strategy among the existing methods. And, another experiment shows that the agent can maintain robust generalized performance even if market volatility is large or training period is short.

با ذکر منبع و مطابق مجوز اثر، به‌طور کامل نمایش داده می‌شود. مجوز: abstract CC0

این خلاصه را عامل پژوهشی Stratmill بر پایه متن اصلی نوشته است؛ نسخه‌ای از اثر منبع نیست.