یادگیری تقویتی عمیق برای مدیریت سبد رمزارز با آگاهی از ریسک
خلاصه
این مقاله عامل یادگیری تقویتی عمیقی را برای مدیریت سبد معرفی میکند که جستوجوی بازده را با مهار ریسک متوازن میسازد. سیاست هدف آن تنظیم میکند که عامل تا چه اندازه از اقدام بهینه پیروی کند و با استفاده از پارامتر قابلتنظیم طمع، انتخابهای کمریسکتر را تشویق میکند. نویسندگان این رویکرد را با دادههای بازار رمزارز ارزیابی میکنند؛ این دادهها بهدلیل فراوانی مشاهدات دقیقهای و نوسانپذیری بالا انتخاب شدهاند.
در دوره آزمون گزارششده، عامل بازده 1800% داشت و در میان روشهای مقایسهشده کمترین ریسک را ثبت کرد. آزمایشهای بیشتر نشاندهنده عملکردی باثبات در نوسانپذیری بالای بازار و دورههای آموزشی کوتاهاند. متن سنجه ریسک، روشهای مقایسه، داراییها، هزینههای معاملاتی یا طراحی ارزیابی را مشخص نمیکند؛ بنابراین ادعاهای عملکردی را نمیتوان با جزئیات ارزیابی کرد یا تعمیم آنها را فراتر از آزمایشهای بیانشده فرض گرفت.
ایدههای کلیدی
- عامل پیشنهادی، مدیریت سبد را با توجه همزمان به سود و مهار ریسک بهینه میکند.
- سیاست هدف قابلتنظیم، ترجیح برای اقدام بهینه را کنترل میکند و هدف آن گرایش به اقدامات کمریسکتر است.
- این رویکرد با دادههای بازار رمزارز و مشاهدات دقیقهای ارزیابی میشود.
- نویسندگان بازده 1800% در دوره آزمون و کمترین ریسک را در میان روشهای مقایسهشده گزارش میکنند.
- آزمایشهای بیشتر از پایداری در برابر نوسان بالا و دوره آموزشی کوتاه حکایت دارند، هرچند متن جزئیات ارزیابی را ارائه نمیکند.
برچسبها
متن کامل
# Automatic Financial Trading Agent for Low-risk Portfolio Management using Deep Reinforcement Learning # Automatic Financial Trading Agent for Low-risk Portfolio Management using Deep Reinforcement Learning The autonomous trading agent is one of the most actively studied areas of artificial intelligence to solve the capital market portfolio management problem. The two primary goals of the portfolio management problem are maximizing profit and restrainting risk. However, most approaches to this problem solely take account of maximizing returns. Therefore, this paper proposes a deep reinforcement learning based trading agent that can manage the portfolio considering not only profit maximization but also risk restraint. We also propose a new target policy to allow the trading agent to learn to prefer low-risk actions. The new target policy can be reflected in the update by adjusting the greediness for the optimal action through the hyper parameter. The proposed trading agent verifies the performance through the data of the cryptocurrency market. The Cryptocurrency market is the best test-ground for testing our trading agents because of the huge amount of data accumulated every minute and the market volatility is extremely large. As a experimental result, during the test period, our agents achieved a return of 1800% and provided the least risky investment strategy among the existing methods. And, another experiment shows that the agent can maintain robust generalized performance even if market volatility is large or training period is short.
با ذکر منبع و مطابق مجوز اثر، بهطور کامل نمایش داده میشود. مجوز: abstract CC0
این خلاصه را عامل پژوهشی Stratmill بر پایه متن اصلی نوشته است؛ نسخهای از اثر منبع نیست.