التعلم المعزز العميق لإدارة محافظ العملات المشفرة مع مراعاة المخاطر
الملخص
تقدم الورقة وكيلاً للتعلم المعزز العميق لإدارة المحافظ، يوازن بين السعي إلى العائد والحد من المخاطر. وتعدّل سياسته المستهدفة مدى تفضيل الوكيل للفعل الأمثل، باستخدام معامل قابل للضبط للجشع لتشجيع الخيارات الأقل خطراً. ويقيّم المؤلفون النهج على بيانات سوق العملات المشفرة، التي اختيروها لتوافر ملاحظات وفيرة على مستوى الدقيقة وارتفاع تقلبها.
حقق الوكيل في فترة الاختبار المذكورة عائداً قدره 1800% وكان الأقل مخاطرة بين الأساليب المقارنة. وتشير تجارب إضافية إلى أداء متين في ظل تقلبات السوق العالية وفترات التدريب القصيرة. ولا يحدد المقتطف مقياس المخاطر أو أساليب المقارنة أو الأصول أو تكاليف المعاملات أو تصميم التقييم؛ لذلك لا يمكن تقييم ادعاءات الأداء بالتفصيل أو افتراض تعميمها خارج التجارب المذكورة.
الأفكار الرئيسية
- يحسن الوكيل المقترح إدارة المحافظ مع مراعاة العائد والحد من المخاطر.
- تضبط سياسة مستهدفة قابلة للتهيئة تفضيل الفعل الأمثل، وتهدف إلى ترجيح الأفعال الأقل خطراً.
- يُقيّم النهج باستخدام بيانات سوق العملات المشفرة ذات الملاحظات على مستوى الدقيقة.
- يذكر المؤلفون عائداً قدره 1800% في فترة الاختبار وأقل مستوى مخاطر بين الأساليب المقارنة.
- تشير تجارب إضافية إلى المتانة أمام التقلب الشديد وقصر فترات التدريب، مع أن المقتطف لا يفصل التقييم.
الوسوم
النص الكامل
# Automatic Financial Trading Agent for Low-risk Portfolio Management using Deep Reinforcement Learning # Automatic Financial Trading Agent for Low-risk Portfolio Management using Deep Reinforcement Learning The autonomous trading agent is one of the most actively studied areas of artificial intelligence to solve the capital market portfolio management problem. The two primary goals of the portfolio management problem are maximizing profit and restrainting risk. However, most approaches to this problem solely take account of maximizing returns. Therefore, this paper proposes a deep reinforcement learning based trading agent that can manage the portfolio considering not only profit maximization but also risk restraint. We also propose a new target policy to allow the trading agent to learn to prefer low-risk actions. The new target policy can be reflected in the update by adjusting the greediness for the optimal action through the hyper parameter. The proposed trading agent verifies the performance through the data of the cryptocurrency market. The Cryptocurrency market is the best test-ground for testing our trading agents because of the huge amount of data accumulated every minute and the market volatility is extremely large. As a experimental result, during the test period, our agents achieved a return of 1800% and provided the least risky investment strategy among the existing methods. And, another experiment shows that the agent can maintain robust generalized performance even if market volatility is large or training period is short.
يُعرض النص كاملًا مع نسبه إلى مصدره وفقًا لترخيصه. الترخيص: abstract CC0
أعدّ وكيل الأبحاث في Stratmill هذا الملخص استنادًا إلى المصدر الأصلي؛ وهو ليس نسخة منه.