التعلم المعزز العميق لإدارة محافظ العملات المشفرة
الملخص
تعرض هذه الورقة إطاراً للتعلم المعزز بلا نموذج لإعادة توزيع المحفظة بين الأصول المالية. وتشمل مكوناته مجموعة من المقيمين المستقلين، وذاكرة لأوزان المحفظة السابقة، والتعلم الدفعي العشوائي عبر الإنترنت، ودالة مكافأة صريحة. ونُفذ الإطار باستخدام شبكات عصبية التفافية ومتكررة وشبكات الذاكرة طويلة قصيرة المدى.
يستخدم التقييم المعلن ثلاثة اختبارات تاريخية على بيانات العملات المشفرة بفترات تداول مدتها 30 دقيقة. ويقارن المؤلفون النماذج باستراتيجيات أخرى لاختيار المحافظ، ويذكرون أن تطبيقاتهم الثلاثة شغلت المراكز الثلاثة الأولى في التجارب، رغم عمولات قدرها 0.25%، مع عوائد لا تقل عن أربعة أضعاف خلال 50 يوماً. هذه نتائج اختبار تاريخي، وليست دليلاً على الأداء المستقبلي. ولا يحدد المقتطف الأصول أو تقسيمات البيانات أو مقاييس المخاطر أو اختبارات المتانة، لذا لا يقدم تفاصيل كافية للحكم على فرط التوافق أو جدوى التداول الفعلي.
الأفكار الرئيسية
- يتعلم الإطار إعادة توزيع المحافظ دون الاعتماد على نموذج مالي صريح للسوق.
- يجمع بين مقيمين مستقلين للأصول وذاكرة لأوزان المحفظة السابقة.
- يشكل التعلم الدفعي العشوائي عبر الإنترنت ودالة المكافأة الصريحة جزءين أساسيين من المنهج.
- يختبر المؤلفون التطبيقات CNN وRNN وLSTM في اختبارات تاريخية للعملات المشفرة.
- لا تثبت مراتب الاختبار التاريخي وعوائده المعلنة الأداء في الأسواق الفعلية.
الوسوم
النص الكامل
# A Deep Reinforcement Learning Framework for the Financial Portfolio Management Problem # A Deep Reinforcement Learning Framework for the Financial Portfolio Management Problem Financial portfolio management is the process of constant redistribution of a fund into different financial products. This paper presents a financial-model-free Reinforcement Learning framework to provide a deep machine learning solution to the portfolio management problem. The framework consists of the Ensemble of Identical Independent Evaluators (EIIE) topology, a Portfolio-Vector Memory (PVM), an Online Stochastic Batch Learning (OSBL) scheme, and a fully exploiting and explicit reward function. This framework is realized in three instants in this work with a Convolutional Neural Network (CNN), a basic Recurrent Neural Network (RNN), and a Long Short-Term Memory (LSTM). They are, along with a number of recently reviewed or published portfolio-selection strategies, examined in three back-test experiments with a trading period of 30 minutes in a cryptocurrency market. Cryptocurrencies are electronic and decentralized alternatives to government-issued money, with Bitcoin as the best-known example of a cryptocurrency. All three instances of the framework monopolize the top three positions in all experiments, outdistancing other compared trading algorithms. Although with a high commission rate of 0.25% in the backtests, the framework is able to achieve at least 4-fold returns in 50 days.
يُعرض النص كاملًا مع نسبه إلى مصدره وفقًا لترخيصه. الترخيص: abstract CC0
أعدّ وكيل الأبحاث في Stratmill هذا الملخص استنادًا إلى المصدر الأصلي؛ وهو ليس نسخة منه.