یادگیری تقویتی عمیق برای مدیریت سبد رمزارز
خلاصه
این مقاله چارچوبی برای یادگیری تقویتیِ بدون مدل ارائه میکند که سبد را میان داراییهای مالی بازتخصیص میدهد. اجزای آن شامل مجموعهای از ارزیابهای مستقل، حافظه وزنهای قبلی سبد، یادگیری دستهای تصادفی برخط و تابع پاداش صریح است. این چارچوب با شبکههای عصبی کانولوشنی، بازگشتی و حافظه بلندکوتاهمدت پیادهسازی شده است.
ارزیابی گزارششده شامل سه بکتست روی دادههای رمزارز با دورههای معاملاتی 30 دقیقهای است. نویسندگان مدلها را با استراتژیهای دیگر انتخاب سبد مقایسه میکنند و گزارش میدهند که سه پیادهسازی آنها، با وجود کارمزدهای 0.25%، در مجموع آزمایشها سه جایگاه نخست را میگیرند و بازدهی دستکم چهاربرابری طی 50 روز دارند. اینها نتایج بکتست تاریخیاند و عملکرد آینده را اثبات نمیکنند. بخش ارائهشده داراییها، تفکیک دادهها، سنجههای ریسک یا آزمونهای استحکام را مشخص نمیکند؛ بنابراین جزئیات محدودی برای ارزیابی بیشبرازش یا امکانپذیری معامله زنده دارد.
ایدههای کلیدی
- چارچوب بدون اتکا به مدل صریح بازار مالی، بازتخصیص سبد را میآموزد.
- این روش ارزیابهای مستقل دارایی را با حافظه وزنهای قبلی سبد ترکیب میکند.
- یادگیری دستهای تصادفی برخط و تابع پاداش صریح، بخشهای اصلی روشاند.
- نویسندگان CNN، RNN و LSTM پیادهسازی را در بکتستهای رمزارز میآزمایند.
- رتبهبندیها و بازده گزارششده در بکتست، عملکرد در بازار زنده را اثبات نمیکنند.
برچسبها
متن کامل
# A Deep Reinforcement Learning Framework for the Financial Portfolio Management Problem # A Deep Reinforcement Learning Framework for the Financial Portfolio Management Problem Financial portfolio management is the process of constant redistribution of a fund into different financial products. This paper presents a financial-model-free Reinforcement Learning framework to provide a deep machine learning solution to the portfolio management problem. The framework consists of the Ensemble of Identical Independent Evaluators (EIIE) topology, a Portfolio-Vector Memory (PVM), an Online Stochastic Batch Learning (OSBL) scheme, and a fully exploiting and explicit reward function. This framework is realized in three instants in this work with a Convolutional Neural Network (CNN), a basic Recurrent Neural Network (RNN), and a Long Short-Term Memory (LSTM). They are, along with a number of recently reviewed or published portfolio-selection strategies, examined in three back-test experiments with a trading period of 30 minutes in a cryptocurrency market. Cryptocurrencies are electronic and decentralized alternatives to government-issued money, with Bitcoin as the best-known example of a cryptocurrency. All three instances of the framework monopolize the top three positions in all experiments, outdistancing other compared trading algorithms. Although with a high commission rate of 0.25% in the backtests, the framework is able to achieve at least 4-fold returns in 50 days.
با ذکر منبع و مطابق مجوز اثر، بهطور کامل نمایش داده میشود. مجوز: abstract CC0
این خلاصه را عامل پژوهشی Stratmill بر پایه متن اصلی نوشته است؛ نسخهای از اثر منبع نیست.