Deep Reinforcement Learning for Cryptocurrency Portfolio Management
Summary
This paper presents a model-free reinforcement learning framework for reallocating a portfolio among financial assets. Its components include an ensemble of independent evaluators, memory of prior portfolio weights, online stochastic batch learning, and an explicit reward function. The framework is implemented with convolutional, recurrent, and long short-term memory neural networks.
The reported evaluation uses three backtests on cryptocurrency data with 30-minute trading periods. The authors compare the models with other portfolio selection strategies and report that their three implementations occupy the top three positions across the experiments, despite commissions of 0.25%, with at least fourfold returns over 50 days. These are historical backtest results, not evidence of future performance. The excerpt does not identify the assets, data splits, risk measures, or robustness checks, so it offers limited detail for judging overfitting or live-trading feasibility.
Key ideas
- The framework learns portfolio reallocations without relying on an explicit financial market model.
- It combines independent asset evaluators with memory of earlier portfolio weights.
- Online stochastic batch learning and an explicit reward function are central parts of the method.
- The authors test CNN, RNN, and LSTM implementations in cryptocurrency backtests.
- The reported backtest rankings and returns do not establish live-market performance.
Tags
Full text
# A Deep Reinforcement Learning Framework for the Financial Portfolio Management Problem # A Deep Reinforcement Learning Framework for the Financial Portfolio Management Problem Financial portfolio management is the process of constant redistribution of a fund into different financial products. This paper presents a financial-model-free Reinforcement Learning framework to provide a deep machine learning solution to the portfolio management problem. The framework consists of the Ensemble of Identical Independent Evaluators (EIIE) topology, a Portfolio-Vector Memory (PVM), an Online Stochastic Batch Learning (OSBL) scheme, and a fully exploiting and explicit reward function. This framework is realized in three instants in this work with a Convolutional Neural Network (CNN), a basic Recurrent Neural Network (RNN), and a Long Short-Term Memory (LSTM). They are, along with a number of recently reviewed or published portfolio-selection strategies, examined in three back-test experiments with a trading period of 30 minutes in a cryptocurrency market. Cryptocurrencies are electronic and decentralized alternatives to government-issued money, with Bitcoin as the best-known example of a cryptocurrency. All three instances of the framework monopolize the top three positions in all experiments, outdistancing other compared trading algorithms. Although with a high commission rate of 0.25% in the backtests, the framework is able to achieve at least 4-fold returns in 50 days.
Shown in full with attribution under the source's licence. Licence: abstract CC0
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.