Skip to content
All library documents

Bitcoin Trading with LSTM Price Models and PPO

Article arXiv papers · Author: Fengrui Liu et al.

Summary

The study proposes an automated high-frequency Bitcoin trading framework that combines a long short-term memory (LSTM) price model with proximal policy optimization (PPO). It represents prices as the agent's states, trades as actions, and returns as rewards. For the price-prediction component, it compares support vector machines, multilayer perceptrons, LSTMs, temporal convolutional networks, and Transformers; the reported experiments favor LSTM.

The authors evaluate the resulting PPO strategy in a simulated environment using synchronized data and compare it with common strategy benchmarks for a single asset. They report higher returns than the strongest benchmark and describe gains during both volatile and rising periods. These results are limited to the study's Bitcoin setting and simulation; the excerpt does not specify transaction costs, market impact, sample design, or out-of-sample robustness. Its claims therefore do not establish that the approach will transfer to live trading or other products.

Key ideas

  • The framework uses PPO to generate Bitcoin trades from price states and return rewards.
  • An LSTM provides the policy's price-model basis after comparison with several other prediction methods.
  • The study evaluates the strategy in a simulated environment with synchronized data.
  • Reported benchmark comparisons favor the proposed method in the tested single-asset setting.
  • The excerpt does not describe trading costs or evidence of live-market robustness.

Tags

Full text
# Bitcoin Transaction Strategy Construction Based on Deep Reinforcement Learning


# Bitcoin Transaction Strategy Construction Based on Deep Reinforcement Learning









The emerging cryptocurrency market has lately received great attention for asset allocation due to its decentralization uniqueness. However, its volatility and brand new trading mode have made it challenging to devising an acceptable automatically-generating strategy. This study proposes a framework for automatic high-frequency bitcoin transactions based on a deep reinforcement learning algorithm-proximal policy optimization (PPO). The framework creatively regards the transaction process as actions, returns as awards and prices as states to align with the idea of reinforcement learning. It compares advanced machine learning-based models for static price predictions including support vector machine (SVM), multi-layer perceptron (MLP), long short-term memory (LSTM), temporal convolutional network (TCN), and Transformer by applying them to the real-time bitcoin price and the experimental results demonstrate that LSTM outperforms. Then an automatically-generating transaction strategy is constructed building on PPO with LSTM as the basis to construct the policy. Extensive empirical studies validate that the proposed method performs superiorly to various common trading strategy benchmarks for a single financial product. The approach is able to trade bitcoins in a simulated environment with synchronous data and obtains a 31.67% more return than that of the best benchmark, improving the benchmark by 12.75%. The proposed framework can earn excess returns through both the period of volatility and surge, which opens the door to research on building a single cryptocurrency trading strategy based on deep learning. Visualizations of trading the process show how the model handles high-frequency transactions to provide inspiration and demonstrate that it can be expanded to other financial products.

Shown in full with attribution under the source's licence. Licence: abstract CC0

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.