跳至正文
返回文库全部文档

利用LSTM价格模型和PPO进行比特币交易

文章 arXiv papers · 作者: Fengrui Liu et al.

总结

该研究提出一种自动化比特币高频交易框架,将长短期记忆(LSTM)价格模型与近端策略优化(PPO)相结合。框架将价格表示为智能体的状态,将交易表示为动作,将回报表示为奖励。在价格预测部分,研究比较了支持向量机、多层感知机、LSTM、时序卷积网络和Transformer;报告的实验结果更支持LSTM。

作者在使用同步数据的模拟环境中评估由此得到的PPO策略,并将其与单一资产的常见策略基准进行比较。他们报告称,该策略的回报高于表现最好的基准,并指出在市场波动和上涨时期均有收益。这些结果仅适用于该研究中的比特币场景和模拟环境;摘录未说明交易成本、市场冲击、样本设计或样本外稳健性。因此,其结论并不能证明该方法可推广到实盘交易或其他产品。

核心观点

  • 该框架使用PPO,根据价格状态和回报奖励生成比特币交易。
  • 比较多种其他预测方法后,使用LSTM为策略提供价格模型基础。
  • 该研究在使用同步数据的模拟环境中评估策略。
  • 在所测试的单一资产场景中,报告的基准比较结果有利于该方法。
  • 摘录未说明交易成本,也未提供该方法在实盘市场中具备稳健性的证据。

标签

全文
# Bitcoin Transaction Strategy Construction Based on Deep Reinforcement Learning


# Bitcoin Transaction Strategy Construction Based on Deep Reinforcement Learning









The emerging cryptocurrency market has lately received great attention for asset allocation due to its decentralization uniqueness. However, its volatility and brand new trading mode have made it challenging to devising an acceptable automatically-generating strategy. This study proposes a framework for automatic high-frequency bitcoin transactions based on a deep reinforcement learning algorithm-proximal policy optimization (PPO). The framework creatively regards the transaction process as actions, returns as awards and prices as states to align with the idea of reinforcement learning. It compares advanced machine learning-based models for static price predictions including support vector machine (SVM), multi-layer perceptron (MLP), long short-term memory (LSTM), temporal convolutional network (TCN), and Transformer by applying them to the real-time bitcoin price and the experimental results demonstrate that LSTM outperforms. Then an automatically-generating transaction strategy is constructed building on PPO with LSTM as the basis to construct the policy. Extensive empirical studies validate that the proposed method performs superiorly to various common trading strategy benchmarks for a single financial product. The approach is able to trade bitcoins in a simulated environment with synchronous data and obtains a 31.67% more return than that of the best benchmark, improving the benchmark by 12.75%. The proposed framework can earn excess returns through both the period of volatility and surge, which opens the door to research on building a single cryptocurrency trading strategy based on deep learning. Visualizations of trading the process show how the model handles high-frequency transactions to provide inspiration and demonstrate that it can be expanded to other financial products.

在遵守原作品许可的前提下,附作者信息全文展示。 许可协议: abstract CC0

此摘要由 Stratmill 研究智能体根据原文撰写,并非原文副本。