LSTM価格モデルとPPOによるビットコイン取引
記事 arXiv papers · 著者: Fengrui Liu et al.
サマリー
この研究は、長短期記憶(LSTM)価格モデルと近接方策最適化(PPO)を組み合わせた、自動高頻度ビットコイン取引の枠組みを提案します。価格をエージェントの状態、取引を行動、リターンを報酬として表します。価格予測部分では、サポートベクターマシン、多層パーセプトロン、LSTM、時間畳み込みネットワーク、Transformerを比較し、報告された実験ではLSTMが優位でした。
著者らは、同期されたデータを用い、シミュレーション環境で構築したPPO戦略を評価し、単一資産向けの一般的な戦略ベンチマークと比較しています。最も強いベンチマークを上回るリターンを報告し、相場の変動期と上昇期の双方で利益が得られたとしています。これらの結果は、対象となったビットコイン市場とシミュレーションに限られます。抜粋には取引コスト、市場インパクト、サンプル設計、アウト・オブ・サンプルでの頑健性が記載されていません。したがって、この手法が実運用や他の商品にも通用することを示すものではありません。
主なアイデア
- PPOを用い、価格状態とリターン報酬からビットコイン取引を生成します。
- 複数の予測手法との比較を経て、LSTMが方策の価格モデルの基盤となっています。
- 同期データを用い、シミュレーション環境で戦略を評価しています。
- 報告されたベンチマーク比較では、テストした単一資産の設定で提案手法が優位でした。
- 抜粋には取引コストや実市場での頑健性を示す根拠は記載されていません。
タグ
全文
# Bitcoin Transaction Strategy Construction Based on Deep Reinforcement Learning # Bitcoin Transaction Strategy Construction Based on Deep Reinforcement Learning The emerging cryptocurrency market has lately received great attention for asset allocation due to its decentralization uniqueness. However, its volatility and brand new trading mode have made it challenging to devising an acceptable automatically-generating strategy. This study proposes a framework for automatic high-frequency bitcoin transactions based on a deep reinforcement learning algorithm-proximal policy optimization (PPO). The framework creatively regards the transaction process as actions, returns as awards and prices as states to align with the idea of reinforcement learning. It compares advanced machine learning-based models for static price predictions including support vector machine (SVM), multi-layer perceptron (MLP), long short-term memory (LSTM), temporal convolutional network (TCN), and Transformer by applying them to the real-time bitcoin price and the experimental results demonstrate that LSTM outperforms. Then an automatically-generating transaction strategy is constructed building on PPO with LSTM as the basis to construct the policy. Extensive empirical studies validate that the proposed method performs superiorly to various common trading strategy benchmarks for a single financial product. The approach is able to trade bitcoins in a simulated environment with synchronous data and obtains a 31.67% more return than that of the best benchmark, improving the benchmark by 12.75%. The proposed framework can earn excess returns through both the period of volatility and surge, which opens the door to research on building a single cryptocurrency trading strategy based on deep learning. Visualizations of trading the process show how the model handles high-frequency transactions to provide inspiration and demonstrate that it can be expanded to other financial products.
出典を明記したうえで、ライセンスに従って全文を掲載しています。 ライセンス: abstract CC0
この要約は原文をもとにStratmillのリサーチエージェントが作成したもので、出典の複製ではありません。