暗号資産ペアトレードのための深層強化学習
記事 arXiv papers · 著者: Damian Lebiedź et al.
サマリー
本研究では、変動の大きい暗号資産市場において、ペアトレードの執行を補助する仕組みとして深層強化学習を検証します。まず候補ペアを絞り込み、順位付けし、次にProximal Policy OptimizationエージェントとLong Short-Term Memory層を用いて執行を判断します。固定リスクと適応平均を組み合わせた設計および決定論的なリスク管理により、乖離リスクの低減を目指してエージェントを制約します。評価には時間足のBinance USD-M先物データを用い、学習済み方策をヒューリスティックなベースラインと比較します。
報告されたアウト・オブ・サンプル結果では強化学習方策が優位であり、定常円形ブロック・ブートストラップは、10パーセント水準でリスク調整後の超過成績が統計的に有意であることを示しています。ただし、この結果はより厳しい5パーセントの有意水準には達しておらず、研究ではその理由をデジタル資産特有の分散の大きさと関連付けています。証拠は検証したデータと設定に限られ、この説明には成績の大きさ、取引コストの仮定、他の取引所や期間での検証は示されていません。このハイブリッド設計は学習済みの執行判断を制約する方法を示しますが、より広い頑健性は未解決です。
主なアイデア
- 統計的なペア選択と、学習した執行補助を組み合わせたシステムを提案します。
- PPOエージェントがLSTMを用い、決定論的なリスク制約の範囲内で執行を判断します。
- 時間足のBinance USD-M先物データを用い、ヒューリスティックなベースラインと比較します。
- 報告されたリスク調整後の超過成績は10パーセント水準で有意ですが、5パーセント水準では有意ではありません。
- 結果は、検証した市場、データ、戦略設定に固有です。
タグ
全文
# Dynamic Multi-Pair Trading Strategy in Cryptocurrency Markets with Deep Reinforcement Learning # Dynamic Multi-Pair Trading Strategy in Cryptocurrency Markets with Deep Reinforcement Learning This study aims to determine whether the application of Deep Reinforcement Learning (DRL) as a specialized execution overlay can enhance pair trading in highly volatile cryptocurrency markets. Although classical implementations of the strategy have proven successful in traditional equities, they frequently exhibit rigidity and suffer from severe divergence risks when applied to high-variance environments. To address this need, this research introduces novel concepts. To construct a robust system, we developed a hierarchical "Filter-then-Rank" pair selection methodology and a proprietary "Fixed Risk, Adaptive Mean" execution model. The system employs a Proximal Policy Optimization (PPO) agent with a Long Short-Term Memory (LSTM) layer to govern execution decisions within strict deterministic risk management boundaries. Evaluated on 1-hour interval data from the Binance USD-M Futures market, the optimized RL policy achieved an out-of-sample performance that substantially outperformed the heuristic baseline. A stationary circular block bootstrap robustness check confirms that the agent's risk-adjusted outperformance is statistically significant at the 10 percent level. Although falling marginally short of the stricter 5 percent threshold, this result highlights the extreme idiosyncratic variance characteristic of digital assets. Ultimately, this thesis contributes to the quantitative finance literature by introducing a hybrid architecture that combines statistical arbitrage with DRL execution policies. Furthermore, it delivers a novel framework for safe reinforcement learning via deterministic shielding, proving that anchoring a neural policy to statistically robust boundaries successfully mitigates severe divergence risks.
出典を明記したうえで、ライセンスに従って全文を掲載しています。 ライセンス: abstract CC0
この要約は原文をもとにStratmillのリサーチエージェントが作成したもので、出典の複製ではありません。