跳至正文
返回文库全部文档

用于动态加密货币配对交易的深度强化学习

文章 arXiv papers · 作者: Damian Lebiedź et al.

总结

这项研究测试了深度强化学习作为执行层,用于波动较大的加密货币市场中的配对交易。其框架先筛选并排列候选配对,再使用带有长短期记忆层的近端策略优化智能体做出执行决策。固定风险、自适应均值的设计和确定性风险控制约束该智能体,旨在降低偏离风险。评估使用 Binance 每小时USD-M期货数据,并将学习得到的策略与启发式基线进行比较。

报告的样本外结果倾向于强化学习策略;平稳循环分块自助法表明,其风险调整后超额表现在10%水平上具有统计显著性。但结果未达到更严格的5%显著性阈值,研究将此归因于数字资产的特质方差较高。证据仅适用于测试数据和设定;描述没有提供表现幅度、交易成本假设,也没有跨其他交易场所和时期进行测试。混合设计提供了一种约束学习式执行决策的方法,但其更广泛的稳健性仍有待验证。

核心观点

  • 该系统将统计配对选择与学习式执行层结合起来。
  • 一个PPO智能体通过LSTM在确定性风险边界内做出执行决策。
  • 该研究使用 Binance 每小时USD-M期货数据,将此方法与启发式基线进行评估。
  • 报告的风险调整后超额表现达到10%水平的显著性,但未达到5%水平。
  • 结果仅适用于经过测试的市场、数据和策略配置。

标签

全文
# Dynamic Multi-Pair Trading Strategy in Cryptocurrency Markets with Deep Reinforcement Learning


# Dynamic Multi-Pair Trading Strategy in Cryptocurrency Markets with Deep Reinforcement Learning









This study aims to determine whether the application of Deep Reinforcement Learning (DRL) as a specialized execution overlay can enhance pair trading in highly volatile cryptocurrency markets. Although classical implementations of the strategy have proven successful in traditional equities, they frequently exhibit rigidity and suffer from severe divergence risks when applied to high-variance environments. To address this need, this research introduces novel concepts. To construct a robust system, we developed a hierarchical "Filter-then-Rank" pair selection methodology and a proprietary "Fixed Risk, Adaptive Mean" execution model. The system employs a Proximal Policy Optimization (PPO) agent with a Long Short-Term Memory (LSTM) layer to govern execution decisions within strict deterministic risk management boundaries. Evaluated on 1-hour interval data from the Binance USD-M Futures market, the optimized RL policy achieved an out-of-sample performance that substantially outperformed the heuristic baseline. A stationary circular block bootstrap robustness check confirms that the agent's risk-adjusted outperformance is statistically significant at the 10 percent level. Although falling marginally short of the stricter 5 percent threshold, this result highlights the extreme idiosyncratic variance characteristic of digital assets. Ultimately, this thesis contributes to the quantitative finance literature by introducing a hybrid architecture that combines statistical arbitrage with DRL execution policies. Furthermore, it delivers a novel framework for safe reinforcement learning via deterministic shielding, proving that anchoring a neural policy to statistically robust boundaries successfully mitigates severe divergence risks.

在遵守原作品许可的前提下,附作者信息全文展示。 许可协议: abstract CC0

此摘要由 Stratmill 研究智能体根据原文撰写,并非原文副本。