跳至正文
返回文库全部文档

MacroHFT:用于加密货币交易的情境感知强化学习

文章 arXiv papers · 作者: Chuqiao Zong et al.

总结

MacroHFT 是一种用于分钟级加密货币交易的强化学习方法。它针对现有方法中提出的两项挑战:智能体可能过拟合,且无法根据金融环境调整策略;市场状况快速变化时,单个智能体的决策也可能产生偏差。该方法使用包括趋势和波动率在内的市场指标来组织数据,并训练多个专业化子智能体,每个子智能体都配有适配器,以根据当前状况调整行为。

第二阶段训练加入一个超智能体,用于整合各子智能体的决策。记忆机制支持这一更高层级的决策过程,使其能够应对市场波动。据报告,在多个加密货币市场上的实验显示,该方法在分钟级任务中达到最先进表现。摘要未说明具体资产、评估时期、基准、成本或稳健性测试,因此该结果本身不能证明其具有实盘盈利能力或可推广性。

核心观点

  • MacroHFT 使用多个专业化强化学习智能体进行加密货币交易。
  • 市场趋势和波动率用于组织训练数据和智能体分工。
  • 条件适配器使子智能体能够根据市场状况调整策略。
  • 配备记忆机制的超智能体整合各子智能体的决策。
  • 报告的实验涉及分钟级任务,而描述中提供的评估细节有限。

标签

全文
# MacroHFT: Memory Augmented Context-aware Reinforcement Learning On High Frequency Trading


# MacroHFT: Memory Augmented Context-aware Reinforcement Learning On High Frequency Trading









High-frequency trading (HFT) that executes algorithmic trading in short time scales, has recently occupied the majority of cryptocurrency market. Besides traditional quantitative trading methods, reinforcement learning (RL) has become another appealing approach for HFT due to its terrific ability of handling high-dimensional financial data and solving sophisticated sequential decision-making problems, \emph{e.g.,} hierarchical reinforcement learning (HRL) has shown its promising performance on second-level HFT by training a router to select only one sub-agent from the agent pool to execute the current transaction. However, existing RL methods for HFT still have some defects: 1) standard RL-based trading agents suffer from the overfitting issue, preventing them from making effective policy adjustments based on financial context; 2) due to the rapid changes in market conditions, investment decisions made by an individual agent are usually one-sided and highly biased, which might lead to significant loss in extreme markets. To tackle these problems, we propose a novel Memory Augmented Context-aware Reinforcement learning method On HFT, \emph{a.k.a.} MacroHFT, which consists of two training phases: 1) we first train multiple types of sub-agents with the market data decomposed according to various financial indicators, specifically market trend and volatility, where each agent owns a conditional adapter to adjust its trading policy according to market conditions; 2) then we train a hyper-agent to mix the decisions from these sub-agents and output a consistently profitable meta-policy to handle rapid market fluctuations, equipped with a memory mechanism to enhance the capability of decision-making. Extensive experiments on various cryptocurrency markets demonstrate that MacroHFT can achieve state-of-the-art performance on minute-level trading tasks.

在遵守原作品许可的前提下,附作者信息全文展示。 许可协议: abstract CC0

此摘要由 Stratmill 研究智能体根据原文撰写,并非原文副本。