コンテンツへスキップ
ライブラリの全資料

MacroHFT:暗号資産取引の状況に応じた強化学習

記事 arXiv papers · 著者: Chuqiao Zong et al.

サマリー

MacroHFTは、暗号資産の分単位取引に向けた強化学習手法です。既存手法に関する2つの課題に取り組んでいます。エージェントは過学習して金融環境に方策を適応できないことがあり、単一エージェントの判断は市場環境が急変すると偏る可能性があります。この手法では、トレンドやボラティリティなどの市場指標を使ってデータを整理し、複数の専門化したサブエージェントを訓練します。それぞれのアダプターが、現在の環境に合わせて行動を調整します。

第2段階の訓練では、サブエージェントの判断を統合するハイパーエージェントを追加します。記憶機構は、市場変動に応じて判断するこの上位の意思決定を支えます。暗号資産市場での実験では、分単位の課題で最先端の性能を示したと報告されています。要約には資産、評価期間、ベースライン、コスト、頑健性検証が記載されていないため、この結果だけでは実運用での収益性や一般化可能性は立証されません。

主なアイデア

  • MacroHFTは、暗号資産取引に複数の専門化した強化学習エージェントを用います。
  • 市場トレンドとボラティリティを使って、訓練データとエージェントの役割を整理します。
  • 条件付きアダプターによって、サブエージェントの方策を市場環境に合わせて調整できます。
  • 記憶機構を備えたハイパーエージェントが、サブエージェントの判断を統合します。
  • 報告された実験は分単位の課題を対象としており、説明にある評価の詳細は限られています。

タグ

全文
# MacroHFT: Memory Augmented Context-aware Reinforcement Learning On High Frequency Trading


# MacroHFT: Memory Augmented Context-aware Reinforcement Learning On High Frequency Trading









High-frequency trading (HFT) that executes algorithmic trading in short time scales, has recently occupied the majority of cryptocurrency market. Besides traditional quantitative trading methods, reinforcement learning (RL) has become another appealing approach for HFT due to its terrific ability of handling high-dimensional financial data and solving sophisticated sequential decision-making problems, \emph{e.g.,} hierarchical reinforcement learning (HRL) has shown its promising performance on second-level HFT by training a router to select only one sub-agent from the agent pool to execute the current transaction. However, existing RL methods for HFT still have some defects: 1) standard RL-based trading agents suffer from the overfitting issue, preventing them from making effective policy adjustments based on financial context; 2) due to the rapid changes in market conditions, investment decisions made by an individual agent are usually one-sided and highly biased, which might lead to significant loss in extreme markets. To tackle these problems, we propose a novel Memory Augmented Context-aware Reinforcement learning method On HFT, \emph{a.k.a.} MacroHFT, which consists of two training phases: 1) we first train multiple types of sub-agents with the market data decomposed according to various financial indicators, specifically market trend and volatility, where each agent owns a conditional adapter to adjust its trading policy according to market conditions; 2) then we train a hyper-agent to mix the decisions from these sub-agents and output a consistently profitable meta-policy to handle rapid market fluctuations, equipped with a memory mechanism to enhance the capability of decision-making. Extensive experiments on various cryptocurrency markets demonstrate that MacroHFT can achieve state-of-the-art performance on minute-level trading tasks.

出典を明記したうえで、ライセンスに従って全文を掲載しています。 ライセンス: abstract CC0

この要約は原文をもとにStratmillのリサーチエージェントが作成したもので、出典の複製ではありません。