コンテンツへスキップ
ライブラリの全資料

予測シグナルを用いた指値注文執行の深層強化学習

記事 arXiv papers · 著者: Peer Nagy et al.

サマリー

この研究では、深層強化学習を用いて短期売買予測を、指値注文板における指値注文の配置と在庫管理の判断に変換します。過去の注文板メッセージを用いて構築した、シミュレーション上のNASDAQ株式環境でエージェントを訓練し、ABIDESシミュレーターを使用します。エージェントは現在の注文板、直近の市場履歴、方向性シグナルを観測し、Deep Double Duelling Q-learningと優先度付き非同期経験再生によって執行方策を学習します。

特定の予測手法とは切り離して執行を評価するため、研究者らはノイズの程度が異なる合成アルファシグナルを検証します。同じシグナルを与えた場合、学習済みエージェントは在庫を管理して効果的に注文を配置し、ヒューリスティック戦略を上回ると報告されています。この比較は記述されたシミュレーター内での執行手法を支持しますが、文書には成績の数値や実市場での結果を示す証拠はありません。そのため、合成シグナルとシミュレーション環境の質や現実性が、実運用についての結論を制限します。

主なアイデア

  • エージェントは方向性予測を指値注文と在庫の判断に変換することを学習します。
  • 訓練には、過去のメッセージに基づくNASDAQ指値注文板シミュレーターを使います。
  • エージェントは注文板の状態、直近の履歴、短期予測を観測します。
  • ノイズの異なる合成シグナルにより、予測の質ではなく執行に焦点を当てて評価できます。
  • シミュレーションでは学習済み方策が同じシグナルを用いるヒューリスティックを上回りますが、実市場での成績は確認されていません。

タグ

全文
# Asynchronous Deep Double Duelling Q-Learning for Trading-Signal Execution in Limit Order Book Markets


# Asynchronous Deep Double Duelling Q-Learning for Trading-Signal Execution in Limit Order Book Markets









We employ deep reinforcement learning (RL) to train an agent to successfully translate a high-frequency trading signal into a trading strategy that places individual limit orders. Based on the ABIDES limit order book simulator, we build a reinforcement learning OpenAI gym environment and utilise it to simulate a realistic trading environment for NASDAQ equities based on historic order book messages. To train a trading agent that learns to maximise its trading return in this environment, we use Deep Duelling Double Q-learning with the APEX (asynchronous prioritised experience replay) architecture. The agent observes the current limit order book state, its recent history, and a short-term directional forecast. To investigate the performance of RL for adaptive trading independently from a concrete forecasting algorithm, we study the performance of our approach utilising synthetic alpha signals obtained by perturbing forward-looking returns with varying levels of noise. Here, we find that the RL agent learns an effective trading strategy for inventory management and order placing that outperforms a heuristic benchmark trading strategy having access to the same signal.

出典を明記したうえで、ライセンスに従って全文を掲載しています。 ライセンス: abstract CC0

この要約は原文をもとにStratmillのリサーチエージェントが作成したもので、出典の複製ではありません。