コンテンツへスキップ
ライブラリの全資料

アクティブな高頻度株式取引のためのPPO強化学習

記事 arXiv papers · 著者: Antonio Briola et al.

サマリー

この文書は、アクティブな高頻度株式取引に対する、エンドツーエンドの深層強化学習アプローチを説明します。エージェントは近接方策最適化を使い、異なる指値注文板特徴量から構成した状態に基づいて、Intel株を1単位取引します。シグナル対ノイズ比を高めるため、大きな価格変動があったデータを訓練用に選び、連続する1か月を検証用に確保し、その後、別の1か月をテストに使います。ハイパーパラメーターは逐次モデルベース最適化で調整します。

著者らは、エージェントが注文板に繰り返し現れるパターンを見つけ、ノイズや変化のある状況でもテスト期間に安定した正のリターンを出したと報告しています。根拠となるのは、1銘柄を数か月にわたって調べた、範囲の限られた実験です。説明にはリターン統計、取引コスト、ベンチマークとの比較、学習した行動が他の期間や資産、実運用でも維持されることを示す根拠はありません。価格変動に基づいて訓練サンプルを選ぶことで学習内容が左右される可能性もあるため、報告結果を広範な収益性の証明と捉えるべきではありません。

主なアイデア

  • 近接方策最適化を使い、単一銘柄を1単位取引します。
  • 状態表現ごとに、含める指値注文板特徴量が異なります。
  • 大きな価格変動があったサンプルを重視し、検証期間とテスト期間を別に設けています。
  • ハイパーパラメーターは逐次モデルベース最適化で調整します。
  • テスト期間の正のリターンが報告されていますが、より広い一般化や執行コストは明らかではありません。

タグ

全文
# Deep Reinforcement Learning for Active High Frequency Trading


# Deep Reinforcement Learning for Active High Frequency Trading









We introduce the first end-to-end Deep Reinforcement Learning (DRL) based framework for active high frequency trading in the stock market. We train DRL agents to trade one unit of Intel Corporation stock by employing the Proximal Policy Optimization algorithm. The training is performed on three contiguous months of high frequency Limit Order Book data, of which the last month constitutes the validation data. In order to maximise the signal to noise ratio in the training data, we compose the latter by only selecting training samples with largest price changes. The test is then carried out on the following month of data. Hyperparameters are tuned using the Sequential Model Based Optimization technique. We consider three different state characterizations, which differ in their LOB-based meta-features. Analysing the agents' performances on test data, we argue that the agents are able to create a dynamic representation of the underlying environment. They identify occasional regularities present in the data and exploit them to create long-term profitable trading strategies. Indeed, agents learn trading strategies able to produce stable positive returns in spite of the highly stochastic and non-stationary environment.

出典を明記したうえで、ライセンスに従って全文を掲載しています。 ライセンス: abstract CC0

この要約は原文をもとにStratmillのリサーチエージェントが作成したもので、出典の複製ではありません。