コンテンツへスキップ
ライブラリの全資料

潜在状態を伴う取引のための再帰型強化学習

記事 arXiv papers · 著者: Andrea Macrì et al.

サマリー

シグナルが平均回帰型のオルンシュタイン・ウーレンベック過程に従い、そのパラメーターが隠れた状態に応じて変化する場合の最適取引を研究しています。再帰型ニューラルネットワークと強化学習を組み合わせ、観測値からシグナルの潜在平均、平均回帰速度、ボラティリティに関する情報を推定し、その情報を取引に活用します。

著者らはDDPGに基づく三つの手法を提案しています。一つは再帰型の隠れ状態を取引エージェントに渡し、残る二つは推定した状態確率または次のシグナル値の予測を与えます。状態変化の複雑さを段階的に高めたシミュレーションと、実データによる株式ペアトレードへの適用では、累積報酬と解釈可能性の点で確率に基づく手法が優れていました。予測を使う方法の追加効果は限定的で、隠れ状態を使う方法は両者の中間でした。これらの結果は、推定情報の表現方法が重要であることを示唆しますが、根拠は指定されたシミュレーションと適用例に限られ、他市場や導入条件での成績を示すものではありません。

主なアイデア

  • 取引シグナルを平均回帰型とし、隠れた状態に応じてパラメーターが変化するモデルを用いています。
  • 再帰型ネットワークがシグナルの観測値から時間的な情報を抽出します。
  • 三つのDDPG設計は、隠れ状態、状態確率、またはシグナル予測をエージェントに提供します。
  • 報告されたシミュレーションとペアトレード適用例では、状態確率を使う手法が最良でした。
  • エージェントに与える情報が、報告された報酬と戦略の解釈可能性に影響します。

タグ

全文
# Deep reinforcement learning for optimal trading with partial information


# Deep reinforcement learning for optimal trading with partial information









Reinforcement Learning (RL) applied to financial problems has been the subject of a lively area of research. The use of RL for optimal trading strategies that exploit latent information in the market is, to the best of our knowledge, not widely tackled. In this paper we study an optimal trading problem, where a trading signal follows an Ornstein-Uhlenbeck process with regime-switching dynamics. We employ a blend of RL and Recurrent Neural Networks (RNN) in order to make the most at extracting underlying information from the trading signal with latent parameters. The latent parameters driving mean reversion, speed, and volatility are filtered from observations of the signal, and trading strategies are derived via RL. To address this problem, we propose three Deep Deterministic Policy Gradient (DDPG)-based algorithms that integrate Gated Recurrent Unit (GRU) networks to capture temporal dependencies in the signal. The first, a one -step approach (hid-DDPG), directly encodes hidden states from the GRU into the RL trader. The second and third are two-step methods: one (prob-DDPG) makes use of posterior regime probability estimates, while the other (reg-DDPG) relies on forecasts of the next signal value. Through extensive simulations with increasingly complex Markovian regime dynamics for the trading signal's parameters, as well as an empirical application to equity pair trading, we find that prob-DDPG achieves superior cumulative rewards and exhibits more interpretable strategies. By contrast, reg-DDPG provides limited benefits, while hid-DDPG offers intermediate performance with less interpretable strategies. Our results show that the quality and structure of the information supplied to the agent are crucial: embedding probabilistic insights into latent regimes substantially improves both profitability and robustness of reinforcement learning-based trading strategies.

出典を明記したうえで、ライセンスに従って全文を掲載しています。 ライセンス: abstract CC0

この要約は原文をもとにStratmillのリサーチエージェントが作成したもので、出典の複製ではありません。