コンテンツへスキップ
ライブラリの全資料

投機取引とペアトレードのための探索的強化学習

記事 arXiv papers · 著者: Yun Zhao et al.

サマリー

本研究では、投機取引を逐次的な最適停止問題として定式化し、一般的な効用関数と価格過程のもとでエントリー時点と決済時点を選びます。まず、エントリーと決済をコックス過程のジャンプとして表し、停止問題を緩和します。上限のある強度制御はエージェントが選択します。

探索的強化学習の定式化では、ランダム化方策がこれらの強度に確率分布を割り当て、シャノン微分エントロピーで目的関数を正則化します。そこから得られる探索的ハミルトン・ヤコビ・ベルマン方程式は、ギブス形式の最適方策を導きます。著者らは誤差評価と、学習目的関数が元の問題の価値関数に収束することを示したうえで、ペアトレードに応用したアルゴリズムを実演します。説明にはその応用のデータや成績結果がなく、ここでいう収束はモデル化された目的関数に関するもので、実運用の成果を示すものではありません。

主なアイデア

  • 投機取引を、エントリー時点と決済時点を選ぶ最適停止問題としてモデル化します。
  • 緩和した定式化では、停止イベントを上限付き強度で制御するコックス過程のジャンプとして表します。
  • ランダム化された強度方策をシャノン微分エントロピーで正則化します。
  • この枠組みから、探索的なHJB方程式とギブス形式の最適方策を導きます。
  • 誤差と収束の結果を示し、ペアトレード向けのアルゴリズムを実演しています。

タグ

全文
# Reinforcement Learning for Speculative Trading under Exploratory Framework


# Reinforcement Learning for Speculative Trading under Exploratory Framework









We study a speculative trading problem within the exploratory reinforcement learning (RL) framework of Wang et al. [2020]. The problem is formulated as a sequential optimal stopping problem over entry and exit times under general utility function and price process. We first consider a relaxed version of the problem in which the stopping times are modeled by the jump times of Cox processes driven by bounded, non-randomized intensity controls. Under the exploratory formulation, the agent's randomized control is characterized via the probability measure over the jump intensities, and their objective function is regularized by Shannon's differential entropy. This yields a system of the exploratory HJB equations and Gibbs distributions in closed-form as the optimal policy. Error estimates and convergence of the RL objective to the value function of the original problem are established. Finally, an RL algorithm is designed, and its implementation is showcased in a pairs-trading application.

出典を明記したうえで、ライセンスに従って全文を掲載しています。 ライセンス: abstract CC0

この要約は原文をもとにStratmillのリサーチエージェントが作成したもので、出典の複製ではありません。