コンテンツへスキップ
ライブラリの全資料

リスクを考慮したペアトレード向け再帰型強化学習

記事 arXiv papers · 著者: Weiguang Han et al.

サマリー

この文書は、双方向ゲート付き再帰型ユニット(GRU)と時間的注意機構を使うペアトレード向け強化学習手法、CREDITを説明しています。このアーキテクチャは、二つの資産の価格変動における時間をまたぐ関係を捉えることを意図しており、市場状態を個別に扱うと見逃す可能性のあるパターンも対象としています。また、取引利益とリスクの両方を考慮する報酬を用い、利益と損失の可能性がともに大きい取引を抑制することを目指しています。

著者らは、この手法を、従来の強化学習手法に見られた頻繁な取引、取引コスト、リスク選好への対応として位置づけています。CREDITは既存の強化学習手法を上回り、米国株の5年分のデータを用いた実験で大きな利益を上げたと報告しています。抜粋にはベンチマークの詳細、コストの仮定、リスク指標、数値結果がないため、報告された優位性の大きさや頑健性は評価できません。他の市場や期間での成績も、ここでは示されていません。

主なアイデア

  • CREDITは、二資産のペアトレードで価格履歴に再帰学習と時間的注意機構を適用します。
  • 報酬は利益と取引リスクの両方を考慮しています。
  • 長期的な価格パターンを学習し、過剰な取引を減らすことを目指した設計です。
  • 米国株の5年分のデータを用いた実験では、他の強化学習手法を上回ったと報告されています。
  • 頑健性や実取引コストを評価するための詳細は、抜粋には十分に示されていません。

タグ

全文
# Mastering Pair Trading with Risk-Aware Recurrent Reinforcement Learning


# Mastering Pair Trading with Risk-Aware Recurrent Reinforcement Learning









Although pair trading is the simplest hedging strategy for an investor to eliminate market risk, it is still a great challenge for reinforcement learning (RL) methods to perform pair trading as human expertise. It requires RL methods to make thousands of correct actions that nevertheless have no obvious relations to the overall trading profit, and to reason over infinite states of the time-varying market most of which have never appeared in history. However, existing RL methods ignore the temporal connections between asset price movements and the risk of the performed trading. These lead to frequent tradings with high transaction costs and potential losses, which barely reach the human expertise level of trading. Therefore, we introduce CREDIT, a risk-aware agent capable of learning to exploit long-term trading opportunities in pair trading similar to a human expert. CREDIT is the first to apply bidirectional GRU along with the temporal attention mechanism to fully consider the temporal correlations embedded in the states, which allows CREDIT to capture long-term patterns of the price movements of two assets to earn higher profit. We also design the risk-aware reward inspired by the economic theory, that models both the profit and risk of the tradings during the trading period. It helps our agent to master pair trading with a robust trading preference that avoids risky trading with possible high returns and losses. Experiments show that it outperforms existing reinforcement learning methods in pair trading and achieves a significant profit over five years of U.S. stock data.

出典を明記したうえで、ライセンスに従って全文を掲載しています。 ライセンス: abstract CC0

この要約は原文をもとにStratmillのリサーチエージェントが作成したもので、出典の複製ではありません。