ペア取引の階層型言語モデル強化学習
記事 arXiv papers · 著者: Polydoros Giannouris et al.
サマリー
本研究は、ペア取引を二つの段階が連動する階層的な意思決定問題として捉えています。長期の時間軸で資産ペアを選び、部分的にしか観測できない状況で、より短い時間軸の取引を執行します。フィードバックが遅れて届き、曖昧な場合もあるため、成績不振はペアの選択、執行ポリシー、またはその両方に起因し得ます。提案手法では、大規模言語モデルを両階層のポリシーとして用い、勾配ベースのファインチューニングではなくプロンプト更新によって適応させます。
軌跡とエピソードの各段階で得られるテキストのフィードバックを使い、ペア選択の抽象化と執行行動を別々に調整します。著者らは、この分離によってパフォーマンス変化の要因を特定しやすくなり、二つのポリシー階層間の不安定性が抑えられると論じています。実市場データを用いた実験では、従来型およびLLMベースのベースラインを一貫して上回る結果が報告されています。データセット、取引コスト、リスク管理、評価設計は明示されていないため、この説明だけでは実運用の可能性や結果の一般化範囲を判断できません。
主なアイデア
- 長期のペア選択と短期の取引執行を分けて扱います。
- 大規模言語モデルを、取引階層の両方のポリシーとして使います。
- プロンプト更新でテキストのフィードバックを反映し、勾配ベースのファインチューニングなしにポリシーを適応させます。
- 実市場データによる実験では従来型およびLLMベースのベースラインを上回りましたが、評価の詳細は示されていません。
タグ
全文
# Moira: Language-driven Hierarchical Reinforcement Learning for Pair Trading # Moira: Language-driven Hierarchical Reinforcement Learning for Pair Trading Many sequential decision-making problems exhibit hierarchical structure, where high-level semantic choices constrain downstream actions and feedback is delayed and ambiguous. Learning in such settings is challenging due to credit assignment: performance degradation may arise from flawed abstractions, suboptimal execution, or their interaction. We study this challenge through pair trading, a domain that naturally combines long-horizon semantic reasoning for asset pair selection with short-horizon execution under partial observability. We formulate pair trading as a hierarchical reinforcement learning problem and propose a language-driven optimization framework in which both high-level and low-level policies are parameterized by large language models (LLMs) and optimized exclusively through prompt updates. Our approach leverages pretrained LLMs as hierarchical policies and uses trajectory- and episode-level textual feedback to adapt abstractions and execution without gradient-based fine-tuning. By explicitly separating abstraction selection from execution, the framework reduces non-stationarity across hierarchical levels and enables targeted adaptation under delayed feedback. Experiments on real-world market data show consistent improvements over traditional and LLM-based baselines, demonstrating the effectiveness of language-driven hierarchical reinforcement learning.
出典を明記したうえで、ライセンスに従って全文を掲載しています。 ライセンス: abstract CC0
この要約は原文をもとにStratmillのリサーチエージェントが作成したもので、出典の複製ではありません。