用于配对交易的分层语言模型强化学习
文章 arXiv papers · 作者: Polydoros Giannouris et al.
总结
本文将配对交易构建为一个分层决策问题,包含两个相互关联的阶段:较长周期内选择资产对,以及在部分可观测条件下于较短周期内执行交易。由于反馈可能延迟且含糊,结果不佳可能源于资产对选择、执行策略,或两者兼有。所提方法在两个层级都使用大型语言模型作为策略,并通过更新提示进行调整,而非基于梯度的微调。
研究使用轨迹和回合层面的文本反馈,分别调整配对选择的抽象表示和交易执行行为。作者认为,这种区分有助于找出表现变化的来源,并减少两个策略层级之间的不稳定性。据报告,在真实世界市场数据上的实验持续优于传统基准和基于 LLM 的基准。文档未说明数据集、交易成本、风险控制或评估设计,因此仅凭此描述无法判断其是否适用于实盘交易,也无法判断结果的普遍性。
核心观点
- 该方法将长期资产对选择与较短周期的交易执行分开处理。
- 大型语言模型在交易层级的两个阶段都作为策略。
- 提示更新利用文本反馈调整策略,无需基于梯度的微调。
- 据报告,真实世界市场数据实验优于传统基准和基于 LLM 的基准,但未提供评估细节。
标签
全文
# Moira: Language-driven Hierarchical Reinforcement Learning for Pair Trading # Moira: Language-driven Hierarchical Reinforcement Learning for Pair Trading Many sequential decision-making problems exhibit hierarchical structure, where high-level semantic choices constrain downstream actions and feedback is delayed and ambiguous. Learning in such settings is challenging due to credit assignment: performance degradation may arise from flawed abstractions, suboptimal execution, or their interaction. We study this challenge through pair trading, a domain that naturally combines long-horizon semantic reasoning for asset pair selection with short-horizon execution under partial observability. We formulate pair trading as a hierarchical reinforcement learning problem and propose a language-driven optimization framework in which both high-level and low-level policies are parameterized by large language models (LLMs) and optimized exclusively through prompt updates. Our approach leverages pretrained LLMs as hierarchical policies and uses trajectory- and episode-level textual feedback to adapt abstractions and execution without gradient-based fine-tuning. By explicitly separating abstraction selection from execution, the framework reduces non-stationarity across hierarchical levels and enables targeted adaptation under delayed feedback. Experiments on real-world market data show consistent improvements over traditional and LLM-based baselines, demonstrating the effectiveness of language-driven hierarchical reinforcement learning.
在遵守原作品许可的前提下,附作者信息全文展示。 许可协议: abstract CC0
此摘要由 Stratmill 研究智能体根据原文撰写,并非原文副本。