Skip to content
All library documents

Hierarchical Language-Model Reinforcement Learning for Pair Trading

Article arXiv papers · Author: Polydoros Giannouris et al.

Summary

This work frames pair trading as a hierarchical decision problem with two linked stages: selecting an asset pair over a longer horizon and executing trades over shorter horizons under partial observability. Because feedback arrives late and may be ambiguous, poor results can stem from the choice of pair, the execution policy, or both. The proposed method uses large language models as policies at both levels and adapts them through prompt updates rather than gradient-based fine-tuning.

Textual feedback at the trajectory and episode levels is used to adjust pair-selection abstractions and execution behavior separately. The authors argue that this separation helps isolate sources of performance changes and reduces instability between the two policy levels. Experiments on real-world market data are reported to show consistent improvements over traditional and LLM-based baselines. The document does not specify the datasets, trading costs, risk controls, or evaluation design, so it is not possible from this description alone to judge live tradability or how broadly the results generalize.

Key ideas

  • The approach separates long-horizon pair selection from shorter-horizon trade execution.
  • Large language models serve as policies at both levels of the trading hierarchy.
  • Prompt updates use textual feedback to adapt policies without gradient-based fine-tuning.
  • Reported experiments on real-world market data outperform traditional and LLM-based baselines, though evaluation details are not provided.

Tags

Full text
# Moira: Language-driven Hierarchical Reinforcement Learning for Pair Trading


# Moira: Language-driven Hierarchical Reinforcement Learning for Pair Trading









Many sequential decision-making problems exhibit hierarchical structure, where high-level semantic choices constrain downstream actions and feedback is delayed and ambiguous. Learning in such settings is challenging due to credit assignment: performance degradation may arise from flawed abstractions, suboptimal execution, or their interaction. We study this challenge through pair trading, a domain that naturally combines long-horizon semantic reasoning for asset pair selection with short-horizon execution under partial observability. We formulate pair trading as a hierarchical reinforcement learning problem and propose a language-driven optimization framework in which both high-level and low-level policies are parameterized by large language models (LLMs) and optimized exclusively through prompt updates. Our approach leverages pretrained LLMs as hierarchical policies and uses trajectory- and episode-level textual feedback to adapt abstractions and execution without gradient-based fine-tuning. By explicitly separating abstraction selection from execution, the framework reduces non-stationarity across hierarchical levels and enables targeted adaptation under delayed feedback. Experiments on real-world market data show consistent improvements over traditional and LLM-based baselines, demonstrating the effectiveness of language-driven hierarchical reinforcement learning.

Shown in full with attribution under the source's licence. Licence: abstract CC0

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.