본문으로 건너뛰기
라이브러리 문서 전체

페어 트레이딩을 위한 계층적 언어 모델 강화학습

기사 arXiv papers · 저자: Polydoros Giannouris et al.

요약

이 연구는 페어 트레이딩을 서로 연결된 두 단계의 계층적 의사결정 문제로 다룹니다. 하나는 더 긴 기간의 자산 쌍 선택이고, 다른 하나는 부분 관측 상태에서 더 짧은 기간에 거래를 실행하는 것입니다. 피드백이 늦게 도착하고 모호할 수 있어, 부진한 결과는 쌍 선택, 실행 정책, 또는 둘 다에서 비롯될 수 있습니다. 제안된 방법은 두 단계 모두에서 대규모 언어 모델을 정책으로 사용하고, 경사 기반 미세 조정 대신 프롬프트를 갱신해 모델을 적응시킵니다.

궤적 및 에피소드 수준의 텍스트 피드백을 사용해 쌍 선택 추상화와 실행 행동을 별도로 조정합니다. 저자들은 이러한 분리가 성과 변화의 원인을 구분하고 두 정책 수준 사이의 불안정성을 줄인다고 주장합니다. 실제 시장 데이터 실험에서 기존 방식 및 LLM 기반 기준선보다 일관되게 개선되었다고 보고합니다. 문서에는 데이터셋, 거래 비용, 리스크 통제, 평가 설계가 명시되지 않아 이 설명만으로는 실거래 가능성이나 결과의 일반화 범위를 판단할 수 없습니다.

핵심 아이디어

  • 이 접근법은 장기 자산 쌍 선택과 단기 거래 실행을 분리합니다.
  • 대규모 언어 모델은 트레이딩 계층 구조의 두 단계에서 정책으로 작동합니다.
  • 프롬프트 갱신은 경사 기반 미세 조정 없이 텍스트 피드백으로 정책을 적응시킵니다.
  • 실제 시장 데이터 실험은 기존 방식 및 LLM 기반 기준선보다 나은 결과를 보고하지만, 평가 세부사항은 제시되지 않습니다.

태그

전문
# Moira: Language-driven Hierarchical Reinforcement Learning for Pair Trading


# Moira: Language-driven Hierarchical Reinforcement Learning for Pair Trading









Many sequential decision-making problems exhibit hierarchical structure, where high-level semantic choices constrain downstream actions and feedback is delayed and ambiguous. Learning in such settings is challenging due to credit assignment: performance degradation may arise from flawed abstractions, suboptimal execution, or their interaction. We study this challenge through pair trading, a domain that naturally combines long-horizon semantic reasoning for asset pair selection with short-horizon execution under partial observability. We formulate pair trading as a hierarchical reinforcement learning problem and propose a language-driven optimization framework in which both high-level and low-level policies are parameterized by large language models (LLMs) and optimized exclusively through prompt updates. Our approach leverages pretrained LLMs as hierarchical policies and uses trajectory- and episode-level textual feedback to adapt abstractions and execution without gradient-based fine-tuning. By explicitly separating abstraction selection from execution, the framework reduces non-stationarity across hierarchical levels and enables targeted adaptation under delayed feedback. Experiments on real-world market data show consistent improvements over traditional and LLM-based baselines, demonstrating the effectiveness of language-driven hierarchical reinforcement learning.

출처의 라이선스에 따라 출처를 표시하고 전문을 공개합니다. 라이선스: abstract CC0

이 요약은 원문을 바탕으로 Stratmill의 리서치 에이전트가 작성했으며, 원문을 복사한 것이 아닙니다.