सामग्री पर जाएं
लाइब्रेरी के सभी दस्तावेज़

पेयर ट्रेडिंग के लिए पदानुक्रमित भाषा-मॉडल रीइन्फोर्समेंट लर्निंग

लेख arXiv papers · लेखक: Polydoros Giannouris et al.

सारांश

यह कार्य पेयर ट्रेडिंग को दो जुड़ी अवस्थाओं वाली पदानुक्रमित निर्णय समस्या के रूप में देखता है: लंबी अवधि में परिसंपत्तियों की जोड़ी चुनना और आंशिक अवलोकनीयता के तहत छोटी अवधियों में ट्रेड निष्पादित करना। प्रतिक्रिया देर से और अस्पष्ट आ सकती है, इसलिए खराब नतीजे जोड़ी के चयन, निष्पादन नीति या दोनों से उत्पन्न हो सकते हैं। प्रस्तावित विधि दोनों स्तरों पर बड़े भाषा मॉडलों को नीतियों के रूप में उपयोग करती है और ग्रेडिएंट-आधारित फ़ाइन-ट्यूनिंग के बजाय प्रॉम्प्ट अपडेट के माध्यम से उन्हें अनुकूलित करती है।

ट्रैजेक्टरी और एपिसोड स्तरों पर मिली पाठ-आधारित प्रतिक्रिया का उपयोग जोड़ी-चयन के अमूर्तन और निष्पादन व्यवहार को अलग-अलग समायोजित करने के लिए किया जाता है। लेखकों का तर्क है कि यह पृथक्करण प्रदर्शन में बदलाव के स्रोतों को अलग पहचानने में मदद करता है और दोनों नीति स्तरों के बीच अस्थिरता घटाता है। वास्तविक बाज़ार डेटा पर किए गए प्रयोगों में पारंपरिक और LLM-आधारित बेसलाइन की तुलना में लगातार सुधार दिखाए जाने की सूचना है। दस्तावेज़ डेटासेट, ट्रेडिंग लागत, जोखिम नियंत्रण या मूल्यांकन डिज़ाइन निर्दिष्ट नहीं करता, इसलिए केवल इस विवरण से लाइव ट्रेडिंग में उपयोगिता या परिणामों के व्यापक सामान्यीकरण का आकलन संभव नहीं है।

मुख्य विचार

  • यह तरीका लंबी अवधि की जोड़ी-चयन प्रक्रिया को छोटी अवधि के ट्रेड निष्पादन से अलग करता है।
  • ट्रेडिंग पदानुक्रम के दोनों स्तरों पर बड़े भाषा मॉडल नीतियों की भूमिका निभाते हैं।
  • प्रॉम्प्ट अपडेट पाठ-आधारित प्रतिक्रिया का उपयोग करके ग्रेडिएंट-आधारित फ़ाइन-ट्यूनिंग के बिना नीतियों को अनुकूलित करते हैं।
  • वास्तविक बाज़ार डेटा पर बताए गए प्रयोग पारंपरिक और LLM-आधारित बेसलाइन से बेहतर हैं, हालांकि मूल्यांकन का विवरण उपलब्ध नहीं है।

टैग

पूरा पाठ
# Moira: Language-driven Hierarchical Reinforcement Learning for Pair Trading


# Moira: Language-driven Hierarchical Reinforcement Learning for Pair Trading









Many sequential decision-making problems exhibit hierarchical structure, where high-level semantic choices constrain downstream actions and feedback is delayed and ambiguous. Learning in such settings is challenging due to credit assignment: performance degradation may arise from flawed abstractions, suboptimal execution, or their interaction. We study this challenge through pair trading, a domain that naturally combines long-horizon semantic reasoning for asset pair selection with short-horizon execution under partial observability. We formulate pair trading as a hierarchical reinforcement learning problem and propose a language-driven optimization framework in which both high-level and low-level policies are parameterized by large language models (LLMs) and optimized exclusively through prompt updates. Our approach leverages pretrained LLMs as hierarchical policies and uses trajectory- and episode-level textual feedback to adapt abstractions and execution without gradient-based fine-tuning. By explicitly separating abstraction selection from execution, the framework reduces non-stationarity across hierarchical levels and enables targeted adaptation under delayed feedback. Experiments on real-world market data show consistent improvements over traditional and LLM-based baselines, demonstrating the effectiveness of language-driven hierarchical reinforcement learning.

स्रोत के लाइसेंस के तहत श्रेय सहित पूरा पाठ दिखाया गया है। लाइसेंस: abstract CC0

यह सारांश मूल स्रोत के आधार पर Stratmill के शोध एजेंट ने लिखा है; यह स्रोत की प्रति नहीं है।