التعلم المعزز المراعي للمخاطر لتداول الأزواج
الملخص
يصف هذا المستند CREDIT، وهو نهج للتعلم المعزز في تداول الأزواج يستخدم وحدة تكرارية مسوّرة ثنائية الاتجاه وانتباهًا زمنيًا. وتهدف البنية إلى التقاط العلاقات عبر الزمن في تحركات أسعار أصلين، بما فيها الأنماط التي قد تفوت عند النظر إلى حالات السوق كلٌّ على حدة. كما يستخدم النهج مكافأة تراعي ربح التداول ومخاطره معًا، بهدف تثبيط الصفقات ذات المكاسب والخسائر المحتملة الكبيرة.
يقدم المؤلفون الطريقة بوصفها استجابة للتداول المتكرر وتكاليف المعاملات والمجازفة في أساليب التعلم المعزز السابقة. ويذكرون أن CREDIT يتفوق على أساليب التعلم المعزز الحالية ويحقق أرباحًا ذات دلالة إحصائية في تجارب استخدمت بيانات أسهم أمريكية لخمس سنوات. ولا يقدم المقتطف تفاصيل المعايير المرجعية أو افتراضات التكاليف أو قياسات المخاطر أو النتائج الرقمية، لذا يتعذر تقييم حجم التفوق المعلن أو متانته. ولم يثبت هنا الأداء في أسواق أو فترات أخرى.
الأفكار الرئيسية
- يطبق CREDIT التعلم المتكرر والانتباه الزمني على سجلات أسعار أصلين لتداول الأزواج.
- تراعي المكافأة الربح ومخاطر التداول معًا.
- يهدف التصميم إلى تعلم أنماط الأسعار طويلة الأجل والحد من التداول المفرط.
- تشير التقارير إلى أن تجارب خمس سنوات من بيانات الأسهم الأمريكية تفوقت على أساليب تعلم معزز أخرى.
- لا يقدم المقتطف تفاصيل كافية لتقييم المتانة أو تكاليف التداول في الواقع.
الوسوم
النص الكامل
# Mastering Pair Trading with Risk-Aware Recurrent Reinforcement Learning # Mastering Pair Trading with Risk-Aware Recurrent Reinforcement Learning Although pair trading is the simplest hedging strategy for an investor to eliminate market risk, it is still a great challenge for reinforcement learning (RL) methods to perform pair trading as human expertise. It requires RL methods to make thousands of correct actions that nevertheless have no obvious relations to the overall trading profit, and to reason over infinite states of the time-varying market most of which have never appeared in history. However, existing RL methods ignore the temporal connections between asset price movements and the risk of the performed trading. These lead to frequent tradings with high transaction costs and potential losses, which barely reach the human expertise level of trading. Therefore, we introduce CREDIT, a risk-aware agent capable of learning to exploit long-term trading opportunities in pair trading similar to a human expert. CREDIT is the first to apply bidirectional GRU along with the temporal attention mechanism to fully consider the temporal correlations embedded in the states, which allows CREDIT to capture long-term patterns of the price movements of two assets to earn higher profit. We also design the risk-aware reward inspired by the economic theory, that models both the profit and risk of the tradings during the trading period. It helps our agent to master pair trading with a robust trading preference that avoids risky trading with possible high returns and losses. Experiments show that it outperforms existing reinforcement learning methods in pair trading and achieves a significant profit over five years of U.S. stock data.
يُعرض النص كاملًا مع نسبه إلى مصدره وفقًا لترخيصه. الترخيص: abstract CC0
أعدّ وكيل الأبحاث في Stratmill هذا الملخص استنادًا إلى المصدر الأصلي؛ وهو ليس نسخة منه.