جوڑیوں کی ٹریڈنگ کے لیے خطرے سے آگاہ تکراری تقویتی سیکھنا
خلاصہ
یہ دستاویز CREDIT بیان کرتی ہے، جو جوڑیوں کی ٹریڈنگ کے لیے تقویتی سیکھنے کا طریقہ ہے اور دو طرفہ گیٹڈ ریکرنس یونٹ اور زمانی توجہ استعمال کرتا ہے۔ اس ساخت کا مقصد دو اثاثوں کی قیمتوں کی حرکات میں وقت کے ساتھ تعلقات کو پکڑنا ہے، جن میں ایسے نمونے بھی شامل ہیں جو مارکیٹ کی حالتوں کو الگ الگ دیکھنے پر نظر سے رہ سکتے ہیں۔ یہ طریقہ ایسا انعام بھی استعمال کرتا ہے جو ٹریڈنگ کے منافع اور خطرے، دونوں کو مدنظر رکھتا ہے اور اس کا مقصد ممکنہ طور پر بڑے منافع اور نقصان والی ٹریڈز کی حوصلہ شکنی ہے۔
مصنفین اس طریقے کو سابقہ تقویتی سیکھنے کے طریقوں میں کثرت سے ٹریڈنگ، لین دین کے اخراجات اور خطرہ لینے کے ردعمل کے طور پر پیش کرتے ہیں۔ وہ رپورٹ کرتے ہیں کہ CREDIT موجودہ تقویتی سیکھنے کے طریقوں سے بہتر ہے اور امریکی حصص کے پانچ سالہ ڈیٹا پر تجربات میں نمایاں منافع کماتا ہے۔ اقتباس میں بینچ مارک کی تفصیلات، لاگت کے مفروضے، خطرے کی پیمائشیں یا عددی نتائج نہیں، اس لیے رپورٹ شدہ برتری کے حجم یا مضبوطی کا اندازہ ممکن نہیں۔ یہاں دوسری منڈیوں یا ادوار میں کارکردگی قائم نہیں کی گئی۔
اہم خیالات
- CREDIT دو اثاثوں کی جوڑیوں کی ٹریڈنگ میں قیمتوں کی تاریخ پر تکراری سیکھنے اور زمانی توجہ لاگو کرتا ہے۔
- اس کا انعام منافع اور ٹریڈنگ کے خطرے، دونوں کو مدنظر رکھتا ہے۔
- ڈیزائن کا مقصد طویل مدتی قیمت کے نمونے سیکھنا اور ضرورت سے زیادہ ٹریڈنگ کم کرنا ہے۔
- امریکی حصص کے پانچ سالہ ڈیٹا پر تجربات میں دوسرے تقویتی سیکھنے کے طریقوں سے بہتر کارکردگی رپورٹ ہوئی ہے۔
- اقتباس مضبوطی یا حقیقی دنیا کے ٹریڈنگ اخراجات جانچنے کے لیے کافی تفصیل نہیں دیتا۔
ٹیگز
مکمل متن
# Mastering Pair Trading with Risk-Aware Recurrent Reinforcement Learning # Mastering Pair Trading with Risk-Aware Recurrent Reinforcement Learning Although pair trading is the simplest hedging strategy for an investor to eliminate market risk, it is still a great challenge for reinforcement learning (RL) methods to perform pair trading as human expertise. It requires RL methods to make thousands of correct actions that nevertheless have no obvious relations to the overall trading profit, and to reason over infinite states of the time-varying market most of which have never appeared in history. However, existing RL methods ignore the temporal connections between asset price movements and the risk of the performed trading. These lead to frequent tradings with high transaction costs and potential losses, which barely reach the human expertise level of trading. Therefore, we introduce CREDIT, a risk-aware agent capable of learning to exploit long-term trading opportunities in pair trading similar to a human expert. CREDIT is the first to apply bidirectional GRU along with the temporal attention mechanism to fully consider the temporal correlations embedded in the states, which allows CREDIT to capture long-term patterns of the price movements of two assets to earn higher profit. We also design the risk-aware reward inspired by the economic theory, that models both the profit and risk of the tradings during the trading period. It helps our agent to master pair trading with a robust trading preference that avoids risky trading with possible high returns and losses. Experiments show that it outperforms existing reinforcement learning methods in pair trading and achieves a significant profit over five years of U.S. stock data.
ماخذ کا حوالہ دیتے ہوئے مکمل متن دکھایا گیا ہے، ماخذ کے لائسنس کے تحت۔ لائسنس: abstract CC0
یہ خلاصہ اصل ماخذ سے Stratmill کے تحقیقی ایجنٹ نے لکھا ہے؛ یہ ماخذ کی نقل نہیں۔