یادگیری تقویتی بازگشتی آگاه از ریسک برای معامله جفتی
خلاصه
این سند CREDIT را توصیف میکند؛ رویکردی مبتنی بر یادگیری تقویتی برای معامله جفتی که از واحد بازگشتی دروازهدار دوسویه و توجه زمانی استفاده میکند. این معماری برای ثبت روابط زمانی در حرکت قیمت دو دارایی طراحی شده است، از جمله الگوهایی که ممکن است با بررسی مستقل وضعیتهای بازار نادیده بمانند. این رویکرد همچنین از پاداشی استفاده میکند که هم سود معامله و هم ریسک را در نظر میگیرد و هدف آن جلوگیری از انجام معاملههایی با سود و زیان بالقوه زیاد است.
نویسندگان این روش را پاسخی به معاملات پرتکرار، هزینههای معامله و ریسکپذیری در رویکردهای پیشین یادگیری تقویتی میدانند. گزارش میکنند که CREDIT در آزمایشهایی با دادههای پنجساله سهام آمریکا از روشهای موجود یادگیری تقویتی بهتر عمل میکند و سود چشمگیری به دست میآورد. گزیده جزئیات معیارهای مبنا، فرضهای هزینه، سنجههای ریسک یا نتایج عددی را ارائه نمیکند؛ بنابراین ارزیابی اندازه یا استحکام برتری گزارششده ممکن نیست. عملکرد در بازارها یا دورههای دیگر در اینجا اثبات نشده است.
ایدههای کلیدی
- CREDIT یادگیری بازگشتی و توجه زمانی را بر تاریخچه قیمت دو دارایی در معامله جفتی اعمال میکند.
- تابع پاداش آن هم سود و هم ریسک معامله را در نظر میگیرد.
- هدف طراحی، یادگیری الگوهای قیمتی بلندمدتتر و کاهش معاملات بیشازحد است.
- طبق گزارش، روشها در آزمایشهای انجامشده روی دادههای سهام آمریکا بهمدت پنج سال، از روشهای دیگر یادگیری تقویتی بهتر عمل میکنند.
- گزیده جزئیات کافی برای ارزیابی استحکام یا هزینههای معامله در دنیای واقعی ارائه نمیکند.
برچسبها
متن کامل
# Mastering Pair Trading with Risk-Aware Recurrent Reinforcement Learning # Mastering Pair Trading with Risk-Aware Recurrent Reinforcement Learning Although pair trading is the simplest hedging strategy for an investor to eliminate market risk, it is still a great challenge for reinforcement learning (RL) methods to perform pair trading as human expertise. It requires RL methods to make thousands of correct actions that nevertheless have no obvious relations to the overall trading profit, and to reason over infinite states of the time-varying market most of which have never appeared in history. However, existing RL methods ignore the temporal connections between asset price movements and the risk of the performed trading. These lead to frequent tradings with high transaction costs and potential losses, which barely reach the human expertise level of trading. Therefore, we introduce CREDIT, a risk-aware agent capable of learning to exploit long-term trading opportunities in pair trading similar to a human expert. CREDIT is the first to apply bidirectional GRU along with the temporal attention mechanism to fully consider the temporal correlations embedded in the states, which allows CREDIT to capture long-term patterns of the price movements of two assets to earn higher profit. We also design the risk-aware reward inspired by the economic theory, that models both the profit and risk of the tradings during the trading period. It helps our agent to master pair trading with a robust trading preference that avoids risky trading with possible high returns and losses. Experiments show that it outperforms existing reinforcement learning methods in pair trading and achieves a significant profit over five years of U.S. stock data.
با ذکر منبع و مطابق مجوز اثر، بهطور کامل نمایش داده میشود. مجوز: abstract CC0
این خلاصه را عامل پژوهشی Stratmill بر پایه متن اصلی نوشته است؛ نسخهای از اثر منبع نیست.