رفتن به محتوا
همه اسناد کتابخانه

یادگیری تقویتی بازگشتی آگاه از ریسک برای معامله جفتی

مقاله arXiv papers · نویسنده: Weiguang Han et al.

خلاصه

این سند CREDIT را توصیف می‌کند؛ رویکردی مبتنی بر یادگیری تقویتی برای معامله جفتی که از واحد بازگشتی دروازه‌دار دوسویه و توجه زمانی استفاده می‌کند. این معماری برای ثبت روابط زمانی در حرکت قیمت دو دارایی طراحی شده است، از جمله الگوهایی که ممکن است با بررسی مستقل وضعیت‌های بازار نادیده بمانند. این رویکرد همچنین از پاداشی استفاده می‌کند که هم سود معامله و هم ریسک را در نظر می‌گیرد و هدف آن جلوگیری از انجام معامله‌هایی با سود و زیان بالقوه زیاد است.

نویسندگان این روش را پاسخی به معاملات پرتکرار، هزینه‌های معامله و ریسک‌پذیری در رویکردهای پیشین یادگیری تقویتی می‌دانند. گزارش می‌کنند که CREDIT در آزمایش‌هایی با داده‌های پنج‌ساله سهام آمریکا از روش‌های موجود یادگیری تقویتی بهتر عمل می‌کند و سود چشمگیری به دست می‌آورد. گزیده جزئیات معیارهای مبنا، فرض‌های هزینه، سنجه‌های ریسک یا نتایج عددی را ارائه نمی‌کند؛ بنابراین ارزیابی اندازه یا استحکام برتری گزارش‌شده ممکن نیست. عملکرد در بازارها یا دوره‌های دیگر در اینجا اثبات نشده است.

ایده‌های کلیدی

  • CREDIT یادگیری بازگشتی و توجه زمانی را بر تاریخچه قیمت دو دارایی در معامله جفتی اعمال می‌کند.
  • تابع پاداش آن هم سود و هم ریسک معامله را در نظر می‌گیرد.
  • هدف طراحی، یادگیری الگوهای قیمتی بلندمدت‌تر و کاهش معاملات بیش‌ازحد است.
  • طبق گزارش، روش‌ها در آزمایش‌های انجام‌شده روی داده‌های سهام آمریکا به‌مدت پنج سال، از روش‌های دیگر یادگیری تقویتی بهتر عمل می‌کنند.
  • گزیده جزئیات کافی برای ارزیابی استحکام یا هزینه‌های معامله در دنیای واقعی ارائه نمی‌کند.

برچسب‌ها

متن کامل
# Mastering Pair Trading with Risk-Aware Recurrent Reinforcement Learning


# Mastering Pair Trading with Risk-Aware Recurrent Reinforcement Learning









Although pair trading is the simplest hedging strategy for an investor to eliminate market risk, it is still a great challenge for reinforcement learning (RL) methods to perform pair trading as human expertise. It requires RL methods to make thousands of correct actions that nevertheless have no obvious relations to the overall trading profit, and to reason over infinite states of the time-varying market most of which have never appeared in history. However, existing RL methods ignore the temporal connections between asset price movements and the risk of the performed trading. These lead to frequent tradings with high transaction costs and potential losses, which barely reach the human expertise level of trading. Therefore, we introduce CREDIT, a risk-aware agent capable of learning to exploit long-term trading opportunities in pair trading similar to a human expert. CREDIT is the first to apply bidirectional GRU along with the temporal attention mechanism to fully consider the temporal correlations embedded in the states, which allows CREDIT to capture long-term patterns of the price movements of two assets to earn higher profit. We also design the risk-aware reward inspired by the economic theory, that models both the profit and risk of the tradings during the trading period. It helps our agent to master pair trading with a robust trading preference that avoids risky trading with possible high returns and losses. Experiments show that it outperforms existing reinforcement learning methods in pair trading and achieves a significant profit over five years of U.S. stock data.

با ذکر منبع و مطابق مجوز اثر، به‌طور کامل نمایش داده می‌شود. مجوز: abstract CC0

این خلاصه را عامل پژوهشی Stratmill بر پایه متن اصلی نوشته است؛ نسخه‌ای از اثر منبع نیست.