Risk-Aware Recurrent Reinforcement Learning for Pair Trading
Summary
This document describes CREDIT, a reinforcement learning approach to pair trading that uses a bidirectional gated recurrent unit and temporal attention. The architecture is intended to capture relationships across time in the price movements of two assets, including patterns that may be missed when market states are considered independently. The approach also uses a reward that accounts for both trading profit and risk, aiming to discourage trades with high potential gains and losses.
The authors frame the method as a response to frequent trading, transaction costs, and risk-taking in earlier reinforcement learning approaches. They report that CREDIT outperforms existing reinforcement learning methods and earns significant profit in experiments using five years of U.S. stock data. The excerpt gives no benchmark details, cost assumptions, risk measurements, or numerical results, so it is not possible to assess the size or robustness of the reported advantage. Performance in other markets or periods is not established here.
Key ideas
- CREDIT applies recurrent learning and temporal attention to price histories for two-asset pair trading.
- Its reward accounts for both profit and trading risk.
- The design aims to learn longer-term price patterns and reduce excessive trading.
- Experiments on five years of U.S. stock data reportedly outperform other reinforcement learning methods.
- The excerpt does not provide enough detail to evaluate robustness or real-world trading costs.
Tags
Full text
# Mastering Pair Trading with Risk-Aware Recurrent Reinforcement Learning # Mastering Pair Trading with Risk-Aware Recurrent Reinforcement Learning Although pair trading is the simplest hedging strategy for an investor to eliminate market risk, it is still a great challenge for reinforcement learning (RL) methods to perform pair trading as human expertise. It requires RL methods to make thousands of correct actions that nevertheless have no obvious relations to the overall trading profit, and to reason over infinite states of the time-varying market most of which have never appeared in history. However, existing RL methods ignore the temporal connections between asset price movements and the risk of the performed trading. These lead to frequent tradings with high transaction costs and potential losses, which barely reach the human expertise level of trading. Therefore, we introduce CREDIT, a risk-aware agent capable of learning to exploit long-term trading opportunities in pair trading similar to a human expert. CREDIT is the first to apply bidirectional GRU along with the temporal attention mechanism to fully consider the temporal correlations embedded in the states, which allows CREDIT to capture long-term patterns of the price movements of two assets to earn higher profit. We also design the risk-aware reward inspired by the economic theory, that models both the profit and risk of the tradings during the trading period. It helps our agent to master pair trading with a robust trading preference that avoids risky trading with possible high returns and losses. Experiments show that it outperforms existing reinforcement learning methods in pair trading and achieves a significant profit over five years of U.S. stock data.
Shown in full with attribution under the source's licence. Licence: abstract CC0
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.