Reinforcement Learning for Statistical Arbitrage in Paired Stocks
Summary
This study presents a model-free reinforcement-learning framework for statistical arbitrage between paired stocks. It first constructs a mean-reverting spread by choosing asset coefficients that minimize an empirical measure of how long the spread takes to revert. This substitutes an observed reversion-time criterion for reliance on a fixed model assumption.
For trading, the framework uses reinforcement learning to select a mean-reversion strategy. Its state includes recent price movement trends, rather than relying only on the spread’s distance from a long-term average, and its reward is tailored to the characteristics of mean-reversion trading. The provided description outlines the design but gives no empirical performance results, training details, or comparison with other strategies. It therefore explains an approach, while leaving its practical effectiveness and robustness unestablished here.
Key ideas
- The framework builds paired-stock spreads by minimizing an empirical reversion-time measure.
- Asset coefficients are optimized as part of the spread construction.
- Reinforcement learning is used to choose the trading actions for the spread.
- The state representation includes recent price trends as well as mean-reversion context.
- The description does not report performance evidence or implementation details.
Tags
Full text
# Advanced Statistical Arbitrage with Reinforcement Learning # Advanced Statistical Arbitrage with Reinforcement Learning Statistical arbitrage is a prevalent trading strategy which takes advantage of mean reverse property of spread of paired stocks. Studies on this strategy often rely heavily on model assumption. In this study, we introduce an innovative model-free and reinforcement learning based framework for statistical arbitrage. For the construction of mean reversion spreads, we establish an empirical reversion time metric and optimize asset coefficients by minimizing this empirical mean reversion time. In the trading phase, we employ a reinforcement learning framework to identify the optimal mean reversion strategy. Diverging from traditional mean reversion strategies that primarily focus on price deviations from a long-term mean, our methodology creatively constructs the state space to encapsulate the recent trends in price movements. Additionally, the reward function is carefully tailored to reflect the unique characteristics of mean reversion trading.
Shown in full with attribution under the source's licence. Licence: abstract CC0
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.