Joint Pair Selection and Trading with Hierarchical Reinforcement Learning
Summary
This paper treats pair trading as a combined decision problem rather than a pipeline that first selects an asset pair and then trades it. The authors argue that separating these stages can discard useful information: selection may ignore whether a pair can be traded effectively, while a trading agent may overfit to chosen assets without learning from the wider asset universe.
Their hierarchical reinforcement learning method assigns pair choice to a high-level policy and trading actions to a low-level policy. The two policies are optimized jointly, allowing selection and execution decisions to inform one another. Experiments on real-world stock data are reported to show improved pair-trading effectiveness compared with existing selection and trading methods. The document provides no specific performance figures or detailed experimental limitations, so it does not establish how the approach generalizes to other markets, asset classes or trading costs.
Key ideas
- Pair selection and trade execution can be learned as connected parts of one task.
- A high-level policy chooses the asset pair, and a low-level policy makes trading decisions.
- Joint learning can pass trading-performance information back into pair selection.
- The study reports comparisons on real-world stock data against existing approaches.
- The provided description gives no detailed performance figures or evidence beyond the stock experiments.
Tags
Full text
# Select and Trade: Towards Unified Pair Trading with Hierarchical Reinforcement Learning # Select and Trade: Towards Unified Pair Trading with Hierarchical Reinforcement Learning Pair trading is one of the most effective statistical arbitrage strategies which seeks a neutral profit by hedging a pair of selected assets. Existing methods generally decompose the task into two separate steps: pair selection and trading. However, the decoupling of two closely related subtasks can block information propagation and lead to limited overall performance. For pair selection, ignoring the trading performance results in the wrong assets being selected with irrelevant price movements, while the agent trained for trading can overfit to the selected assets without any historical information of other assets. To address it, in this paper, we propose a paradigm for automatic pair trading as a unified task rather than a two-step pipeline. We design a hierarchical reinforcement learning framework to jointly learn and optimize two subtasks. A high-level policy would select two assets from all possible combinations and a low-level policy would then perform a series of trading actions. Experimental results on real-world stock data demonstrate the effectiveness of our method on pair trading compared with both existing pair selection and trading methods.
Shown in full with attribution under the source's licence. Licence: abstract CC0
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.