Machine Learning for Selecting Mean-Reverting Pairs
Summary
This article presents an unsupervised learning framework for narrowing the search for equity pairs that may exhibit mean reversion. It first applies principal component analysis to asset returns to represent shared risk exposures, then uses density-based clustering, such as DBSCAN or OPTICS, to group assets whose behavior is similar. This data-driven grouping is intended to reduce the number of candidate pairs and statistical tests without limiting the search to familiar categories such as securities in the same sector.
Pairs within clusters are screened using Engle–Granger cointegration, a Hurst exponent below 0.5, a mean-reversion half-life between one day and 365 days, and at least twelve crossings of the spread’s mean per year. The article contrasts this approach with distance and correlation screening, noting that correlation alone does not establish an equilibrium relationship and that testing every possible pair raises computational and false-discovery concerns. It describes a proposed framework and its implementation, but supplies no performance results here. The filters depend on historical estimates, and the article does not establish that selected pairs will remain mean-reverting or profitable.
Key ideas
- PCA on returns can summarize shared risk exposures before candidate pairs are formed.
- DBSCAN and OPTICS can cluster assets without relying on predefined sector membership.
- Clustering reduces the number of pairwise tests, helping manage computational cost and false discoveries.
- The proposed screening rules combine cointegration, Hurst exponent, half-life, and spread mean crossings.
- Correlation or low distance alone does not establish that a pair has a stable equilibrium relationship.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.