Machine Learning Clustering and Filters for Pairs Selection
Summary
The document outlines a pairs selection framework that first reduces security-return features with principal component analysis, then clusters the compact representations using OPTICS or DBSCAN. OPTICS can identify clusters without a fixed cluster count; DBSCAN is an alternative when domain knowledge can guide its sensitive distance parameter. The aim is to limit an otherwise unwieldy search while finding related securities that may not share a sector. The component count is chosen empirically, with a stated upper bound of 15 to limit sparse high-dimensional distances.
Candidate pairs are screened for cointegration, mean-reverting behavior, practical reversion speed, and enough crossings of the spread’s mean. The text discusses the Engle-Granger test’s dependence on which asset is treated as dependent, and presents orthogonal regression as a way to estimate a more order-symmetric hedge relationship. It also describes a Hurst exponent below 0.5 and bounds on half-life as filters, with monthly mean crossings as an opportunity criterion. These are proposed selection rules, not evidence of profitable trades; implementation choices, data periods, and out-of-sample validation remain important.
Key ideas
- PCA compresses security-return features before unsupervised clustering, with dimensionality balanced against useful variance and clusterability.
- OPTICS avoids specifying a cluster count, while DBSCAN offers an alternative whose distance threshold needs careful selection.
- Pairs are screened for cointegration, a Hurst exponent below 0.5, suitable mean-reversion half-life, and frequent spread crossings.
- Orthogonal regression is presented as a hedge-ratio approach that accounts for variation in both pair constituents.
- The framework describes selection criteria but does not establish that selected pairs will be profitable.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.