Selecting Pairs with PCA and Density-Based Clustering
Summary
This module describes a candidate-selection process for pairs trading based on dimensionality reduction and clustering. It starts from a panel of asset prices, converts prices to returns, standardizes them, and applies principal component analysis to create lower-dimensional features for each asset. OPTICS or DBSCAN then groups assets with similar feature representations, and the module generates all pairwise combinations within each cluster. OPTICS is presented as a less parameter-sensitive option, while DBSCAN allows more direct control over clustering settings; a nearest-neighbor distance plot can help guide the latter's distance threshold.
The code also includes PCA and t-SNE visualizations and a sector-based way to generate candidate pairs. These outputs identify assets for further investigation, but the excerpt does not describe spread construction, cointegration or other validation tests, trading rules, costs, or performance results. Clustering similarity alone therefore does not establish that a pair will mean-revert or be profitable.
Key ideas
- The workflow converts price histories into standardized returns before feature extraction.
- PCA compresses the asset return panel into a lower-dimensional representation for clustering.
- OPTICS and DBSCAN group assets, and pairs are formed from combinations within each cluster.
- A nearest-neighbor distance plot can inform DBSCAN's distance setting.
- Cluster membership proposes candidates but does not demonstrate spread stability or profitability.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.