Using K-Means to Find Candidate Pairs for Statistical Arbitrage
Summary
The article explains K-means clustering as an unsupervised method for grouping stocks with similar feature values, then applies it to candidate selection for statistical arbitrage. It contrasts clustering with manually choosing pairs by sector or industry, and describes testing candidate spreads for cointegration with an Augmented Dickey-Fuller test. The proposed stock features include dividend yield, valuation, market capitalization, earnings, and operating measures. A backtest is presented as a follow-on to selecting a cointegrated pair from a cluster.
The examples illustrate that apparent business similarity does not guarantee a stable spread, and that clustering may surface less obvious candidates. The reported pair results apply to historical data and do not establish that the relationship will persist. The article itself notes that a strategy can stop performing as cointegration weakens and that statistical arbitrage is not risk-free. Its demonstrations and parameter choices are illustrative; they do not provide evidence of robust performance across markets or time periods.
Key ideas
- K-means groups stocks by selected features to generate candidate pairs for further analysis.
- Sector similarity alone does not establish that two stocks have a cointegrated relationship.
- The article uses an Augmented Dickey-Fuller test on candidate spreads before backtesting.
- A historically cointegrated pair may lose that relationship, so statistical arbitrage carries risk.
- The examples do not establish that the selection method generalizes across periods or markets.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.