Clustering Stocks with Return Data, Features, and Distance Metrics
Summary
The document discusses building an affinity matrix to cluster stocks, with the aim of studying cohesion within and between ETFs and possibly supporting arbitrage research or risk modeling. It compares clustering on summary statistics such as average return and volatility with using full return histories, and considers adding characteristics such as market capitalization, industry, dividends, valuation ratios, or factor betas.
The answers recommend choosing features that are stable and informative, scaling mixed features before calculating distances, and considering robust distance measures when returns contain outliers. One response suggests dimension reduction methods such as principal component analysis or manifold learning for feature extraction and visualization. Another describes clustering stocks by correlations of daily changes and using K-means, reporting industry-separated groups in one Brazilian-stock experiment. The suggested number of clusters depends on the sample and purpose; there is no universal rule. These examples are exploratory, and the reported clustering results do not establish predictive value or arbitrage profitability.
Key ideas
- Use return histories when temporal co-movement matters, since similar averages and volatilities can hide different return sequences.
- Scale features measured in different units so high-variance inputs do not dominate distance calculations.
- Choose features and distance functions based on the intended use and the information each feature contributes.
- Dimensionality reduction can help extract features and visualize high-dimensional stock clusters.
- Cluster count depends on the sample and objective, and should not be treated as a universal constant.
Tags
Full text
# How to cluster stocks and construct an affinity matrix?
# How to cluster stocks and construct an affinity matrix?
My goal is to find clusters of stocks. The "affinity" matrix will define the "closeness" of points. This article gives a bit more background. The ultimate purpose is to investigate the "cohesion" within ETFs and between similar ETFs for arbitrage possibilities. Eventually if everything goes well this could lead to the creation of a tool for risk modelling or valuation. Currently the project is in the proposal/POC phase so resources are limited.
I found this Python example for clustering with related docs. The code uses correlations of the difference in open and close prices as values for the affinity matrix. I prefer to use the average return and standard deviation of returns. This can be visualised as a two dimensional space with the average and standard deviation as dimensions. Instead of correlation, I would then calculate the "distance" between data points (stocks) and fill the affinity matrix with the distances. The choice of the distance function is still an open issue. Is calculating the distance between data points instead of correlations valid?
If it is can I extend this approach with more dimensions, such as dividend yield or ratios such as price/earnings?
I did a few experiments with different numbers of parameters and different distance functions resulting in different numbers of clusters ranging from 1 to more than 300 for a sample size of 900 stocks. The sample consists of large and mid cap stocks listed on the NYSE and NASDAQ. Is there a rule of thumb for the number of clusters one should expect?
## Answer by Ram Ahluwalia (score 12, accepted)
https://quant.stackexchange.com/a/2265
You should consider an unsupervised learning algorithm such as K-nearest neighbor ('KNN').
KNN will measure the distance amongst the observations in your space. You can and probably should consider alternative distance functions (besides euclidean) particularly if you are clustering on features such as returns which have outliers. There are quite a few unsupervised clustering algorithms out there - see here. You can certainly include features such as stock characteristics with these algorithms. You can also include the betas of the securities with respect to various risk factors as well. This would allow you to capture the distances in correlation space since a security based covariance matrix can be expressed as the : cross-product of (betas for factors) * covariance matrix of factor returns * transposed(betas for factors).
I would spend time thinking about the appropriate choice of features (which features are stable? which features predict risk or return? which sets of features are contributing unique sources of information? what are the invariants?) and choice of distance function.
Also, if you are mixing features with different unit scales (i.e. returns, betas, variances) then you need to normalize/pre-process your inputs otherwise the features with the highest variance will be the primary basis for clustering. Alternatively, you can stick to one class of features for your your clustering so you have some more intuition on interpreting the results.
## Answer by Flake (score 5)
https://quant.stackexchange.com/a/2283
Quant Guy's answer is quite informative for your question already.
Just to add few other things: instead of figuring out the choice of features by your own brain, you could also use machine learning techniques to help in extracting the 'features' for your specific purpose, e.g. risk modeling or returns forecasting or portfolio construction as mentioned by Tal.
Take a look at Principal component analysis and manifold learning (e.g. isomap). Even more interesting, the Unsupervised Feature Learning and Deep Learning. The first two methods both have implementation in the scikit-learn, the library you are currently looking into.
The first two methods mentioned above, could help you not only to extract the more important components from you features, but also to visualize your clustering given your feature dimension is bigger than 2.
## Answer by Tal Fishman (score 2)
https://quant.stackexchange.com/a/2297
Rather than suggesting alternative clustering techniques, as Quant Guy and Flake have (great advice, btw), I'll offer my thoughts on the method you've proposed.
On the characteristics used to cluster stocks: You propose using sample statistics (mean and standard deviation of returns). I would suggest you use the entire return (not price) series. For example, if stock A's returns on two successive days are (+1%,-1%) and B's returns are (-1%,+1%), your method would rank these two very closely based on mean and standard deviation, when in fact they should be quite far apart, particularly if most stock pairs in your sample are positively correlated (which I believe the vast majority are). Potential elaborations on this method include volatility-adjusted returns, market-beta adjusted returns, and excess returns relative to some risk model. I would shy away from using too many non-return characteristics, particularly fundamentals such as dividend yield or P/E, but you may want to introduce size (market cap) and industry. If your sample is exclusively ETFs, I strongly reiterate my advice to shy away from all fundamentals, including size and industry (which are irrelevant for broad ETFs and potentially misleading for narrow ones).
On the distance function: The function referenced in the paper you linked seems very complicated so I can't comment on it, but in general you should consider that if you use more than one type of input (e.g. returns and fundamentals), you should upweight differences which should be close together and downweight differences which are expected to be far apart, using something like Mahalanobis.
On the number of clusters: You will need to think more about your sample to figure out the optimal number of clusters. If you are looking at the entire market, I think statistical techniques are generally good at identifying no more than about 5 independent sources of variation. Assuming each one of these has 2 potential states, that implies $2^5=32$ clusters. If you repeat the analysis using only ETFs in your sample, given all the overlap, I'd go for even fewer, perhaps 10 clusters.
Best of luck. Your question does make sense to me now; the edits helped immensely.
## Answer by Felipe Martins Melo (score 0)
https://quant.stackexchange.com/a/37016
I was once trying to do something similar. My idea as to find baskets of stocks that behaved similarly.
Similar stocks, in my understanding, are stocks that vary together in a temporal way, i.e., if they go they same direction by a similar amount in the same days, than they can be considered similar.
What worked for me was this:
1) I picked one year of data from the stocks of interest.
2) Calculated the correlation between the daily variations only, generating a correlation matrix with a value of 1 all over the main diagonal.
3) Transformed the columns into feature vectors (could've been the lines as well)
4) Fed those vectors into a simple Kmeans (I used Apache Spark's implementation)
5) Starting from K = 2, I kept iterating until my within set sum of squared errors found a valley (in my case, they best K was 6).
6) There the clusters were, interestingly separated by industry, with banks in one cluster, siderurgics in another, etc. Basically, I clustered the stocks that correlate the most with the same set of other stocks.
I've only tried it for a set of Brazilian stocks, though.
## Answer by Alex C (score 0)
https://quant.stackexchange.com/a/37025
You may be interested in a paper by Marcos Lopez de Prado called Building Diversified Portfolios that Outperform Out-of-Sample, in Journal of Portfolio Management, 2016
He uses a clustering technique (described in the paper) to group stocks for the purpose of forming a diversified risk parity portfolioShown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.