Skip to content
All library documents

Combining Correlation and Fundamental Features in Stock Clustering

Article Quant Q&A · Author: FredNgu

Summary

The document considers clustering a large stock universe using both return correlations and company characteristics such as valuation and size. Correlation-based distances can group stocks that move in similar directions, while feature-based clustering groups firms by measured attributes. The question is how to combine those distinct inputs in one clustering approach.

One response suggests applying principal component analysis and clustering on the resulting multidimensional representation, which may reduce the need to set feature importance directly. Another proposes estimating cluster exposures in a factor-style regression before calculating feature coefficients, and notes that resulting groups may resemble sectors or industries. These are brief suggestions, not a tested comparison: no implementation details, dataset, validation criteria, or empirical performance are supplied. PCA also does not by itself determine appropriate scaling or weighting across heterogeneous inputs, so practical use would require careful preprocessing and evaluation.

Key ideas

  • Correlation distances can group stocks by how their returns move together.
  • Fundamental characteristics such as valuation and size provide a different basis for clustering.
  • PCA is suggested as a way to create a multidimensional representation for clustering.
  • A factor-style regression using cluster exposures is offered as another approach.
  • The suggestions include no empirical comparison or validation results.

Tags

Full text
# Mixture of a similarity-based clustering and a feature-based clustering


# Mixture of a similarity-based clustering and a feature-based clustering












Take for example the S&P500 universe with 500 stocks. Something interesting would be to create clusters based on stocks' correlations in order to have clusters that have the same "direction" in the market. That's easily done by a clustering algorithm, with a well-defined distance matrix from the correlation matrix.

But if we want to cluster the 500 stocks based on both stocks' correlations and features, how would we do that ? As far as I know the clustering algorithms only works with either similarity measures (often correlations) or features (P/E ratio, size...) but not both at the same time.

Do you have any idea ?

## Answer by Hao Zhang (score 1)

https://quant.stackexchange.com/a/54233

Often before running these types of clustering algorithms, it would be good to run a PCA on them and create those as your feature set. This allows you to run clustering on multi-dimensional data without worrying about feature importance too much.

## Answer by chrisaycock (score 0)

https://quant.stackexchange.com/a/54229

You're describing something similar to a Fama-French model with cluster exposure instead of market exposure. So regress-out the cluster-specific beta before computing your feature coefficients.

I think that, in practice, your clusters will simply be a sector or industry.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.