Predicting Stock Correlations to Build Dynamic Equity Clusters
Summary
The article describes a data-driven way to group stocks for portfolio diversification. Instead of relying only on fixed business-sector labels, it trains models to predict future pairwise return correlations, then clusters stocks using the predicted correlation matrix. Inputs cover historical daily and monthly returns, fundamental factors, text features from company filings, news co-occurrence embeddings, and existing GICS classifications. For each stock pair, feature distances and cosine similarities become model inputs; ridge regression, neural networks, and XGBoost are among the candidate models.
The reported comparisons use out-of-sample returns and benchmark clusters with GICS at several levels. The article says machine-learning clusters generally improve some measures of within-cluster return similarity or between-cluster separation, while GICS can retain lower within-group dispersion for some fundamental factors. Daily returns and GICS are identified as influential features. These findings are specific to the described universe and evaluation, and the excerpt omits full tables, sample details, and implementation choices, so it does not establish that the approach will improve diversification in other settings.
Key ideas
- The method predicts future pairwise stock return correlations from multiple data sources.
- Pairwise feature distances and similarities are used to train correlation models.
- Hierarchical clustering of predicted correlations creates dynamic groups at GICS-like levels.
- The article reports mixed comparisons across similarity, stability, and diversification measures.
- Historical returns and existing industry classification are described as important predictors.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.