Skip to content
All library documents

Unsupervised Learning for Fund Visualization, Stock Clustering, and Factor Premia

Article BigQuant

Summary

This article surveys three uses of unsupervised learning in investment research: manifold learning, clustering, and matrix factorization. For visualization, it applies t-SNE to fund returns, projecting high-dimensional observations into two dimensions so funds with similar returns appear near one another. The article reports that t-SNE performed best among the tested dimensionality-reduction methods on a handwritten-digit dataset, and that similar-return funds formed visible groups in the projection.

For stock analysis, it compares K-Means, hierarchical, and spectral clustering using industry-concept affiliations. K-Means and hierarchical clustering performed similarly and better than spectral clustering, while the hierarchical results grouped stocks with related concepts. For factor pricing, it describes a three-step PCA approach: extract return components, estimate their premia with cross-sectional regression, then use time-series regression to estimate the target factor premium. The source presents this as more accurate than traditional estimation, but provides no detailed quantitative results here. It also cautions that unsupervised methods reveal structure rather than directly forecasting future asset performance.

Key ideas

  • Unsupervised learning can reveal structure in unlabeled financial data, though it is not presented as a direct forecasting method.
  • t-SNE can visualize fund-return similarities by projecting return data into two dimensions.
  • The reported stock-clustering comparison favors K-Means and hierarchical clustering over spectral clustering.
  • A three-step PCA and regression procedure is described for estimating factor premia when controls are omitted.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.