Skip to content
All library documents

K-Means Clustering for Unsupervised Analysis of Stocks

Article SuperMind

Summary

This article introduces unsupervised clustering and contrasts it with supervised classification: cluster labels are not known in advance, so judging whether a grouping is correct is difficult. The goal is to group observations so that members of each cluster resemble one another while differing from members of other clusters. Such groupings can organize data or help explore possible hidden categories.

It explains K-means as an iterative procedure: select K initial centers, assign each observation to its nearest center, recompute centers from the resulting groups, and repeat until assignments stop changing. Because the initial centers can affect the result, the article suggests running several fits with different starts and combining assignments by voting. It gives an example clustering constituents of the SSE 50 into four groups using two features from a single day. That illustration is not a return test, and the article does not discuss feature scaling, choosing K, or validating whether the clusters are useful for investment decisions.

Key ideas

  • Clustering groups observations without known labels, so result quality can be difficult to evaluate.
  • K-means alternates between nearest-center assignment and recalculating cluster centers.
  • Initial center selection can materially affect the resulting groups.
  • Repeated runs with different initializations can be combined by voting.
  • The stock example uses two features for SSE 50 constituents on a single date and does not test trading performance.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.