Using K-Means to Cluster Daily OHLC Bars and Explore Market Regimes
Summary
The article introduces K-means as a hard clustering method and applies it to daily open, high, low, and close data. It explains that the algorithm assigns each observation to the nearest cluster centroid and iterates assignments and centroid updates to reduce within-cluster variation. Since the optimization can settle on a local solution, repeated initializations are used to seek a lower-variation result. A synthetic example compares cluster assignments when the chosen cluster count differs from the number of generated groups.
For market data, the tutorial describes normalizing bar features relative to the open, clustering S&P 500 candles, visualizing the resulting groups, and tabulating which cluster tends to follow another. Such groupings may help researchers investigate recurring bar shapes or regimes, but the article does not establish a profitable signal. K-means requires a predefined cluster count, assigns every point including outliers, and can be sensitive to initialization and sample choice. Noisy financial data may yield apparent clusters that are artifacts rather than stable distributions.
Key ideas
- K-means creates hard assignments by repeatedly associating observations with their nearest centroid.
- The number of clusters must be selected in advance and can materially change the resulting partition.
- The tutorial uses normalized OHLC features to group daily bars and inspect cluster sequences.
- Random initialization can affect the local solution, so repeated fits can help find a lower within-cluster variation.
- Noise, outliers, and sensitivity to the sample make apparent financial clusters uncertain and potentially unstable.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.