Using K-Means Clustering for Index Market Timing
Summary
The strategy applies K-Means clustering to daily index features, such as recent average volatility, MACD, and average trading volume, to group trading days into market regimes. With two clusters, it compares each day’s feature vector with cluster centers, assigns it to the nearer group, and repeatedly updates the centers until they stabilize or reach an iteration limit. It then labels the clusters by examining subsequent returns in the training data: a cluster associated with stronger returns is treated as bullish, and the other as weak. The proposed timing rule enters at the next open when the forecast regime is bullish and exits at the next open when it is weak.
The example reports cumulative returns for five training observations, illustrating how the two clusters could be interpreted; this is not evidence of out-of-sample effectiveness. The document advises standardizing features because Euclidean distance is scale-sensitive, selecting features carefully, and balancing cluster count against overfitting. Regime labels must be evaluated against lagged returns to avoid using future information in a live decision.
Key ideas
- K-Means groups index days using engineered market features and distance from cluster centers.
- Cluster labels can be assigned by comparing their training-period subsequent returns.
- The proposed rule trades at the next open according to the predicted regime.
- Feature standardization matters because K-Means uses Euclidean distance.
- Too many clusters can overfit, while too few may fail to separate useful regimes.
- The example is small and does not demonstrate out-of-sample performance.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.