Implementing K-Means Clustering in MQL5 with Rectilinear Distance
Summary
The article introduces unsupervised clustering and implements k-means in MQL5. It walks through selecting initial centroids, measuring each observation's distance to each centroid, assigning observations to the nearest cluster, grouping observations, and updating centroids from cluster means. For distance, the example uses rectilinear distance, the sum of coordinate-wise absolute differences, and includes a small two-dimensional dataset with intermediate distance and assignment outputs.
It also describes storing unequal-sized clusters in a fixed-width matrix and applying mean normalization to price data before plotting. The author notes that centroid selection can be random when the number of clusters is known, but recommends avoiding random initialization when the routine is used with the elbow method to select a cluster count. The examples illustrate algorithm mechanics rather than demonstrating trading performance; the article provides no evaluation of predictive value or investment results.
Key ideas
- K-means is an unsupervised method that assigns observations to one of a chosen number of clusters.
- The example initializes centroids, assigns each point by minimum rectilinear distance, then updates centroids using cluster means.
- A matrix with one row per cluster can hold cluster members when cluster sizes differ.
- Mean normalization is applied to price data to make its plotted values more comparable in scale.
- Random centroid selection is discouraged in the article's elbow-method use case.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.