Skip to content
All library documents

Using K-Means to Group Mutual Funds by Volatility

Article Quant Q&A · Author: Tomás Ayala

Summary

The document considers how to classify mutual funds into low, medium, and high volatility groups without choosing fixed volatility thresholds in advance. The answer proposes unsupervised clustering with k-means: provide the funds’ volatility measurements and choose three clusters, then use the resulting cluster assignment to place each fund into a group. The cluster centers, calculated as the means of the groups, can help describe each group’s typical volatility.

A small example demonstrates the approach on one-dimensional numeric data, where the algorithm separates observations into three clusters and returns assignments and centers. The example illustrates the mechanics, not fund classifications or investment results. K-means requires the number of groups to be chosen, and the answer does not discuss how to estimate volatility, select a measurement period, handle outliers, or assess cluster quality. Therefore, the groups are data-dependent and should not be treated as universal volatility cutoffs.

Key ideas

  • K-means can group funds using their volatility measurements without preset threshold values.
  • The analyst specifies the desired number of clusters, such as three groups.
  • Cluster assignments map each observation to a group, while cluster centers summarize group averages.
  • The resulting boundaries depend on the dataset and do not define universal volatility categories.

Tags

Full text
# How to group mutual funds by volatility?


# How to group mutual funds by volatility?












I want to group Mutual Funds by their volatility.

Ideally, I would like to end up with the mutual funds beings attached to different groups:

- High volatility

- Medium volatility

- Low Volatility

My questions is : What numbers could be consider like low , medium , high volatility ?

Maybe some intervals: 0 -5 % is low 5 - 15% is Medium and Higher is High......

I'm little bit confused on how to tackle this problem...

## Answer by SRKX (score 3, accepted)

https://quant.stackexchange.com/a/4508

What you are looking for is an unsupervised learning algorithm algorithm: i.e an algorithm that will by itself determine the 3 most rational groups from your dataset. This method will allow you to choose the boundaries of the groups based on the dataset you provide and not by choosing some given fixed values.

The algorithm I suggest you to use is the K-means algorithm. You provide it with the data, and the number $k$ of clusters (groups) that you want to have. The algorithm will then split the data into the $k$ groups you would like. Note that this algorithm can handle points with several features, whereas you will be using only one (volatiliy).

Here is an idea of how it works in Matlab:

```
test=[0 1 2 3 100 105 98 1000 1001 997]';
[idx,C] = kmeans(test,3);
```

The value returned for `idx` is a vector where each point in `test` is attributed a cluster number (representing its group):

```
idx =

 2
 2
 2
 2
 3
 3
 3
 1
 1
 1
```

You can then look at the variable `C` which contains the mean of each cluster which could be undrestood as "the perfect point for each cluster"

```
c =

  999.3333
  1.5000
  101.0000
```

So it found three groups in `test` one around 999.33, one around 1.5, and one around 101.

That should do the trick.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.