Skip to content
All library documents

Clustering Similar Time Series with Features and Curve Structure

Article Quant Q&A · Author: fooledbypattern

Summary

This document surveys approaches to grouping time series that look similar, including curves with similar means but different variance or shape. One suggestion is to represent each series with informative features such as higher distribution moments, autocorrelation statistics, or spectral summaries, then inspect whether groups separate in that feature space before using k-means. A Gaussian mixture model is also mentioned as a possible alternative when the data structure warrants it.

Other responses describe treating each discretized curve as a vector, applying competitive learning vector quantization, or using functional-data methods. Under a Gaussian-process assumption, optimal quantization is linked to principal directions of the covariance operator. A neural-network approach based on summary statistics is reported, but its forecasting-model classification results were disappointing. The document offers methods and conceptual pointers rather than a comparative benchmark. The suitable representation and algorithm depend on the data-generating process, the desired notion of similarity, and whether shape, variance, or forecast behavior matters.

Key ideas

  • K-means can cluster curves by their overall vector shape, not only by a constant mean level.
  • Feature engineering with moments, autocorrelation, or spectral summaries can expose differences that raw means miss.
  • Gaussian mixtures, vector quantization, neural networks, and functional-data methods are possible alternatives.
  • The neural-network example had weak results, and no method is recommended universally.
  • For Gaussian-process curves, functional quantization connects representative curves to leading covariance directions.

Tags

Full text
# How to group timeseries showing similar curve


# How to group timeseries showing similar curve












I am trying to classify similar looking curves of a timeseries and was wondering what is the best algorithm to research. Reading R, it looks like k-means clustering could be applied - but I don't know if there are better algorithms. Any pointers is much appreciated. My concern is if k-means could group points around a neighborhood of specific mean, but what if two curves with similar mean show different variance.

## Answer by Ram Ahluwalia (score 12)

https://quant.stackexchange.com/a/2527

If the means are similar, then K-means will not do a great job. I would generate new features, perhaps based on higher moments of the distribution or some other properties (auto-correlation, summary of spectral density, etc.).

Using this new set of features, If you see separation of two curves when you plot draws in feature space then k-means would be an effective grouping algorithm.

It's hard to diagnose the best classification tool without better understanding the data generating process. For example, it might make sense to use a mixture of gaussians model instead.

This R task view page on clustering will provide a broad list of various tools you can use.

## Answer by Samik R (score 5)

https://quant.stackexchange.com/a/2535

We have used Neural Networks for this purpose. What we did was this: select a set of characteristics which can be directly calculated from the time series (e.g., mean, standard deviation, skewness, kurtosis, p-value from a normal fit, ACF stats etc.), and then run an NN to learn from a set of timeseries about how to classify them. We were trying to classify which among a set of models would produce the best forecasts for a series. We had the forecast answers, since we used the M-Competition data.

Unfortunately, the results weren't great. May be we were looking at wrong set of characteristics, may be our classification were wrong. But this is definitely an way to solve the problem you mention. I would be interested in knowing if you get any success in this work.

I would also be interested if someone wants to collaborate on this investigation. Please see my profile for contact information.

## Answer by Quant (score 1)

https://quant.stackexchange.com/a/9641

> My concern is if k-means (AKA lloyd algorithm) could group points around a neighborhood of specific mean, but what if two curves with similar mean show different variance.

K-means will groups curves which are close to an "average curve", but not necessarily close to a constant mean. Three very separated clusters can have the same mean. One being a cluster whose average is oscillating, the other one around an affine function, and the third one a constant function. In a sense you need to consider that each curve is a n-dimensional vector in the case where n is the number of points of discretization.

If you really don't want to use K-means, there is always the so-called "CLVQ algorithm" (competitive learning vector quantization) It should give similarresults.

More generally, quantization / clustering has extended to the functional case in "Functional quantization of Gaussian processes" by Luschgy and Pagès (Journal of functional analysis). The main result is that if you assume that your curves are independent draws of a Gaussian process, the optimal quantizer of that Gaussian process will span a principal space of its covariance operator, in other words, it will span the same finite-dimensional space as the first Karhunen-Loève eigenfunctions. This result as then been used by Tarpey and Kinateder in "Clustering Functional Data" (Journal of classification).

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.