Skip to content
All library documents

Why Unsupervised Clustering Still Needs Out-of-Sample Evaluation

Article Quant Q&A · Author: Dionysios Georgiadis

Summary

The discussion considers a trading research process that clusters historical intraday observations, such as daily vectors of hourly exchange rates, and then measures the next day's average return for each cluster. A researcher might use the return associated with the current observation's cluster as a forecast. Although the clustering algorithm does not use future returns when forming groups, the full analysis still links clusters to subsequent outcomes, so it is not automatically free of in-sample bias.

The answer distinguishes fitting clusters from evaluating how useful they are on new data. If the clusters are built using the entire sample, some patterns may only have appeared later in the period, and their associated forward returns would not have been available at the time a historical forecast was made. A held-out or later chronological sample is therefore needed to assess whether new observations fit the learned clusters and whether the return relationship generalizes. The exchange does not give a detailed split design, account for repeated model choices, or report empirical performance; it explains why unsupervised feature construction does not remove the need for out-of-sample evaluation.

Key ideas

  • Clustering can be unsupervised while the later return analysis still creates a predictive claim.
  • Clusters fitted on the full sample may encode patterns that emerged only later.
  • Use new or held-out observations to assess cluster assignment and associated forward returns.
  • Out-of-sample validation tests whether a discovered relationship generalizes beyond its fitting data.

Tags

Full text
# Unsupervised learning and in out of sample


# Unsupervised learning and in out of sample












Assume we are given $N$ samples, let's say small timeseries of 1 hour resolution daily exchange rates - for the sake of argument. Each sample is a $24$ element vector $x$.

Then we proceed to do clustering using our favourite unsupervised learning algorithm, say K-means. Assume we use $k$ classes.

Afterwards, we observe the mean return values of the following days conditioned on the class. Meaning, for the days in class $k=1$ we have on average returns $\mu_1$ on the following day.

The statement then is the following, if a day looks like class $c$ then tommorrow you can expect returns $\mu_c$.

Here is my question:

Have we broken any rule in this process? Should we have stored some of the $N$ days for out of sample performance? It seems to me that even talking about out of sample and in sample is meaningless, since the algorithm only uses the vectors $x$ and is absolutely agnostic about the returns on the following day.

## Answer by RRG (score 1, accepted)

https://quant.stackexchange.com/a/35558

You form your clusters on in-sample data. How well new data conforms to these clusters going forward can only be determined from out-of-sample data. So you are estimating the generalisation of your clusters.

If you don't separate the data then you might find clusters which only appeared in the future, and any associated forward returns would have been linked to those clusters.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.