Skip to content
All library documents

Standardizing Returns Before PCA for Pairs Trading Clusters

Article Quant Q&A · Author: Lakshya

Summary

The document discusses a preprocessing choice in a pairs trading workflow that applies principal component analysis (PCA) to ETF returns and clusters the resulting factor loadings. The question compares a paper’s stated procedure, which standardizes returns before PCA, with an implementation that appears to fit PCA on unscaled returns and scale afterward. The answer argues that PCA on variables with very different variances can make the leading components reflect high-volatility series, potentially affecting the clusters.

The suggested experiment is to standardize the returns matrix before fitting PCA, then inspect the clustering results. The response also cautions that scaling components after PCA can give low-variance components influence comparable to high-variance ones, undermining PCA’s variance-based ordering. These are methodological observations, not a reproduced study: the document provides no controlled comparison, performance results, or evidence that standardization will recover the paper’s findings. The suitable preprocessing may depend on whether volatility itself is intended to shape the clustering.

Key ideas

  • PCA on unscaled returns can be dominated by assets with higher variance.
  • Standardizing returns before PCA is proposed as an experiment for the clustering workflow.
  • Scaling principal components after fitting PCA may alter their relative influence.
  • The question raises a trade-off between equalizing inputs and retaining volatility differences as clustering information.
  • The recommendation is not validated by comparative performance evidence in the document.

Tags

Full text
# Standardising before PCA for clustering


# Standardising before PCA for clustering












https://premio-vidigal.inesc.pt/pdf/SimaoSarmentoMSc-resumo.pdf Referring to this paper and author's implementation here https://github.com/simaomsarmento/PairsTrading I am looking to recreate this paper but something is going wrong. When I take the ETFs mentioned here and clean the data, it doesn't perform as the author claimed. Primarily, my clustering step was failing miserably. I did some research and looked at their code. It turns out that the author is scaling "after" fitting PCA on a returns matrix. The PCA factor loadings are then used in a clustering algorithm. Now the paper that has been provided, states that returns are standardised beforehand. This is also somewhat unintuitive for me since we lose out on a potential clustering factors i.e. std deviation. Which I feel is a necessity for good clustering to detect pairs trading.

## Answer by IsmaelGoodpeople (score 2)

https://quant.stackexchange.com/a/85388

Looking quickly at the github, I See this part of code :

```
X, explained_variance = series_analyser.apply_PCA(N_PRIN_COMPONENTS, df_returns, 
                                                  random_state=0)#12)
```

So PCA was performed directly on the returns matrix, This will capture the the first components with the highest volatilies which i think can distort the clustering, you can only perform PCA directly on the original vector if it's components have the same Variance.

Now you can try using standarscaler before applying PCA and look at your results.

N.B : scaling after fitting PCA is somewhat counter-productive as you make the less variable components at the same scale as the most variable ones which i guess voids PCA, but i can be corrected on this.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.