Skip to content
All library documents

Why PCA Factors Can Remain Highly Correlated Across Assets

Article Quant Q&A · Author: Baba_Yaga

Summary

The document describes a factor-research workflow: build long-short strategy returns from cross-sectional z-scores of predictive factors, run principal component analysis on the time series of those returns, and retain a subset of components. The researcher then applies the component transformation to the asset-level factor matrix and observes that the resulting features are highly correlated across assets, despite the components having low correlation in the PCA dataset.

The text poses this discrepancy but supplies no answer, derivation, data, or proposed diagnostic. A key issue for analysis is that PCA’s low-correlation property applies to the observations and covariance matrix used to fit it; it does not by itself ensure low correlation in a different transformed dataset or across a different dimension. The document is therefore useful as a research question about the scope of PCA decorrelation, but it does not establish why the observed pattern occurs or how to correct it.

Key ideas

  • The workflow applies PCA to historical returns from long-short strategies built on predictive factors.
  • The PCA components are then used to transform asset-level factor observations.
  • Low component correlation in the PCA input does not guarantee low correlation in a distinct transformed dataset.
  • The document presents the observed correlation pattern as a question and provides no resolution or empirical evidence.

Tags

Full text
# High correlation between aggregated features constructed with principal components


# High correlation between aggregated features constructed with principal components












I have $k$ predictive factors constructed for $N$ assets using differing underlying data sources. For a given date, I compute the daily returns over a lookback window of long/short strategies constructed by sorting these factors. The long/short strategies are constructed in a simple manner by computing a cross-sectional z-score. Once the daily returns for each factor are constructed, I run a PCA on this $T\times k$ dataset (for a lookback window of $T$ days) and retain only the first $m$ principal components (PCs).

Generally I see that, as expected, the PCs have a relatively low correlation. However, if I were to transform the predictive factors for any given day using the PCs i.e. going from a $N \times k$ matrix to a $N \times m$ matrix, I see that the correlation between the aggregated "PC" features is quite high. Why does this occur?

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.