Skip to content
All library documents

Choosing PCA Components for Correlated Return Predictors

Article Quant Q&A · Author: NickF

Summary

The document raises two questions about using principal component analysis on index and sector returns as predictors of a single stock’s return. It asks whether the first component, or a set of leading components, should explain a minimum share of variance before PCA is worthwhile. It also asks whether PCA should be applied only to a subset of highly correlated predictors, then combined with the remaining original variables for regression.

The questioner describes an alternative workflow: run PCA on all predictors, regress the outcome on a few leading components, and use their regression coefficients with the eigenvectors to infer sensitivities to the original variables. This leaves open whether the individual original factors are significant, since significance was assessed only for the retained components. The document contains no answers, data, or empirical comparison, so it offers no universal variance threshold and does not establish which workflow is preferable. Its useful contribution is identifying the distinction between component-level regression results and inference about original factor sensitivities.

Key ideas

  • The document asks whether PCA has a minimum explained-variance threshold for retaining components.
  • It considers applying PCA only to a subset of correlated return predictors.
  • A proposed workflow regresses the outcome on leading components and maps loadings back to original variables.
  • Significant components do not by themselves establish that every original factor is significant.
  • No threshold or preferred modeling workflow is established in the document.

Tags

Full text
# Is there a considered floor for variation the 1st principal component must explain?


# Is there a considered floor for variation the 1st principal component must explain?












I am wondering if there is a considered floor to the percentage variation the 1st principal component must explain in general for PCA - ie. any lower and it is not worth doing PCA at all? Is the floor near 75%, 80% or should the 1st 3 explain a minimum of 90% or what?

As a follow on, if I have 10 X variables (index & sector returns) and only 6 are highly correlated (I take a correlation above 0.8 to be highly correlated - or is that too high?) should I just do PCA on those 6, then combine the 1st two principal components with the 4 remaining original variables and use that as my X for regression?

What I was doing was doing PCA on all 10 variables, taking the 1st 2 or 3 principal components, regressing those on Y (which is a single stock's return) then taking those PC's betas and matrix multiplying them by the eigenvectors to back out sensitivities to the original 10 factors but I am left with the situation of not knowing if all 10 factors are significant (all I know is that the 1st 2 PCs are significant)

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.