Using Squared PCA Loadings to Assess Original Factor Contributions
Summary
The document explains how to relate principal components back to the original variables after principal component analysis. Once eigenvalues identify components that account for a chosen share of total variance, the entries of each corresponding eigenvector indicate how the original inputs contribute to that component.
The suggested measure is to square each eigenvector entry. Because the squared entries sum to one, they can be read as normalized contributions to that eigenvector, and therefore used to inspect which original variables are associated with a selected component. Plotting these values across components can show how the input contributions change. The answer does not provide an example dataset or a procedure for ranking factors across multiple components into a single cumulative variance share. It also does not establish that a rotation method such as Jacobian rotation is more suitable; the practical guidance is to inspect squared loadings for the components already selected.
Key ideas
- Select significant components using their eigenvalues and a variance threshold or another criterion.
- For a given component, square each eigenvector entry to measure normalized input contributions.
- The squared entries of an eigenvector sum to one.
- Comparing squared loadings across selected components can show how original variables contribute.
Tags
Full text
# After PCA on original factors, how to tell which original factors are dominant? # After PCA on original factors, how to tell which original factors are dominant? When doing the PCA analysis, you end up with eigenvalues which are ordered by how much variance they explained for each eigenvector. Say, the eigenvectors since they are orthogonal, do not represent the reality - I want to stick with the original factors only, but determine which are the most important and which add up to providing, say, 90% of the variance? Also, is Jacobian rotation more conducive to finding dominant original factors? ## Answer by John (score 2, accepted) https://quant.stackexchange.com/a/11371 When I use PCA, I follow a few typical steps. First, I would apply PCA to the covariance matrix, I would then designate certain eigenvalues as dominant or significant (such as by those that contribute up to $x\%$ of variance or by RMT), and then I would identify the eigenvectors that match up with those significant eigenvalues. I think you're with me at this point. It appears you want to know how to determine which of the inputs to the covariance matrix match up with the eigenvectors (i.e. how much does year $Y$ contribute to eigenvalue $N$, in your example). One way to make that determination is to square the eigenvector $N$. This squared eigenvector should sum up to $1$. Thus, you can consider each of these squared values like a percent of contribution to the eigenvector (and thus the eigenvalues). For your example, you could plot these squared values against the years for each to get a sense of how it changes as you change eigenvectors.
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.