How Scaling Choices Affect PCA Components
Summary
The document explains how centering and scaling affect principal component analysis (PCA), in response to a question about why a credit research report recommends normalization for credit default swap data but not for other products. PCA on unscaled variables gives more influence to variables with larger variance, so differences in units or scale can dominate the components. Standardizing variables can remove this scale effect, but it also gives low-variance variables equal influence, potentially allowing noise to shape the results.
The answer interprets scaling only some variables as a deliberate way to reduce their influence relative to unscaled variables, while noting that this changes the relative explanatory power in an arbitrary way. It gives no empirical comparison or detailed justification for the report’s specific CDS choice, and concludes that such mixed preprocessing is difficult to interpret and may be poor practice. The appropriate choice therefore depends on the data and research objective; the document does not establish that CDS variables should always be normalized or that one preprocessing rule applies universally.
Key ideas
- Unscaled PCA gives greater influence to variables with higher variance.
- Standardization reduces scale effects but can give noisy variables disproportionate influence.
- Scaling only selected variables changes their relative contribution to the components.
- The document argues that mixed scaling can be hard to interpret and should be justified by the analysis.
Tags
Full text
# Why normalize only data for CDSs for PCA? # Why normalize only data for CDSs for PCA? I'm reading a Credit Suisse Research Report on PCA. The report says that to preprocess the data, you should "Centre data (and normalize when considering CDS data)." Why would you only normalize data for CDS, and not for other products? ## Answer by Lucas Morin (score 1) https://quant.stackexchange.com/a/17380 The choice of normalization depends on your data set: Without normalization : variable with high variance will have more impact on the PCA. You will have size effects. For exemple if you have one variable in meters and the other one in kilometers the one in meters will have way more impact. To avoid that you can normalize but now every variable will have the same power of explnanation. Noise may have a lot of power explanation. Normalising some variable but not the others could be interpreted as a way to reduce the impact of a high variance variable without changing the order of power of explanation of the others. By doing this, the explanation power of the normalized variable will be arbitrary changed with respect to the order of power explanation of the others variable. Plus that this hybrid pre-process is shady to non-statistician people. I think this should be concidered bad practice.
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.