Interpreting PCA Factors in Statistical Risk Models
Summary
This discussion distinguishes named factors, defined by observable economic or financial variables, from statistical factors extracted from asset returns with principal component analysis (PCA). The first PCA component often resembles a broad market factor, while later components may reflect groups of assets that moved together in the chosen sample. PCA identifies patterns of co-movement but does not explain their economic cause.
The answers note that statistically identified factors can vary with the data and time period, which makes them harder to interpret and may limit their usefulness for studying persistent sources of risk. They also suggest that PCA can be robust to small data changes, though the discussion offers no supporting test or details about how robustness should be assessed. The takeaway is to distinguish stability under small input changes from stability across different market periods; the latter is not established here.
Key ideas
- Named factors correspond to defined economic or financial variables, while PCA factors are inferred from return data.
- The first PCA component often tracks broad market movement.
- Later components capture sample-specific co-movement and may lack a clear economic interpretation.
- PCA factors can depend on the selected data and period, so their stability should be assessed.
- The discussion claims robustness to small data changes but supplies no empirical evidence.
Tags
Full text
# Answer by nbbo2 (score 1) # What happens if my risk factor caught by statistical risk model using PCA turns out to be totally different from other PM's risk factor? In order to explain systematic risk we use risk factors and I've learned that since they try to explain 'systematic' risk, risk factors are relatively well-known. However, what happens if the risk factors I got from PCA (Statistical Risk Model) turn out to be really different from other PM's risk factors who also used PCA to get risk factors. Depending on which data to use, and for which period to analyze, it seems likely to have different risk factors according to individual PM. It seems that if different risk factors try to capture the same concept 'systematic risk', it no longer captures systematic risk anymore since they use 'different' factors. Is there any part I misunderstand? ## Answer by nbbo2 (score 1) https://quant.stackexchange.com/a/62025 There are two kinds of factors. Named or defined factors are related to observable economic or financial variables, such as FamaFrench HMB, or the market factor or an oil price factor. Unnamed or statistically identified factors are the result of a PCA using only stock prices. Although the first PCA factor is usually close the Market factor mentioned above (i.e. overall movement of all stocks), the other factors are not easy to describe, and they are sensitive to the time period studied. Basically the PCA algorithm is picking up that some stocks are moving together during this period but cannot tell us why, and in another period the same comovement may not occur. To many, that is a disadvantage of using PCA factors in stock research. ## Answer by Bob Jansen (score 0) https://quant.stackexchange.com/a/62024 If I understand correctly your question is: #### Question If I ran PCA and someone else runs PCA and the result is completely different then what good is PCA? Does it make sense to use it for risk management? #### Answer If you would get a different outcome due to small changes in the data it would be bad measure. However, what people (hopefully) found that use PCA in their work is that the method is robust to small changes in the data they use.
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.