Preparing Return and Factor Data for PCA Across Changing Volatility Regimes
Summary
The discussion addresses whether stock returns and estimated factor exposures should be combined for principal component analysis, and whether each dataset should be standardized first. The replies emphasize that the answer depends on the intended model, beliefs about whether stock and factor volatility is stationary, and how much data is needed to estimate covariance reliably. A single covariance estimate across a long sample may mix distinct market regimes, so the sample window and volatility scaling matter.
One suggested approach is to scale each stock or factor on its own timescale, reducing observations to innovations before estimating covariance. The other reply notes that PCA on a correlation matrix implicitly uses standardized variables, since correlation is equivalent to covariance computed from standardized data. It also suggests considering factor analysis when a model explicitly separates common factor effects from idiosyncratic residuals. The thread gives no universal standardization rule or empirical comparison; choices should reflect the data structure, target horizon, and assumptions about changing volatility and correlations.
Key ideas
- PCA preparation depends on the model and assumptions about volatility stationarity.
- A long sample can combine distinct covariance regimes and distort estimates relevant to the current horizon.
- Scaling stocks or factors separately can express observations as innovations before covariance estimation.
- PCA on a correlation matrix is equivalent to applying covariance analysis to standardized variables.
- Factor analysis may suit models that separate common exposures from idiosyncratic risk.
Tags
Full text
# Pca on multidimensional data # Pca on multidimensional data My data has stock returns over n periods for x stocks and m factor exposures for each stock ( ex: value, momentum) for n periods(output of regressions ) . Can I club this data together and then compute the correlation matrix (x by x matrix) and then run pca Do I need to standardize each data set separately ( returns and factor exposure)? ## Answer by lehalle (score 2) https://quant.stackexchange.com/a/61291 It is matter of choice of model - do you believe that the volatility of each of your stock / factors is stationary? - what is, according to you, the needed number of observation to compute the coef of your covariance matrix with accuracy? For instance, results on rough volatility suggest that if you want to use you volatility estimate over the next $N$ days, you should use the last $N$ days to estimate it. It is probably a lower number of days than the ones you want to estimate your covariance... Moreover, for the covariance: do you really want to use days spanning from 6 months before the financial crisis (say early 2008) to one year later (says 2010)? This (relatively short) period is made of very different covariance regime... Choose your model, can be one different scale for each stock / factor, hence you reduce the returns to their "innovation", and after that you compute the covariance of what you have. If you want more formalized details, I suggest you have a look at Practical Volatility and Correlation Modeling for Financial Market Risk Management, by Torben G. Andersen, Tim Bollerslev, Peter F. Christoffersen, and Francis X. Diebold. ## Answer by stans (score 0) https://quant.stackexchange.com/a/60599 When running PCA on the correlation matrix, you do not need to standardize the data. Correlation is the same as covariance of the standardized data. You may also want to experiment with other types of factor analysis since model $$ R = F\beta + \varepsilon, $$ where idiosyncratic risks $\varepsilon$ are unexplained by any factors, makes more sense. For example, you can try maximum likelihood factor analysis or principal axis factoring. Stata and SPSS are quite convenient for this purpose. R has it as well, of course.
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.