Using PCA and Incremental SVD to Track Cross-Asset Market Moves
Summary
This discussion considers how to use principal components from a cross-asset ETF portfolio to describe future market moves and assess whether factor weightings remain stable. One response warns that transposing a data matrix with many time observations relative to assets can create zero eigenvalues in the resulting covariance or correlation matrix. It recommends singular value decomposition as a way to identify days with similar loading patterns and notes that random matrix effects, including the Marchenko–Pastur framework, may matter.
A second response suggests incremental SVD: update component scores as each new market observation arrives, rather than repeatedly calculating the decomposition from scratch. The discussion offers methods and cautions, but no empirical results or complete procedure for selecting and validating forward-looking factor weights. It also notes that projecting orthogonal components into future market moves requires additional design choices.
Key ideas
- Transposing a time series with many dates relative to assets can produce zero eigenvalues in the resulting covariance matrix.
- SVD can reveal days that load similarly on principal components.
- Incremental SVD can update component scores as new market observations arrive.
- Orthogonal principal components still require a separate method for projecting into future returns.
- Random matrix effects may complicate interpretation of estimated components.
Tags
Full text
# Optimizing Principal Component factor weightings over time
# Optimizing Principal Component factor weightings over time
I was given the returns of a cross-asset class portfolio of ETFs and I conducted PCA to obtain factors on dates from T-n, T-3, T-2,..., T. What I would like to do is decompose the market moves from T+1, T+2, ... onwards into combinations of the PCs.
My questions is, what sort of algorithm or optimization method can I use to obtain a set of factor weightings that explain the maximal amount of variance in the market moves going forward. Furthermore, this optimization method should be capable of outputting a list of maxima , not just one, so that the stability of the optimal factor weightings can be assessed over time, since the optimal factor weightings for one day/time period may not be the same as the optimal factor weightings over a larger period.
## Answer by user6430 (score 1, accepted)
https://quant.stackexchange.com/a/9888
Your approach is a good one. But before you venture too far, you should be aware of issues related to zero eigenvalues (positive semi-definiteness) of your correlation matrix $\mathbf{R}$ or covariance matrix $\mathbf{C}$. Let $p$ be the number of assets, and $t$ the number of, for example, day or bars. You probably have many more times in the time series than you do assets, and thus, $t\gg p$. Since you are turning the dataset on its side, and treating days like variables, and assets like observations, there will be $t-p$ zero eigenvalues in your covariance matrix $\mathbf{C}$.
Given the above, don't ever lose sight of the basic premise of, for example, Markowitzian portfolio optimization, where the covariance between assets (not days) is determined, and the number of days is exceedingly large (250 trading days per year) when compared with the number of assets. By default, Markowitz did not need to worry too much about zero eigenvalues of $\mathbf{C}$ when introducing his theory since $t\gg p$.
As a solution, you could therefore use singular value decomposition (SVD) on your $\mathbf{C}$, which is ideal for large dimension and low example datasets. When done, what you will observe is that certain days (bars) will "load" on certain principal components -- so you essentially be identifying groups of trading days which are similar. Since principal components are by definition orthogonal (zero correlation between them), you will need to think of a way to project your results into the future. Overall, however, there will be high merit in what you are doing for the ETFs involved, since you will be able to understand the structure of trading days for the basket of ETFs and days used.
It's a complex undertaking, and since so many issues are involved surrounding the problem of $t \gg p$, you will undoubtedly run into the Marcenko-Pastur law involving $\gamma=p/t$ and random data matrices. Look at, for example, Johnstone's talk
## Answer by ashleyA (score 2)
https://quant.stackexchange.com/a/9934
Have you considered using 'incremental' singular value decomposition to calculate your component scores? Each future market move (or increment) forces a recalculation of component scores given the new data.
This paper outlines an algorithm to do this Fast Low-Rank Modifications of the Think Singular Value Decomposition
> This paper develops an identity for additive modifications of a singular value decomposition (SVD) to reflect updates, downdates, shifts, and edits of the data matrix. This sets the stage for fast and memory-efficient sequential algorithms for tracking singular values and subspaces.Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.