PCA Factors Describe Returns but Do Not Forecast Them on Their Own
Summary
The document explains why extracting principal components from constituent stock returns does not by itself forecast a particular stock’s next-day return. PCA represents each asset return as a weighted combination of components estimated from historical data. The components are uncorrelated and ordered by variance, but this decomposition describes covariance structure rather than predicting the future value of any component.
A regression of a stock’s historical returns on contemporaneous component scores estimates its relationship to those factors, but forecasting requires an independent method for predicting future factor values. In an equity setting, components may resemble a broad market factor and long-short portfolios; if the market factor can be forecast, that forecast could contribute to an asset-return estimate. The answer also warns that PCA loadings are not necessarily equivalent to stock betas. It offers no predictive performance results or out-of-sample validation, and suggests that forecasting the broad market with betas may be more direct.
Key ideas
- PCA decomposes historical returns into uncorrelated components ordered by variance.
- The decomposition estimates factor exposures but does not predict future component values.
- A next-day stock forecast requires a separate forecasting method for the components.
- PCA loadings should not automatically be interpreted as conventional stock betas.
Tags
Full text
# Forecasting next day return of a stock using PCA of index constituents
# Forecasting next day return of a stock using PCA of index constituents
I am trying to predict the return of BN4.SI ( a singapore stock ) and part of Strait Times 30 component index using principal component Analysis. I have written my code in python.
My Question is i have got factors loading how can i predict the next day return of BN4.SI based on this PCA Factors.
Please help me.
The steps i have taken are
- Get the matrix of standardised returns of all stocks. Standardised log returns are [(x - mean)/std for x in array]
- Generate covariance matrix of the return matrix and that should give me a square matrix ( m X m)
- I am using numpy linear algebra eig function to calculate eigen values and eigen vectors. Since my returns are standardised then sum of eigen values should be equalled to number of components.
- Then sort the eigen values in decending order to get first 3 eigen values which explains almost 80% of the variance. This is shown in scree plot in the picture added below.
- Next step is to get top 3 PCA vector which is acheived by getting the dot product of eigen vector and the original return matrix data. pca1 = ti.np.dot(newDF,eVector1.reshape(-1,1)).reshape(1,-1) pca2 = ti.np.dot(newDF,eVector2.reshape(-1,1)).reshape(1,-1) pca3 = ti.np.dot(newDF,eVector3.reshape(-1,1)).reshape(1,-1)
- To perform linear regression with BN4.SI standarised return data. I need to get the matrix of the transpose of PCA vectors. np.column_stack([pca1.T,pca2.T,pca3.T])
- I am using sklearn to do linear regression. from sklearn import linear_model model = linear_model.LinearRegression() model.fit(ti.np.column_stack([pca1.T,pca2.T,pca3.T]),newDF["BN4.SI"]) model.score(ti.np.column_stack([pca1.T,pca2.T,pca3.T]),newDF["BN4.SI"]) print(model.coef_) print(model.intercept_) [ 0.2802088 0.37944899 0.13462393] -1.15582251414e-18
## Answer by Richi Wa (score 0)
https://quant.stackexchange.com/a/34665
I decided to write an answer in order not to write too many comments.
What do we get by PCA? Let us assume we have $n$ random variables. We get a representation of the data of the form $$ X_i = e_{1,i} P_1 + \cdots + e_{n,i} P_n $$ where the PCAs $P_1,\ldots,P_n$ are uncorrelated (not necessarily independent) and ordered (descending) by variance.
The above relation is estimated from past data.
In the world of interest rates $X_i$ could be the change (rather not the level) of vertex $i$ and it turns out that the change of the i-th rate can be decomposed into a parallel shift ($P_1$) and a steppening ($P_2$) and so on. This conditional on the future realization of $P_1$ we can estimate the change of vertex $i$ as it is estimated to be proportional by a factor $e_{1,i}$ to the parallel shift. The relation above does not tell us whether $P_{1,t+1}$ the future pf $P_1$ will be positive or negative.
In the case of stocks we usually take $X_i$ to be the return of asset $i$. And the principle components are a market factor and several long/short portfolios. If we can predict the market then we can plug this into the above equation.
Note that as far as I remember $e_{1,i}$ is not the beta of the stock in the stock setting. Thus we would get a forecast cheaper by looking a betas and forecasting the broad market.Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.