Interpreting PCA Variance Plots in Stock Prediction
Summary
The question describes applying principal component analysis to many correlated stock features while trying to predict the FTSE 100. It raises confusion about a result suggesting that no components explain 99% of variability, whether additional features such as recent gradients might help, and what a test target labeled Y_TEST represents. It also asks how a time-series model can use historical observations without continually expanding its input size.
The answer offers only a tentative interpretation: the plotted result may show cumulative variance explained by components ordered from largest to smallest. It also notes that Y_TEST commonly denotes either test-set labels or predictions, depending on the code. The response does not inspect the actual plot, data preprocessing, model, or train-test setup, so these explanations are possibilities rather than a diagnosis. It provides no forecasting method or evidence that PCA improves equity index prediction.
Key ideas
- A PCA plot may show cumulative variance explained across components sorted by their contribution.
- The meaning of a variable named Y_TEST depends on the surrounding implementation.
- Y_TEST can refer to test-set outcomes or model predictions on the test set.
- The answer cannot diagnose the PCA result or time-series design without the plot and procedure.
Tags
Full text
# Using PCA to predict Stock Prices
# Using PCA to predict Stock Prices
I would like to model an index e.g. FTSE 100. I have a list of all companies that make up the index and their stock values (daily high, low, volume, close values). In total, this time series data has 400+ features.
I would like to build a neural network similar to this one, but first I ran a Pearson correlation and found that there is high correlation between stocks.
I want to predict the value of the FTSE 100.
First of all, I scaled the data and then applied PCA to remove correlations and found that 0 components account for 99% of variability. Now my columns look like the following:
```
Date FTSEOpen FTSEClose Stock1Open Stock1Close Stock2Open Stock2Close
01/01/2006 2880 2890 144 130 300 333
...
08/01/2018 3862 3851 204 311 134 154
```
I did PCA on all columns, and got the following result (which doesn't make sense!)
I am currently thinking of adding extra features such as 5day gradients etc. to mix things up.
Also, I'm not sure what Y_TEST is. I understand this is next days data, but I'm still trying to understand what the network input is. If I use all past data, the input dimensions keep increasing with every day.
Let's say I computed PCA, now I have just 1 vector...(1 date column and 1 PCA vector), this now looks like very little data to actually predict stock prices.
## Answer by madilyn (score 2)
https://quant.stackexchange.com/a/37654
Without knowing everything about your procedure, your graph is most probably % cumulative variance explained against each principal component, sorted starting from the largest. See this SE post for a similar plot.
It is conventional to refer to your target vector as $\bf{y}$ so one can only guess that `Y_TEST` refers to either the labels on the test set (out of sample) or your predictions on the test set.Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.