Skip to content
All library documents

Interpreting PCA Variance Plots in Stock Prediction

Article Quant Q&A · Author: GRS

Summary

The question describes applying principal component analysis to many correlated stock features while trying to predict the FTSE 100. It raises confusion about a result suggesting that no components explain 99% of variability, whether additional features such as recent gradients might help, and what a test target labeled Y_TEST represents. It also asks how a time-series model can use historical observations without continually expanding its input size.

The answer offers only a tentative interpretation: the plotted result may show cumulative variance explained by components ordered from largest to smallest. It also notes that Y_TEST commonly denotes either test-set labels or predictions, depending on the code. The response does not inspect the actual plot, data preprocessing, model, or train-test setup, so these explanations are possibilities rather than a diagnosis. It provides no forecasting method or evidence that PCA improves equity index prediction.

Key ideas

  • A PCA plot may show cumulative variance explained across components sorted by their contribution.
  • The meaning of a variable named Y_TEST depends on the surrounding implementation.
  • Y_TEST can refer to test-set outcomes or model predictions on the test set.
  • The answer cannot diagnose the PCA result or time-series design without the plot and procedure.

Tags

Full text
# Using PCA to predict Stock Prices


# Using PCA to predict Stock Prices












I would like to model an index e.g. FTSE 100. I have a list of all companies that make up the index and their stock values (daily high, low, volume, close values). In total, this time series data has 400+ features.

I would like to build a neural network similar to this one, but first I ran a Pearson correlation and found that there is high correlation between stocks.

I want to predict the value of the FTSE 100.

First of all, I scaled the data and then applied PCA to remove correlations and found that 0 components account for 99% of variability. Now my columns look like the following:

```
Date         FTSEOpen FTSEClose Stock1Open Stock1Close Stock2Open Stock2Close
01/01/2006   2880     2890      144        130         300        333
...
08/01/2018   3862     3851      204        311         134        154
```

I did PCA on all columns, and got the following result (which doesn't make sense!)

I am currently thinking of adding extra features such as 5day gradients etc. to mix things up.

Also, I'm not sure what Y_TEST is. I understand this is next days data, but I'm still trying to understand what the network input is. If I use all past data, the input dimensions keep increasing with every day.

Let's say I computed PCA, now I have just 1 vector...(1 date column and 1 PCA vector), this now looks like very little data to actually predict stock prices.

## Answer by madilyn (score 2)

https://quant.stackexchange.com/a/37654

Without knowing everything about your procedure, your graph is most probably % cumulative variance explained against each principal component, sorted starting from the largest. See this SE post for a similar plot.

It is conventional to refer to your target vector as $\bf{y}$ so one can only guess that `Y_TEST` refers to either the labels on the test set (out of sample) or your predictions on the test set.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.