Why Yield Curve PCA Factors May Not Match Classic Curve Spreads
Summary
The document examines why principal components extracted from yield curve data may correlate less strongly with familiar level, slope, and curvature measures than a reference study suggests. The questioner fits PCA to par yields and compares the resulting component scores with selected yield and spread series. The reported correlations are high for the first two components but lower for the third, which is intended to represent curvature.
The response points to a likely mismatch between the maturities in the PCA data and those used to construct the comparison butterfly spread. It recommends matching maturities to the observed third-component loading shape and trying alternative maturity combinations. It also advises checking the variance explained by each component and whether PCA was applied to yield levels or changes at a particular sampling frequency. The discussion offers plausible diagnostics rather than a definitive reanalysis; results depend on the dataset, maturity coverage, transformations, and sample period. It further notes that convexity and historical-data choices can affect interpretations used for hedging.
Key ideas
- PCA components need not align closely with standard level, slope, and curvature spreads when the maturity choices differ.
- Compare the butterfly spread’s maturities with the loading pattern of the third principal component.
- Check the variance explained by each component when interpreting a low curvature correlation.
- The results can depend on whether the PCA uses yield levels or changes and on the sampling frequency.
- Historical data choices and convexity can affect how PCA factors are applied to hedging.
Tags
Full text
# Correlation between Yield Curve PCA components and Level - Slope - Curvature
# Correlation between Yield Curve PCA components and Level - Slope - Curvature
I found a piece of the following book:
Yield Curve Modeling and Forecasting
That states that when extracting the PCA components from the yield curve and projecting the data along these axis we find a very strong correlation between:
- Yield projected along the first axis and the 10Y yield (shift)
- Yield projected along the second axis and the 10Y-6M spread (slope)
- Yield projected along the third axis and the 6M+10Y-2*5Y butterfly spread (butterfly)
I know this is "classic" theory of yield curve modelling, but I wanted to check it by myself. The data used can be found at the following link: Nominal Yield Curve - Board of Governors. I used the SVENPYXX columns (Par Yield) from 1st Jan 85 to dec 08 as stated.
I ran the following code:
```
from sklearn.decomposition import PCA
import pandas as pd
from scipy.stats import pearsonr
#Read File and drop na_rows
df=pd.read_excel('feds200628.xlsx',index_col=0).dropna()
#Fit a PCA on 3 components
pca = PCA(n_components=3)
pca.fit(df)
#Projection on PCA components
projected = pd.DataFrame(pca.fit_transform(df))
eigen_vectors=pd.DataFrame(pca.components_).transpose()
eigen_vectors.plot()
#Correlation analysis
correls=[]
correls.append(pearsonr(projected[0],df['SVENPY05'])[0])
correls.append(pearsonr(projected[1],df['SVENPY10']-df['SVENPY01'])[0])
correls.append(pearsonr(projected[2],
df['SVENPY01']+df['SVENPY10']-2*df['SVENPY05'])[0])
```
I get the following plot for my PCA components
Which is the "classical" Shift-Twist-Butterfly graph - but my correlations are way lower than what we can see in the book.
As you can see on page 12, projected data and respectively 10Y yield, 10Y - 0.5Y spread, 10Y+0.5Y-2*5Y seem to be "strongly" correlated:
Diebold and Li Forecasting Term Structure through a similar approach also find strong correlations (0.97, -0.99, 0.99):
My correlations are the following: -0.99: Satisfying -0.96: Satisfying -0.44 => ??
Not sure what I do wrong, and why do I find such a low correlation with my butterfly component. Note that instead of using the 0.5Y I used the 1Y as the 0.5 was not in the data.
## Answer by r-learning-machine (score -1)
https://quant.stackexchange.com/a/82034
You need to think about how the given strategies align with the data (including the range of maturities) as well as your results:
- it looks like you conducted your PCA over maturities up until 8 years. But your trading strategy references up to the 10Y.
- think closely about 10Y+0.5Y-2x5Y and how it matches up with your graph for the loadings of the 3rd PC. Curvature (the 3rd PC) reaches its lowest in the 2 to 4 year range, not the 5 year range and your PCA seems to cover up to 8Y not 10 years. Specifically, you should try 8Y+6M-2x4Y and see if it increases the correlation. You may also try other ranges and see what fits your data better, 8Y+[6M,1Y]-2[3.5Y,5Y]. Since the paper used the 6M and you only have the one year, you can try to interpolate to the 6M and check that although it looks like you may have done that already.
- what was the proportion and cumulative variance for the first 3 PCs? Were you using daily changes, weekly changes? I’m not that surprised that you are not seeing a super high correlation with the butterfly spread, I’ve seen weirder things in the past. Especially if your 3rd PC explains very little variance. I’d be more concerned about the correlations with the first 2 PCs as they tend to explain a very high proportion of the variance.
- for practical purposes: there is a brand new video out that discusses regression and PCA for hedging, which is highly relevant to the question you asked. Many of the concepts discussed in the video including historical data dependency, convexity (which may not fully be captured in the given strategies), using yield over level changes (i.e., discussion relating to spurious correlation and advantages, disadvantages of using level vs slope changes) may help you digest the differences you are seeing with respect to your data and the original paper you were referencing.Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.