Ranking Kalman Filter Pairs by Beta Stability and Prediction Error
Summary
The document considers how to rank pairs of securities modeled with a time-varying linear relationship. In a state-space regression, a Gaussian Kalman filter produces state covariance estimates and prediction errors. The example applies this setup to daily returns for QQQ and XLK, plots the filter outputs, and compares the idea with a more volatile VXX and TLT relationship. The proposed goal is to favor pairs whose estimated beta is stable and whose prediction errors are small.
It does not provide a finished scoring formula or demonstrate that any ranking predicts future trading results. The example is code for calculating and visualizing filter quantities, while the question of how to combine them into a comparable numeric criterion remains open. Since covariance and prediction error can depend on units, scaling, model specification, and time period, a practical ranking would need normalization and out-of-sample evaluation. The document frames a research question rather than validating a pairs-trading selection method.
Key ideas
- A Kalman filter can estimate a time-varying beta in a state-space regression.
- State covariance and prediction error are candidate measures of beta stability and fit.
- The example compares QQQ and XLK with the more volatile VXX and TLT pair.
- The document does not specify a scoring rule for ranking candidate pairs.
- Any ranking should account for scale and be evaluated out of sample.
Tags
Full text
# Good criteria to sort state-space $\beta_{t}$ according to Kalman filter output
# Good criteria to sort state-space $\beta_{t}$ according to Kalman filter output
Let's assume the usual state-space linear model without constant term for simplicity:
$y_{t}=\beta_{t} X_{t}+\epsilon_{t}$
If we apply Gaussian Kalman filter to estimate $\beta_{t}$ we get $P_{t}$, which is the covariance matrices of predicted states, and $v_{t}$, which is the prediction error.
The following simple `R` code allows you to download pair of tickers (`QQQ` and `XLK` for instance) from Yahoo Finance and estimate $P_{t}$ and $v_{t}$ while plotting them:
```
# ======================================== #
# Kalman filter errors and states variance #
# ======================================== #
op <- par(no.readonly = TRUE)
Sys.setenv(TZ = 'UTC')
# Contents:
# 1. Installing packages
# 2. Loading packages
# 3. Downloading and plotting data
# 4. Kalman filtering of linear regression Beta
# *********************************
# 1. Installing packages
# *********************************
#install.packages('KFAS')
#install.packages('latticeExtra')
#install.packages('quantmod')
# *********************************
# 2. Loading packages
# *********************************
require(compiler)
require(latticeExtra)
require(KFAS)
require(quantmod)
# *********************************
# 3. Downloading and plotting data
# *********************************
Symbols <- c('QQQ', 'XLK')
getSymbols(Symbols, from = '1950-01-01')
data <- na.omit(merge(Cl(QQQ), Cl(XLK)))
colnames(data) <- Symbols
xyplot(data)
# *********************************
# 4. Kalman filtering of linear regression Beta
# *********************************
y <- na.omit(merge(ClCl(QQQ), ClCl(XLK)))[,1]
X <- na.omit(merge(ClCl(QQQ), ClCl(XLK)))[,2]
model <- regSSM(y = y, X = X, H = NA, Q = NA)
object <- fitSSM(inits = rep(0, 2), model = model)$model
KFAS <- KFS(object = object)
P <- xts(as.vector(KFAS$P)[-1], index(y))
v <- xts(t(KFAS$v), index(y))
Z <- cbind(P, v)
colnames(Z) <- c('Covariance of predicted state', 'Prediction error')
xyplot(tail(Z, 1000))
```
Now let's iterate this procedure over several pairs of securities to estimate their $\beta_{t}$, $P_{T}$ and $v_{t}$ and you want to sort these pairs by the stability and accuracy of $\beta_{t}$, that is, low variance and low prediction error.
I would like to know suitable criteria to make this ranking system having available $P_{t}$ and $v_{t}$, i.e. how to penalize a linear relationship because of too high variance and prediction errors?
For instance:
Replacing `QQQ` and `XLK` in my code with `VXX` and `TLT`, you will see greater $P_{t}$ and $v_{t}$, which are linearly related between `VXX` and `TLT` and are more volatile and have weaker predictive power than the one between `QQQ` and `XLK`.
This is similar to a ranking system and I would like to know how to produce some numeric criteria.Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.