Using a Kalman Filter to Estimate a Time-Varying Pairs Hedge Ratio
Summary
The document presents an attempted Kalman filter implementation for estimating a changing hedge ratio between two cointegrated stock log-price series. The proposed state is a single beta that evolves over time, with the second stock’s log price treated as the observation and the first stock’s log price used in the observation relationship. The estimated beta is then used to calculate a spread for a pairs trade.
The response flags implementation issues involving the observation matrix’s dimensions, the use of the update method versus filtering a full series, and matching state dimensions to the filter’s expectations. These are practical cautions, but the response does not provide a corrected implementation or establish that the proposed model is statistically appropriate. It also gives no evidence about trading performance, parameter selection, or whether the resulting spread remains mean reverting out of sample.
Key ideas
- A Kalman filter can be used to estimate a hedge ratio that changes over time between two assets.
- The example models one beta as the hidden state and uses paired log prices in the observation relationship.
- The observation matrix and state arrays must match the Kalman filter’s expected dimensions.
- The answer raises concerns about filtering method usage but does not provide a complete corrected implementation.
- A time-varying hedge ratio alone does not demonstrate profitable or persistent pairs-trading behavior.
Tags
Full text
# Kalman Filter Implementation using PyKalman
# Kalman Filter Implementation using PyKalman
I am trying to apply a simple Kalman filter to pair trading. My underlying stock pair is cointegrated with no constant term. Can someone kindly advise if i am going on the right track with the Kalman filter?
```
from pykalman import KalmanFilter
import numpy as np
import pandas as pd
# Log prices of stock_1 and stock_2
stock_1_price = np.log(df_combined[f"Close_{stock_1}"]) # independent variable
stock_2_price = np.log(df_combined[f"Close_{stock_2}"]) # dependent variable
# Transition matrix for beta (1x1 identity matrix because we're only estimating beta)
transition_matrix = np.eye(1) # 1x1 identity matrix
# Observation and transition covariances
observation_covariance = 1 # Observation noise, can be tuned or optimized
transition_covariance = np.array([[0.001]]) # Process noise covariance (small to keep beta stable)
# Initial estimates
initial_state_mean = np.array([0]) # Initial guess for beta
initial_state_covariance = np.array([[1]]) # Initial covariance
# Set up the Kalman Filter
kf = KalmanFilter(
transition_matrices=transition_matrix,
initial_state_mean=initial_state_mean,
initial_state_covariance=initial_state_covariance,
observation_covariance=observation_covariance,
transition_covariance=transition_covariance
)
state_means = np.zeros(len(stock_2_price)) # beta over time
state_covariances = np.zeros(len(stock_2_price))
state_mean = initial_state_mean
state_covariance = initial_state_covariance
# Run the Kalman filter manually in a loop
for t in range(len(stock_2_price)):
observation_matrix = np.array([[stock_1_price[t]]]) # stock_1_price for beta
state_mean, state_covariance = kf.filter_update(
state_mean,
state_covariance,
observation_matrix=observation_matrix,
observation=stock_2_price[t]
)
state_means[t] = state_mean[0]
state_covariances[t] = state_covariance[0, 0]
# Extract the hedge ratios (beta)
hedge_ratios = state_means # beta values over time
# Calculate the spread using the estimated hedge ratios
kalman_spread = stock_2_price - hedge_ratios * stock_1_price
plt.plot(kalman_spread)
```
```
## Answer by Sane (score 1)
https://quant.stackexchange.com/a/80487
A few considerations:
In your code, the `observation_matrix` is a 1x1 matrix based on `stock_1_price[t]`, which is incorrect. For the Kalman filter to work correctly, you need to ensure that the `observation_matrix` and `observation` are correctly sized. In your case, since you're estimating a single `beta`, the observation matrix should be a 1-dimensional array rather than a 2D matrix.
The `kf.filter_update` method is typically used to update the state estimate in each iteration. However, this method is meant to be used within the Kalman filter's internal loop, and you should use the `kf.filter` method for the entire time series data.
You should ensure that the dimensions and initializations of `state_mean` and `state_covariance` match the expected dimensions of your state vector.Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.