Why OLS Pair Spreads Can Fail Out of Sample
Summary
The document asks why a spread formed from an in-sample OLS regression of one asset on another diverges out of sample, even though the assets appear cointegrated and their simple price difference stays relatively bounded. The proposed spread scales one series by the fitted regression coefficient before subtracting the other; the reported fit has a high R-squared, but the out-of-sample behavior does not revert as expected.
The response explains that a regression coefficient is chosen to fit a particular sample and may differ across periods. A relationship between assets can also change after events, so apparent cointegration and a visually linear scatterplot do not establish a stable hedge ratio. The document offers a caution rather than a diagnostic procedure: it does not examine the underlying data, establish that simple subtraction is a valid spread, or test for structural breaks. Its examples are limited to the questioner's reported results and do not show a general comparison of spread methods.
Key ideas
- An OLS hedge ratio is optimized for the sample on which it is estimated.
- A fitted coefficient may change across time, making an in-sample spread unreliable out of sample.
- Events can alter the relationship between two assets.
- A bounded simple difference does not by itself establish cointegration or a sound trading spread.
Tags
Full text
# Why is OLS based spread not reflective of actual difference? # Why is OLS based spread not reflective of actual difference? I'm trying to define and track the spread between two time series (data available here), for the purpose of learning pair trading basics. When running a cointegration test the two series seem to be cointegrated (using python's `statsmodels.tsa.stattools.coint` the results are `(-3.3744261541141616, 0.04527140003070947, array([-3.98694225, -3.38584874, -3.07883012]))`, and when generating a scatter plot the relation looks linear without much noise. Most posts I've read recommend computing the spread by regressing one time series on the other (for an initial 'train' time window) and then extracting the coefficient and using it to subtract between the two (after multiplying one of them by it). When I do that the results of the OLS are `rsquared=0.88 coeff=2.554344` (using python's `statsmodels.api.OLS`) but the spread on the time period that's outside of the one the OLS was computed on drifts apart without converging back. On the other hand when I subtract the two time series (beyond the initial period used to compute the OLS) the difference seems to be bounded narrowly. See following plot (Blue line is the OLS spread, Orange line is simple subtraction) My question is - given that the simple subtraction yields a pretty consistent spread (in and out of sample), why isn't the in-sample regression able to extrapolate well on out of sample data? ## Answer by spar7453 (score 1) https://quant.stackexchange.com/a/55456 If my understanding is correct, then you are asking why in-sample coefficient is not working for out-of-sample data. - It is hard to tell whether two stocks are co-integrating without specific reason. - Although Stock A and Stock B are co-integrating with some coefficient, if you took regression on some particular time period, you might get some different coefficient. This is because regression always tries to find of the best fit. - Possibly, there may be some events announced that change co-integrating relation.
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.