Comparing Regression Residuals and Log-Ratio Spreads in Pairs Trading
Summary
The document compares two spread definitions used to trigger Z-score entries in automated pairs strategies: residuals from an ordinary least squares regression of log returns, and the log ratio of the two contract prices. The author reports that the return-residual series is noisy and seems uninformative, while the log-ratio spread has produced more reliable signals and risk management in their experience.
The question is whether the residual calculation is being applied correctly, whether smoothing it would help, and whether using 30-minute returns or a longer interval explains the difference. No data, formal comparison, or explanation of the two measures' statistical assumptions is supplied, so the reported performance is anecdotal. The discussion also does not establish that either spread is generally superior: results can depend on the pair, regression specification, hedge ratio, sampling, and how the series is normalized. It is a useful prompt to distinguish price-level relationships from return regressions and to validate spread construction out of sample before using threshold signals.
Key ideas
- The author compares return-regression residuals with a log ratio of paired contract prices.
- The return-residual spread is described as noisy, while the log-ratio spread is reported to behave more reliably in the author's strategy.
- The strategy opens positions when a spread crosses a Z-score threshold.
- Changing the candle interval did not resolve the reported noise.
- The document raises questions about smoothing and correct spread construction but provides no formal test or general conclusion.
Tags
Full text
# Is my spread calculation correct? # Is my spread calculation correct? I have automated Pair trading strategies running on both CME Futures and cryptocurrency perpetual pairs. I can chose between different spread calculation type and I noticed that the one I thought was the more "efficient" was in fact really bad. All position entries are triggered when the spread cross over or under a Zscore threshold. Here one of the few calculation type I use: - Spreased based on the residuals of the Ordinary Least Square of the Log returns of the two contracts -> I find this one extremly noisy and in fact really not meaningful, so my question is: Am I doing it correctly ? Should I smoothen the spread signal by averaging it ? The Log return is currently calculated based on the return between two successive "candles" of 30min timeframe. Increasing the timeframe doesn't solve the issue really. - Spread based on a simple Log(y/x) where x is the independant variable and y is the dependant variable. This one is quite reliable and gives really good returns and risk management. Would you be kind to lead me in a good direction, maybe I'm doing something wrong ? Thanks !
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.