Choosing a Dependent Variable for Multi-Asset Mean-Reversion Spreads
Summary
This note asks how to choose the dependent asset when fitting an ordinary least squares hedge relationship among correlated equities. Using different stocks as the regression dependent variable produces different hedge ratios and therefore different spreads. The intended strategy is to sell a spread when it is above estimated fair value and buy it when it is cheap, on the assumption that the spread mean-reverts.
The author seeks an intuitive selection rule and asks whether to compare separate models through backtesting or follow common industry practice. No method, empirical results, or answer is included, so the document is best read as a concise statement of a modeling problem rather than a complete trading recipe. Its central lesson is that regression orientation affects the constructed residual, and thus the signal being traded. Any model comparison would need to assess stability and realistic trading performance; the note itself does not specify validation procedures, costs, or criteria for selecting among candidate spreads.
Key ideas
- Changing the dependent asset in an ordinary least squares regression changes the estimated hedge ratios.
- Different hedge ratios produce different spread series and potentially different mean-reversion signals.
- The proposed trade sells a spread above fair value and buys it when it appears cheap.
- The note raises model selection and industry practice questions but provides no answer or backtest evidence.
Tags
Full text
# Determine Dependent Variable Product # Determine Dependent Variable Product Let's say I have three products that are correlated (e.g. AAPL, MSFT, and AMZN). I would like to construct a spread between these products and trade the mean-reverting spread. Specifically, sell the spread when it exceeds the fair value and buy the spread when it becomes cheap. Now, depending on how I develop the hedge ratios between these products, the resulting spread would be different. For example if I use OLS with the returns of AAPL as the dependent variable (DV), the resulting spread would be different than MSFT as the DV. My question is whether there exists an intuitive way to determine the DV or whether I should run three separate models and determine the DV based on the backtested results? What are some common practices in the industry? Any resources to read more about this?
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.