Skip to content
All library documents

Using Normalized Price Differences to Select Pairs

Article Quant Q&A · Author: Travis Liew

Summary

The document asks how to calculate the sum of squared deviations (SSD) used in a pairs-selection approach attributed to Gatev and coauthors. It contrasts two interpretations: summing squared deviations of each stock’s observations from its own mean, and summing squared differences between two price series after each is normalized by its starting value. The latter compares the paths of the two stocks directly and gives smaller values to pairs whose normalized prices move more closely together.

The questioner reports choosing the normalized-series comparison based on a cited monograph. The discussion does not provide a worked calculation, empirical results, or a detailed treatment of implementation choices. It also acknowledges that there may be no uniquely correct definition without specifying the intended methodology. The measure is therefore best understood as a particular distance-based pairing criterion, not a general statistical definition of sum of squares or proof that a selected pair will remain related.

Key ideas

  • Normalize each price series by its value at the start of the identification period before comparing paths.
  • The pair distance is calculated by summing squared differences between the two normalized series.
  • A smaller distance indicates that the normalized price paths tracked each other more closely during the measurement period.
  • This distance criterion is method-specific and does not guarantee a persistent relationship.

Tags

Full text
# Calculating the Sum of Squared Deviations between two Normalized Price Series


# Calculating the Sum of Squared Deviations between two Normalized Price Series












How can I calculate the sum of square deviations between two normalized price series according to (Gatev et. co 2006)? My normalized price series of stocks $X$ and $Y$ consist of the cumulative total daily (log) returns of adjusted close series for stock data taken from Yahoo Finance.

According to some research on the internet it seems I am to:

> Divide the time series of the share prices by the price on the first day of the pairs identification period and then subtract the two new time series from each other, square these differences and sum the result. The smaller the value, the more likely the two stocks are to be a good pair.

If that were the case, I would end up with the formula:

$$SSD=\sum(\frac{x_1,x_2,..,x_n}{x_1}-\frac{y_1,y_2,..,y_n}{y_1})^2$$ Where $X=(x_1,x_2,..x_n)$ and $Y=(y_1,y_2,..y_n)$ are stock price series.

But according to other research, sum of squared deviations refers simply to sum of squares, so then the formula is: $$SS = \sum((x_1,x_2,..,x_n)-\frac{x_1+x_2+..+x_n}{n})^2,or\sum(X-\bar{X})^2$$ But how would I incorporate $Y$ into the second one? Subtract the two results so $SS(X)-SS(Y)$?

If anyone could shed some light on the issue, that'd be fantastic!

Thanks.

## Answer by Travis Liew (score 0, accepted)

https://quant.stackexchange.com/a/10062

Choosing to mark this question "Answered" as @user2763361 has pointed out that there is no "correct" answer. I've chosen to go with the first equation according to a UBS Monograph. Thanks.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.