Skip to content
All library documents

Use Returns, Not Price Levels, to Compare Stock Correlations

Article Quant Q&A · Author: A.L. Verminburger

Summary

The document considers whether stock prices should be z-score standardized or log transformed before calculating correlations or clustering distances. It explains that correlation is already invariant to linear shifts and positive rescaling, so z-scoring each series does not change its correlation with another series. This makes z-score scaling unnecessary when correlation itself is the comparison measure.

For comparing stocks, the response recommends calculating correlations from returns rather than raw price levels. Applying logarithms to price levels changes the variables being compared: correlation after that transformation measures linear association between log prices, which is appropriate only when that relationship is the intended one. The document does not establish that log transformation is universally wrong; its usefulness depends on the model and question. It also leaves the broader clustering design, including choice of return definition, sampling window, and distance conversion, unspecified.

Key ideas

  • Correlation is unchanged by linear shifts and positive rescaling of its inputs.
  • Z-scoring price series therefore does not alter their pairwise correlations.
  • Stock correlations are generally calculated from returns rather than price levels.
  • Log-transforming prices changes the relationship being measured and targets linear association between log prices.
  • Choose transformations according to the relationship the clustering analysis is meant to capture.

Tags

Full text
# Answer by msitt (score 0, accepted)


# z-score versus log standardisation of stock prices for calculating correlation; which to use (in ML clustering, distance measure)?












I need to compare (get correlation between) different financial instruments (stocks).

The problem is that different stocks will have different price scales.

I was thinking of using z-score standardization on my price time series vectors $\boldsymbol{x_{j}}$:

$$\boldsymbol{x_{j}'} = \frac{\boldsymbol{x_{j}} - \bar{\boldsymbol{x_{j}}}}{\sigma} $$

Now a paper I read uses natural log standardization to achieve the same goal:

$$\boldsymbol{x_{j}''} = ln(\boldsymbol{x_{j}})$$

Is one approach correct and the other incorrect; are both usable, if so which one is preferred and what are the nuances?

Additional info based on answers and comments:

Let me add some context where this is coming from (more of a statistics / machine learning perspective). I want to do classification of different equity markets. Standardisation is a "standard" part of data pre-processing for forecasting or clustering (this is a clustering problem). And I am guessing if I were to use things like expected return and volatility AND Euclidean distance as my measure, it would make sense. However, I have chosen to use correlation as my distance measure. And this is where the question arises. I do not understand why, statistically I should use returns. I can kind of see how z-score is already incorporated into correlation (rather than covariance), although not 100%, not quite sure about log transformation. Since I am doing correlation I am measuring by default the linear relationship; I thought there would be no difference in the linear relationship between X and Y or ln(X) and ln(Y), it just makes sure the scales are the same. But then again the scales do not matter here since we are "standardizing" in the denominator of the correlation equation. Here is the link to the paper that used ln(price).

## Answer by msitt (score 0, accepted)

https://quant.stackexchange.com/a/33367

First let me say that correlation between two stocks is almost always taken in return space. First you would transform your price series to a return series and take the correlation.

Now let me address your question about correlations in general. Note the formula for correlation: $$ \rho_{XY} = \frac{E[(X-\mu_X)(Y-\mu_Y)]}{\sigma_X\sigma_Y} $$ From this, you can see that correlation is a normalized value that is invariant to scaling/shifting of the inputs. So using your "z-score standardization" method, will actually give you the exact same correlations!

The log standardization will measure the linear relationship between the log transformed variables. You only want to do this if you think there should be a linear relationship between $\textrm{ln}(X)$ and $\textrm{ln}(Y)$.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.