Skip to content
All library documents

Choosing a Stock Benchmark with Returns, Regression, and Correlation

Article Quant Q&A · Author: RiffRaffCat

Summary

The discussion considers how to choose between two indices as a benchmark for a stock. It cautions against regressing price levels because financial price series may be nonstationary, making a levels regression vulnerable to spurious results. Instead, compare stock returns with index returns, using simple or logarithmic returns, and assess the relationship through regression fit or correlation.

For a single stock and a short sample, a purely statistical ranking may be unstable. Familiarity with the company and its likely market exposure can inform benchmark choice. When selecting benchmarks systematically across many securities, regression or correlation offers a simple starting point, but strong historical fit may reflect chance, a particular market regime, or stock-specific moves. Regression also provides estimated beta and other relationship details, while neither measure alone establishes that an index is an economically appropriate benchmark.

Key ideas

  • Compare stock and index returns rather than their price levels to reduce the risk of spurious regression.
  • Correlation and regression fit provide closely related measures of co-movement.
  • Regression can add information such as the stock’s estimated sensitivity to an index.
  • Short samples and individual-stock effects can make a statistically selected benchmark unreliable.
  • Large-scale automated selection is feasible as a starting point but requires judgment about misleading relationships.

Tags

Full text
# Testing which index is a better benchmark to track stock prices


# Testing which index is a better benchmark to track stock prices












Let's say a Hedge Fund is tracking a stock price. Now the fund has three columns of data, Stock price, Index 1, and Index 2. All of these have data from 2016/01/01 - 2017/01/01. If the fund is to decide which Index it is going to use as the main benchmark for the price of the stock, then which is a better way to make such a decision?

- Simple Regression for P = beta0 + beta1 * Index1 + error and P = beta0 + beta1 * Index2 + error

However, I feel this approach is wrong, because the data we are using is time series data, and this violates many assumptions of OLS model, so we should not use this, am I understanding this correctly?

- Time series regression for the above model.

Could we run this as a time series model?

- Correlation Coefficient between (P, Index1), and (P, Index2).

Could you please inform me which is better and the reasons behind it?

Thank you very much!

## Answer by nbbo2 (score 2, accepted)

https://quant.stackexchange.com/a/42926

To avoid "violating the assumptions of the OLS model" it is important to do the regression using returns $\frac{P_t-P_{t-1}}{P_{t-1}}$ (or logarithmic returns $\ln P_t - \ln P_{t-1}$) for both the index and the stock and not the price level $P_t$. A regression in levels is statistically invalid (so called unit root problem).

Whether you look at the $R^2$ from the regression or calculate the correlation $\rho$ directly makes little difference, the concept is the same.

## Answer by mperlow (score 1)

https://quant.stackexchange.com/a/42925

I have seen this done mostly qualitatively. For example, this is a large cap tech name, so I will benchmark it against the QQQ.

With short time periods of data and a single stock, running a historical regression and choosing a benchmark without taking into consideration your intuition could lead to sporadic results.

Now, if you are doing this wholesale across thousands of securities and trying to algorithmically pick the best benchmark, then running a regression and looking at the R^2 OR running a correlation is a very simple solution to the problem, with full understanding that you could land on outlandish relationships because of market regime, idiosyncratic performance masking itself as some unrelated systematic performance, or dumb luck. I would likely start with the OLS regression as it provides a bit more information as to the relationship of the stock vs. index.

I don't think this violates the OLS model - it is exactly how CAPM calculates stock betas / alphas / residuals.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.