Skip to content
All library documents

Avoiding Look-Ahead Bias in CRSP Value-Weighted Returns

Article Quant Q&A · Author: Max Rolfes

Summary

The document addresses why a researcher’s value-weighted CRSP returns may exceed a published market return. Its main diagnosis is a timing error in portfolio weights: monthly returns from the prior month-end to the current month-end must be weighted using information available at the prior month-end. Using current month-end share prices and shares outstanding to construct weights introduces look-ahead bias and can inflate measured performance.

The answer also recommends checking whether the stock universe matches the Fama-French market return definition. That definition screens for U.S.-incorporated firms listed on major U.S. exchanges, eligible share codes, valid beginning-of-month price and share data, and usable returns relative to the Treasury bill rate; it also excludes some microcaps. The response gives a conceptual timing rule and universe criteria, but no debugging steps or reproduction of the researcher’s calculations. Delisting returns are mentioned in the question, not identified by the answer as the source of the discrepancy.

Key ideas

  • Weights for a monthly return should use market capitalization measured at the start of the return period.
  • Using end-of-period prices and shares to weight that period’s returns introduces look-ahead bias.
  • A replication should match the reference portfolio’s exchange, incorporation, share-code, and data-quality screens.
  • Some microcaps may be excluded when matching the Fama-French market return universe.
  • The answer offers likely causes but does not verify which one explains the reported difference.

Tags

Full text
# What are necessary adjustments to returns in CRSP?


# What are necessary adjustments to returns in CRSP?












I guess this is a pretty straight forward and basic question. I am using the entire CRSP universe from 1962-2016 and my goal is to replicate a research paper. However, I realized that the average (value weighted) return of my downloaded CRPS data is too high in comparison to e.g. the original paper, or market return data from the Kenneth R. French website. The difference is significant ~ 3% higher, thus I am a bit clueless as to what I did wrong.

What I did: I used holding period return data for the entire data base and took the PERMNO to identify single stocks. Moreover I only select those issues with a share code of 10 or 11 and I have added the delist return to the appropriate date. I never worked with CRSP and I was under the impression that I accounted for everything that should be accounted for.

My question therefore is: Is there anything I missed are there any adjustments that should have been done to the data to derive the appropriate results.

I am grateful for any advise and if the conclusion is that I must have messed up, well I would be grateful for that inside too.

## Answer by Matthew Gunn (score 3)

https://quant.stackexchange.com/a/36896

I'm going to guess that you might be getting the timing mismatched when computing value weights. (When I was a TA for a first year finance PhD class, I was surprised at how common this error was.)

- Let $s_{it}$ be the share price of firm $i$ at the end of month $t$.

- Let $n_{it}$ be the number of shares outstanding of firm $i$ at the end of month $t$.

- Let $r_{it}$ be the return of firm $i$ from the end of month $t-1$ to the end of month $t$.

If you have portfolio weights $w_{it}$ such that $\sum_i w_{it} = 1$, the portfolio return is:

$$ r^{(p)}_t = \sum_{i} w_{it} r_{it} $$

If you choose weights $w_{it}$ proportional to $s_{it} n_{it}$, you will be using information from the future to choose portfolio weights! This is an epic no no that will lead to unachievable, high returns. For the return from $t-1$ to $t$, you only have access to information available at time $t-1$. The value weight $w_{it}$ for month $t$ should be proportional to the market cap $s_{i, t-1} n_{i, t-1}$. For example, the value weight return for February should use the market caps as of the end of January.

#### Other possibilities (if you're doing the weights properly)

Another possibility is that you aren't matching the universe of stocks properly.

Make sure you're matching the universe of Fama-French market return, "... all CRSP firms incorporated in the US and listed on the NYSE, AMEX, or NASDAQ that have a CRSP share code of 10 or 11 at the beginning of month t, good shares and price data at the beginning of t, and good return data for t minus the one-month Treasury bill rate (from Ibbotson Associates)." Basically you want to keep only share codes 10 and 11 and then exclude some microcaps (this last part is less quantitatively important in modern data).

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.