CRSP Data and the Pre-1962 Fama–French Selection Bias
Summary
The post asks whether the historical stock data behind Kenneth French’s factor library avoids the pre-1962 selection bias that Fama and French identified in earlier Compustat data. The concern is that the older Compustat sample overrepresented large firms that had survived and succeeded, which could distort historical results.
The answer says French’s earlier-period data comes from CRSP rather than Compustat. It describes CRSP as assembling historical prices from multiple sources, including newspaper records for 1926–1933 and dividend records for 1934–1961, and cites CRSP’s quality-control practices. On that basis, the answer concludes the data is free of the stated Compustat selection problem and can be used back to 1926. The post offers this as a source-based explanation, not an independent audit: it does not examine other possible biases, survivorship effects, or the details of how the factor series were constructed.
Key ideas
- The cited Fama–French studies avoided pre-1962 Compustat data because of a serious selection bias.
- The answer attributes the earlier French-library history to CRSP records rather than Compustat.
- The described source records include newspaper and dividend data from different historical periods.
- The answer’s conclusion addresses the stated Compustat bias but does not evaluate other potential biases.
Tags
Full text
# Answer by JorgeT (score 3) # Is the Fama-French website data free of the serious selection bias pre-1962 where it's tilted toward big historically successful firms? Fama and French use data starting in 1963 in both 1) "Common risk factors in the returns on stocks and bonds" (1993) and 2) "The cross-section of expected stock returns" (1992) and mention in (2) that "the COMPUSTAT data for earlier years have a serious selection bias; the pre-1962 data are tilted toward big historically successful firms". (link) In French's website where he posts data, his data goes back to 1926. Is this data better and doesn't have this selection bias or does it still have it? http://mba.tuck.dartmouth.edu/pages/faculty/ken.french/data_library.html ## Answer by JorgeT (score 3) https://quant.stackexchange.com/a/37174 The data that is now available in Kenneth Frenchs' website is now free of that problem so we can use all the available data which starts in 1926. The data comes from Chicago Booths' Center for Research in Security Prices (CRSP) which uses multiple sources of data to create the most accurate database for stock prices. In a report titled "Data descriptions guide" updated in march 2017 the CRSP highlights that "CRSP stock files are designed for research and educational use and have proven to be highly accurate. Considerable resources are expended in ongoing efforts to check and improve data quality both historically and in the current update". Specifically, their data prior to 1962 does not come from COMPUSTAT like in Fama and French's paper. Instead, it comes the New York Times newspaper (1926-1933), Moody's Quarterly Dividend Record (1934-1961).
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.