Building a Fama–French-Style Factor Model from Firm Data
Summary
The document outlines the main steps for constructing a factor model from firm characteristics and stock returns, using a Fama–French-style approach. It recommends cleaning the data, forming characteristic-sorted portfolios for factors such as SMB, HML, RMW, and CMA, calculating portfolio returns, and estimating stocks’ exposures to those factors through regressions. It also notes that the study needs stock returns, a risk-free rate, and a local market index in addition to accounting data.
The discussion treats missing observations as a normal consequence of firms entering and leaving the market or reporting data at different times. Its practical suggestion is to omit a stock from a portfolio for periods when the needed data is unavailable. The document does not provide a complete implementation, sorting rules, or detailed guidance on accounting-data timing and survivorship bias. It points to programming knowledge as necessary and leaves model interpretation and specific estimation choices to the researcher.
Key ideas
- A factor replication requires cleaning and aligning firm characteristics with returns over time.
- Characteristic-sorted portfolios are used to construct factor returns such as SMB, HML, RMW, and CMA.
- Estimate stock factor exposures by regressing excess stock returns on factor portfolio returns.
- Stocks with unavailable characteristic data can be excluded from the relevant portfolio period.
- The study also needs stock returns, a risk-free rate, and a local market index.
Tags
Full text
# How to build Factor model like Fama & French (2014)? # How to build Factor model like Fama & French (2014)? I would like to conduct a study where I build a factor model based on the characteristics/variables I have collected about firms, using a couple of countries. I have acquired two Excel data files from CompuStat : - Monthly index prices (MSCI WORLD INDEX) (1950-2017) - Monthly financial statement data (like P / E ratio, B / P ratio) of different firms over the period 1950-2017. The companies all have a company key (GVKEY) as a filter option in Excel. So I think I first need to merge the data based on the GVKEY and date. That would not be any problem, but the main problems are: - Some companies do not have values for one or more variables at certain time points and not all firms have the time period 1950-2017, some start at 1993 for example. Does anyone have any suggestions of how to deal with the problem and merging the data and how I can conduct my study? Does someone have a step-by-step plan or something for me? Could this task be as simple as regressing average returns for a stock with its different factors? I heard that I need to do something with creating risk factors through regressions, making portfolio's and sorting on characteristics and using something like a fama-macbeth regression. The problem is that I do not know the hard-programming to figure out how to construct portfolio's etc.. I'd be grateful for any help, I want to use STATA! ## Answer by Alexandre Oliveira (score 2) https://quant.stackexchange.com/a/33871 Replicating Fama & French 2014 model will require some programming language knowledge, to save time. Major steps includes: - Data cleaning & handling missing data - Creating portfolio compositions (SMB, HML, RMW, CMA) for each period - Computing portfolio returns for each period - Compute stock exposures (betas) to each factor (the regression of stock's excess returns against portfolio excess returns) - Interpret results You will also need to have stock returns datasets, the risk-free rate and local equity market index for the target market. The fundamental data will be used to create the portfolios. Missing data will always exists as companies enters and leaves the market all the time, but it may happens also due to delays in financial statement releases and other issues. if data is not available for a good reason, just let the stock out of the spread portfolio for that period.
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.