Skip to content
All library documents

Constructing Fama–French Size and Value Portfolios and Factors

Article Quant Q&A · Author: Zabbka osckey

Summary

The question outlines a procedure for constructing the size and value portfolios used to form the Fama–French three-factor model from monthly stock data. At each period, stocks are split into small and big groups using market capitalization, then each size group is divided into high, middle, and low book-to-market groups. This creates six portfolios whose returns can be used to calculate size and value factor series, alongside the market excess return, for a regression estimating factor loadings and an intercept.

The answers advise automating the repeated calculations with statistical software such as R rather than doing them manually, and point to an existing implementation as a place to study the factor construction steps. They do not provide code or resolve methodological details in the proposed procedure. In particular, the question’s use of a cross-sectional mean size cutoff and its description of portfolio returns as sums warrant scrutiny; standard implementations require carefully specified breakpoints, return weighting, timing, and handling of missing data. The material is a starting point, not a complete recipe.

Key ideas

  • Sort stocks by size and book-to-market to form six size-value portfolios at each period.
  • Use portfolio returns to construct SMB and HML factor return series.
  • Regress stock or portfolio excess returns on market, size, and value factors to estimate exposures.
  • Automate repeated portfolio calculations with a scriptable statistical tool.
  • Specify breakpoints, weighting, timing, and missing-data rules before treating the procedure as a standard factor replication.

Tags

Full text
# Could someone teach me how to construct the portfolios by compute (like using R, Excel or Eviews)


# Could someone teach me how to construct the portfolios by compute (like using R, Excel or Eviews)












Recently, I am doing my dissertation that covers asset pricing theory. The empirical test of Fama 3 factors model is an important part of this dissertation. Please let me review the fama model.

Fama 3 factors model is $r-R_f=α+β_m(K_m−R_f)+\beta_s⋅SMB+\beta_v⋅HML+e$

where $R_f$ is risk free return, ($K_m−R_f$) is premium return and $K_m$ is market return, SMB is the "Small Minus Big" market capitalization risk factor. HML is is the "High Minus Low" value premium risk factor.

Please let mt detail my question. I have over 500 stocks and their monthly return ,monthly market value and monthly book-to-market ratio (call it b/m). The time interval is from 01.2010 to 12.2015, then because the data is monthly so I have 60 periods. Those data is in one excel document.

its methodology and what I want to do is:

- For all stocks at each period, compute the mean for market value.

- Divide them into two groups (say big and small size) by comparing each stock with mean. If stock's market value greater than mean, put them into big size group. If small than mean, put them into small size group. Thus, for each period we have two groups, big and small group in terms of market value.

- For each period, we compare all the stocks' b/m within each group, then furtherly divide them into 3 groups , named high b/m (top 30%) group, medium b/m (middle hierarchy from 30% to 70%) group and low b/m (bottom 30%) group. Namely, for each period, we finally divide 6 groups. They are S/L, S/M, S/H and B/L, B/M, B/H. For example, S/L group contains all stocks that are small market value and low b/m ratio simultaneously.

These 6 groups are just 6 portfolio. So for each period, we have 6 different portfolios.

So far this is my first stage. Then I have second stage:

- After we get 6 different portfolios at each period, we need to compute weighted average return for all stocks at each period.

- We need to compute the SMB and HML for each period.

$$SMB=\frac{\left[\left(\frac{S}{L}+\frac{S}{M}+\frac{S}{H}\right)-\left(\frac{B}{L}+\frac{B}{M}+\frac{B}{H}\right)\right]}{3}$$ $$HML=\frac{\left[\left(\frac{S}{H}+\frac{B}{H}\right)-\left(\frac{S}{L}+\frac{B}{L}\right)\right]}{2}$$.

Where S/L,S/M....B/L is the summation return of each group. Like S/L is the summation of each stock's return.

Finally, we got time series data, then can regress this series data, and figure out α and 3 $\beta$.

So, basically those are want I want to do.

Now I compute those manually although I use excel, this is really a huge amount of workload, so tired to do this. I am still in step 1 at first stage. Is anyone interesting in telling me how to achieve dividing stocks into portfolios by compute, and compute those data by compute?

People usually just download the data from French's website. But my goal is to check Fama model in Asian stock market, like Shanghai and Hong Kong. All answers welcome, really really crazy and tired to do this manually.

Anyway, thanks everyone in advanced.

## Answer by simmy (score 5)

https://quant.stackexchange.com/a/26303

First of all, it is not conceivable to do all that work by hand! You are crazy to have just thought it!

Second, if you want to repeat your work with different datasets, I suggest you to use R, since, once you have written a script, you can use it all the times you want. But, there's a 'but': you cannot think we are going to write some code for you (you should write our names on your dissertation as co-authors). You can ask here about theoretical problems you encounter, or you can post the section of your code that is not working and we would be happy to help you. In every kind of research you will have to conduct, you are going to analyze data, and the best way to do this is using softwares like R (I'm telling you that, sooner or later, you will have to learn to code some (even small) scripts). Once you begin, you will use it for everithing.

Given this, you can find plenty of YouTube tutorials, pdf guides, examples of scripts and communities (such as stack exchange ones) on the web ready to help you on your path to the dark side of statistical and econometrical analyses with any software you want (R, gretl, Stata, Eviews, or even more programming oriented as MATLAB or Python).

Your analysis seems quite simple (in the sense that you do not need strange packages or functions to compute your calculations) and you will discover by yourself how useful, powerful and not that difficult to learn R is.

## Answer by Cyurmt (score 3)

https://quant.stackexchange.com/a/33051

I've posted on my website R code that replicates the Fama-French factors (plus momentum) from scratch. It assumes you have access to WRDS but if you have your own data, you can begin using the code where ever you see fit. In your case, you'd want to start in the Construct Fama-French Factors section of my Main_Fama_French file and also look at the Form_CharSizePorts2 function in the Support_Functions file.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.