Why Fama–French Portfolio Replications Change Across Data Vintages
Summary
The discussion examines why a replication of the Fama–French five factor study for Europe produced different average excess returns for the same stated period and portfolio sorts. The researcher subtracts the Europe dataset's risk free rate from each of 25 size and book to market portfolio returns, then compares the resulting means and regressions with the paper. A respondent repeats the calculation using the current library files and finds small differences from the published table. Using archived files from January 2017 brings the computed means much closer to the paper's results.
The evidence points to dataset revisions as a plausible explanation: equity histories can be updated for corporate actions and other corrections, and the data provider warns that files may change. The exchange does not establish whether the historical differences came only from data revisions or also from changes in the production methods. It offers a useful replication check, while cautioning that matching current files alone may not reproduce results calculated from an older data vintage.
Key ideas
- A replication should align portfolio definitions, dates, and risk free rate series before comparing means.
- Subtracting the monthly risk free rate from each portfolio return gives the excess return series.
- Archived data produced means much closer to the original paper than the current files did.
- Historical revisions may affect replication, and the discussion cannot isolate data changes from methodology changes.
Tags
Full text
# Mean of Excess Returns: Why Do Fama-French Library and F&F (2017) Differ for the Same Time Period and Portfolios?
# Mean of Excess Returns: Why Do Fama-French Library and F&F (2017) Differ for the Same Time Period and Portfolios?
I'm currently having a hard time finding the solution to my problem. I'm trying to replicate the 5 Factor Model from Fama & French (2017) for Europe. I'm using Data from the French Library and tried to calculate the Means of Excess Returns for the same time period 07.1990 - 12.2015. I also use the same Portfolio sortings. But my results differ from the paper. I calculated the Excess Returns by subtracting the riskfree rate (obtained from the Europe 5 factor Dataset) from the Portfolio Returns.
From what I've read, the Breakpoints for Portfolio construction are the same. Does someone know the answer? Or am I miscalculating something
Edit1:
Thanks for your reply, I'll try to be more precise. All Data I used is for Europe. I used the 5x5 Size, B/M Portfolio Dataset. So I had Returns on 25 Portfolios. With Data ranging from 07.1990 to 12.2015. The time period used in F&F (2017). Then I took the Riskfree rate from the "5 Factor for Europe" Dataset and subtracted it from the portfolio Returns. Month by Month for all 306 Datapoints. Thus I had the Excess Returns for all 25 Size, B/M Portfolios. Then I calculated the Mean of all 25 Excess Returns and compared it to the Paper. I did this in Excel and R. And in both cases there was difference in results of Means. For the means of the 5 factor Returns, I get pretty similar results. There are only very minor differences, which I think result from rounding differences. But I'm not able the calculate the same Excess Returns for the Portfolios. I also calculated the Regression and I only get approximately similar results to the paper. My supervisor told me to keep trying to get closer to the paper. But I really don't know how
Edit 2: Here's my R-Script for calculating the Means of Excess Returns:
> Returns <- read.csv("C:5x5_Size_BM.csv") Factor <- read.csv("C:/Europe_5_Factors.csv")
#Loading 5 Factor Dataset and Value-Weighted Portfolio Returns
> Factors <- Factor[1:306, 2:7] Returns <- Returns[1:306, 2:26]
#Removing datapoints after 12.2015 and first column with monthly dates
> Riskfree_Rate <- Factors[, 6]
#extracting the Riskfree rate data column into a single vector
> Riskfree_rate_Matrix <- matrix(rep(Riskfree_Rate, times = 25), nrow = 306, ncol = 25)
#turning vector into 25x306 matrix with identical datapoints
> Returns_Matrix <- as.matrix(Returns)
> Excess_Returns <- Returns - Riskfree_rate_Matrix
> Means_of_Excess_Returns <- colMeans(Excess_Returns)
> Means_Matrix <- matrix(Means_of_Excess_Returns, nrow = 5, ncol = 5, byrow = TRUE)
> print(Means_Matrix)
```
[,1] [,2] [,3] [,4] [,5]
```
[1,] -0.1181046 0.2912418 0.3668301 0.5107190 0.6726471
[2,] 0.1800000 0.3951307 0.4839216 0.6519281 0.7114706
[3,] 0.3244771 0.5082026 0.4926471 0.5074837 0.6722549
[4,] 0.4579739 0.4795425 0.4864379 0.5061765 0.6013399
[5,] 0.3447712 0.5021569 0.5128431 0.6157190 0.5027451
## Answer by Enrico Schumann (score 2)
https://quant.stackexchange.com/a/81590
How close do you come with your results? Equity data are often revised (corporate actions are adjusted/corrected, ...). And IIRC, even on K. French's website, there is a warning that datasets can change. That being said, I just tried this (in R):
```
library("NMOF") ## https://enricoschumann.net/R/packages/NMOF/index.htm
E5 <- French("~/Downloads/French",
"Europe_5_Factors_CSV.zip")
i <- rownames(E5) >= "1990-07-31" & rownames(E5) <= "2015-12-31"
E5 <- E5[i, ]
E25 <- French("~/Downloads/French",
"Europe_25_Portfolios_ME_BE-ME_CSV.zip")
i <- rownames(E25) >= "1990-07-31" & rownames(E25) <= "2015-12-31"
E25 <- E25[i, ]
E25.xreturns <- E25 - E5[, "RF"]
m.xreturns <- colMeans(E25.xreturns)
100 * matrix(m.xreturns, byrow = TRUE, nrow = 5)
## [,1] [,2] [,3] [,4] [,5]
## [1,] -0.1181046 0.2915359 0.3668301 0.5102614 0.6728431
## [2,] 0.1800000 0.3950654 0.4839216 0.6519281 0.7114706
## [3,] 0.3162092 0.5082026 0.4927124 0.5071242 0.6724837
## [4,] 0.4588235 0.4795425 0.4864379 0.5075817 0.5997059
## [5,] 0.3447386 0.5021569 0.5128431 0.6157190 0.5027451
```
(Note that you'll need at least version 2.10-2 of the NMOF package to run the example.)
Comparing with the table on page 447, I get a maximum absolute difference of about 5 bp, which seems not too far off to me.
Update: I just fetched the two datasets from an archived version of Kenneth French's website, dated January 2017. When I redo the exact steps from above, I get the following matrix:
```
## with csv files as of Jan 2017
## [,1] [,2] [,3] [,4] [,5]
## [1,] -0.1040523 0.2826144 0.4181699 0.4937255 0.6769608
## [2,] 0.2295425 0.3887255 0.4793464 0.6386275 0.7223529
## [3,] 0.2982353 0.5266993 0.5023203 0.5062418 0.6611765
## [4,] 0.4430719 0.4733660 0.4801634 0.5588235 0.6221242
## [5,] 0.3440196 0.4841503 0.5157516 0.6111438 0.4943464
```
Now all absolute differences are below 1 bp. (But of course there is no way of telling whether only the data changed, or the scripts that produced the data as well.)
## Answer by phdstudent (score 1)
https://quant.stackexchange.com/a/81593
It can indeed change over time. Here's a paper about it:
- Akey et al. Noisy Factors? The Retroactive Impact of Methodological Changes on the Fama-French Factors (2021)
And replicating FF factors can be very very difficult even for the U.S. What is your data source for the individual stock data?Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.