Skip to content
All library documents

Computing R-Squared for Fama–MacBeth Regressions

Article Quant Q&A · Author: NYC07

Summary

The document explains that Fama–MacBeth estimation does not require a special way to aggregate period-specific R-squared values. After estimating coefficients with the procedure, use the chosen coefficient estimate and the relevant observations to calculate fitted values, total and explained sums of squares, and R-squared by the usual regression definition.

It frames Fama–MacBeth as a way to estimate coefficients and conduct inference when errors are correlated across assets within a period but independent across periods. The method runs cross-sectional regressions over time and combines the resulting coefficient estimates; time-clustered standard errors are presented as an alternative that can yield similar inference with different weighting. The document does not fully resolve which observations should define the R-squared in every application, so the researcher must specify the sample and fitted-value construction consistently with the model being evaluated.

Key ideas

  • Fama–MacBeth estimation provides coefficient estimates that can be used to form fitted values for an ordinary R-squared calculation.
  • Calculate R-squared from actual outcomes and fitted values rather than averaging period-specific R-squared values by default.
  • The procedure addresses cross-sectional error correlation by estimating regressions separately for each period and analyzing the coefficient series over time.
  • Time-clustered standard errors offer an alternative approach, though weighting can differ.

Tags

Full text
# How to compute a Fama-Macbeth R-Squared (R2)?


# How to compute a Fama-Macbeth R-Squared (R2)?












I'm reaching out regarding the R-Squared of a Fama-Macbeth regression. This is often reported in econometric results but I have yet to find a good explanation of how it is computed.

Specifically, if I consider the second stage of a Fama-Macbeth regression, where we are potentially running hundreds of regressions, how are the R-Squareds of these hundreds of regressions aggregated into a final R-Squared for the entire procedure? I understand that the coefficients are aggregated by a simple averaging, but was unclear about the R-Squareds.

I understand that there are codes to do this, but am trying to understand what's under the hood.

Thanks!

EDIT:

From the Fama-Macbeth regression we specify a model where each return $y_{i,t}$ of portfolio $i$ in time period $t$ can be priced by: $$ y_{i,t}=\gamma_0 + \gamma_1 \beta_{1, i}+ \gamma_2 \beta_{2, i} + \dots + \gamma_j \beta_{N,i} $$

where $N$ represents the total number of factors.

When calculating the predicted values to calculate our $R^2$, do we take the residuals of each time period or the mean portfolio returns, i.e. are our residuals $y_{i,t}-\hat{y}_{i,t}$ ($i \times t$ number of resids) or only $y_{i}-\hat{y}_{i}$ (i number of resids).

## Answer by Matthew Gunn (score 3)

https://quant.stackexchange.com/a/36981

There's nothing different here. To compute $R^2$, you need the actual values $y_i$ and the fitted (i.e. model predicted) values $\hat{y}_i$. Think of the Fama-Macbeth procedure as just another way to get fitted values $\hat{y}_i$.

Once you have your coefficient estimate $\hat{\mathbf{b}}$ from running Fama-Macbeth. Calculate $R^2$ the usual way: calculate the total sum of squares, obtain the fitted values $\hat{y}_i = \mathbf{x}_i \cdot \hat{\mathbf{b}}$, calculate the explained sum of squares, and then compute $R^2$.

#### Quick econometrics review

Imagine you have the following panel regression.

$$ y_{it} = \mathbf{x}_{it} \cdot \mathbf{b} + \epsilon_{it} $$

Now let's imagine that the error terms are cross-sectionally correlated (i.e. $\operatorname{E}[\epsilon_{it}\epsilon_{jt}] \neq 0$) but across time, the error-terms are independent (and we have $\operatorname{E}[\epsilon_{it_1}\epsilon_{j t_2}] = 0$ for $t_1 \neq t_2$). Because of the cross-sectional correlation, the typical OLS standard errors are going to be understated. In typical finance settings, they will be massively understated because cross-sectional correlation is big.

What to do?

- Option 1: Compute clustered standard errors, clustering on the time variable. (This is arguably a more modern approach.)

- Option 2: Fama-Macbeth procedure

(Note: in typical situations, you should obtain similar results but the two approaches involve different weighting.)

Back in 1973, cluster robust, Rogers standard errors weren't around yet, and instead Fama and Macbeth developed their immensely intuitive procedure. The basic intuition is that:

- Each time period is independent, so we can use our regular Stats 1 techniques to estimate the mean and t-stat for a stationary time series.

- We can run period by period cross-sectional regressions to obtain a time-series of estimates $\{\hat{\mathbf{b}}_t\}$

Combine (1) and (2) and you have the Fama-Macbeth procedure.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.