Skip to content
All library documents

Clustered and GMM Standard Errors in Asset-Pricing Panels

Article Quant Q&A · Author: Richard Hardy

Summary

The document examines how to choose standard errors for asset-pricing panel regressions, such as CAPM or multifactor models estimated with monthly data. It frames time-clustered errors as a way to account for cross-sectional dependence and asks how they compare with GMM-based errors, including a simplified GMM case with no autocorrelation. It also points to Fama–MacBeth estimation as an alternative discussed in the cited literature.

The author then questions whether a panel regression written with realized market returns adequately represents a model based on expected market returns. Replacing the expectation with the observed return may create measurement error, while time effects do not straightforwardly resolve the issue because the market shock enters with asset-specific beta loadings. The document raises these questions but supplies no resolution or empirical comparison, so it serves as a statement of modeling and inference concerns rather than a recommendation.

Key ideas

  • Time clustering is presented as a response to cross-sectional correlation in finance panel data.
  • The document asks whether clustered and GMM-based standard errors are related or preferable under specified assumptions.
  • Fama–MacBeth errors are cited as another approach discussed in the referenced literature.
  • Using realized market returns in place of expected returns raises a potential measurement-error concern.
  • Time fixed effects may not resolve the stated issue because market shocks interact with asset-specific betas.

Tags

Full text
# Clustered vs. GMM-based standard errors: which ones to use in asset pricing?


# Clustered vs. GMM-based standard errors: which ones to use in asset pricing?












Consider estimating an asset pricing model such as the CAPM or a multifactor model using monthly data. Petersen (2009) section "Asset pricing application" suggests use of standard errors clustered by time, as this addresses the problem of cross-sectional correlation. (Meanwhile, serial correlation is not much of an issue.) Petersen (2009) discusses how time-clustered standard errors fare against alternative approaches such as Fama-MacBeth (works fine) and several other. However, he does not seem to consider GMM estimation that is discussed extensively in Cochrane "Asset Pricing" (2005) Part II, e.g. Chapter 11.

On the other hand, Cochrane does mention clustered standard errors and Petersen's paper in his video lecture on the topic; see here and again here. However, he does not compare between clustered standard errors and GMM-based ones either.

So how do clustered standard errors and GMM-based standard errors compare? (If needed, we may consider (1) clustering by time and (2) assuming zero autocorrelation in GMM for concreteness.) Is one a special case of the other? Which ones makes more sense (given that 13 years have passed since Petersen's paper, I guess there should be some consensus on the matter)?

Due to lack of answers, this has been reposted on Cross Validated Stack Exchange.

Update (now posted as a separate question) After some thinking, I am not sure if I can even formulate a statistically adequate panel regression model. Take the case of the CAPM written in terms of excess returns (thus the asterisks): $$ R_{i,t}^*=\beta_i\mathbb{E}(R_{m,t}^*)+\varepsilon_{i,t}. $$ We can generalize it to allow for nonzero Jensen's $\alpha$ $$ R_{i,t}^*=\alpha_i+\beta_i\mathbb{E}(R_{m,t}^*)+\varepsilon_{i,t} $$ (and later test $H_0\colon \alpha_1=\dots=\alpha_N=0$), but we have the problem of $\mathbb{E}(R_{m,t}^*)$ not being observable. Just replacing $\mathbb{E}(R_{m,t}^*)$ with $R_{m,t}^*$ would introduce measurement error and bias the estimate of $\beta_i$. Adding time fixed effects would not help with the problem. To see that, express $R_{m,t}^*$ as $R_{m,t}^*=\mathbb{E}(R_{m,t}^*)+\varepsilon_{m,t}$, yielding $$ R_{i,t}^*=\alpha_i+\beta_iR_{m,t}^*+\beta_i\varepsilon_{m,t}^*+\varepsilon_{i,t} $$ where we see that the time fixed effect is convoluted with $\beta_i$. So how does one even start working out standard errors clustered by time if the model itself is inadequate? Or did I miss something?

References

- Cochrane, J. (2005). Asset Pricing: Revised Edition. Princeton University Press.

- Petersen, Mitchell (2009). "Estimating Standard Errors in Finance Panel Data Sets: Comparing Approaches". Review of Financial Studies 22 (1): 435–480. (A freely accessible working paper version can be found here.)

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.