Estimating Barra Factor Returns from Stock Returns and Exposures
Summary
The document explains how to recover factor returns from observed individual stock returns and a Barra factor exposure matrix. The suggested approach is a cross-sectional linear regression: stock returns are modeled as factor exposures multiplied by factor returns, plus a residual specific return. The fitted coefficients are the factor returns for that observation date.
A simplified linear algebra solution is given for the case where the exposure matrix has sufficient rank and the model has no additional risk sources. Repeating the regression across dates produces factor return time series. The factor covariance matrix is then estimated in a later step, using a half-life to weight observations. The discussion distinguishes factor exposures from factor covariance: the covariance matrix is not inserted into the regression equation to obtain factor returns. The account is conceptual and does not specify practical choices such as regression weighting, constraints, or treatment of missing data.
Key ideas
- Factor returns are estimated by regressing stock returns on the stocks’ factor exposures.
- The fitted regression coefficients represent factor returns for a given observation date.
- Residual returns capture stock-specific variation not explained by the factors.
- Repeated cross-sectional fits produce time series that can be used to estimate factor covariance.
Tags
Full text
# How to calculate factor return given a barra model and individual stock returns?
# How to calculate factor return given a barra model and individual stock returns?
Assuming that we have k factors and n stocks
Then Barra provides a k x k factor covariance matrix F, and a k x n factor exposure matrix E.
At runtime, we also observed a n x 1 vector r of individual stock return
My question is, how should we calculate the k x 1 vector of factor return f?
```
r = E^t * f
```
or
```
r = E^t * F * f
```
basically, we need to solve a regression problem to get f . But I am not sure which of these 2 formula should I use.
IIRC it should be the second form, but I am not so sure.
Can someone please help to clarify?
Thanks
## Answer by Nucular (score 2, accepted)
https://quant.stackexchange.com/a/80130
Fit a linear regression model where `y = r` and `X = E'`. As you said,
```
r = E^t * f
```
The coefficients of the linear model will be `f`.
## Answer by Kermittfrog (score 3)
https://quant.stackexchange.com/a/79976
IMHO, without additional risk sources, this looks like an exercise in linear algebra.
Given observed $r$ of dimension $n\times 1$, known $E$ of dimension $k \times n $ and unknown $f$ of dimension $k \times 1$, where $k\leq n$, we find:
$$ \begin{align} r&=E^Tf\\ \Rightarrow\left(EE^T\right)^{-1}Er&=f \end{align} $$
## Answer by Ethantr (score 1)
https://quant.stackexchange.com/a/83743
Barra gives the exposure matrix E $\in M_{k, n}(\mathbb{R})$.
To construct the vector of factor return f $\in M_{k, 1}(\mathbb{R})$ at each time of observation, they solve the regression of equation $r = E^t f + u$ with $u$ being the specific part of the returns, not explained by the factors (which constitute the systematic part).
Thanks to this step and repeating it at each date for which they want to get the vector of factors returns, they end up with $k$ time series. And after that comes the part where they construct the covariance matrix and specific risk matrix using amongst other steps, their parameter "half-life" to weight the time observations.Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.