Skip to content
All library documents

Market-Capitalization Weights in Cross-Sectional Weighted Regression

Article Quant Q&A · Author: mattaj247

Summary

The document explains how to apply market-capitalization weights when regressing stock returns on factor scores. In weighted least squares, each observation's squared residual is multiplied by its weight in the objective function. If the desired observation weight is the square root of market capitalization, use that square-root value as the regression weight; the fact that residuals are squared does not mean the weight itself should be square-rooted again.

The answer illustrates the distinction with example implementations in Python and R. It does not present empirical results or discuss why square-root market capitalization is an appropriate weighting scheme for a particular factor model. The key practical caveat is to confirm how a chosen software package defines its weights: WLS weights multiply squared residuals, while some routines may accept differently interpreted quantities, such as inverse variances.

Key ideas

  • Weighted least squares minimizes a sum of squared residuals multiplied by observation weights.
  • A target weight of square-root market capitalization should be passed as that weight, rather than converted to a fourth root.
  • Ordinary least squares is the special case where every observation has equal weight.
  • Check a regression package's weight convention before supplying market-cap-based values.

Tags

Full text
# Weighting stocks by market capitalization in a cross-sectional weighted regression


# Weighting stocks by market capitalization in a cross-sectional weighted regression












I am trying to regress stock returns on a series of factor scores to get factor coefficients. I want to weight the regression by the square root of market cap which I'm doing by applying a weighting function to my x and y variables before running the regression.

I'm a bit confused, however, but whether I should be using square root or fourth root as in the end I'm trying to minimize the square of the errors. Does this mean I need to use fourth root or have I over-complicated it?

Thanks!

## Answer by skoestlmeier (score 1)

https://quant.stackexchange.com/a/51200

Let $\text{SSR}$ denote the sum of squared residuals and $\text{WSSR}$ the weighted $\text{SSR}$. Standard OLS-regression approach minimizes the $\text{SSR}$ with $y$ as the dependent and $y$ as the independent variable:

$$\text{SSR}(\beta) = \sum^n_{i=1}{(y_i - \hat{x}_i \cdot \beta)^2}$$

The WLS-approach adds a weight $w$ for each of the observations $x_i$. OLS-regression is the special case of WLS when applying $w=1$ for all $x_i$. WLS minimizes the weighted $\text{SSR}$:

$$\text{SSR}(\beta, w) = \sum^n_{i=1}{w \cdot (y_i - \hat{x}_i \cdot \beta)^2}$$

If your weights $w_i$ are the squared root of firms $i$ market cap, this results in:

$$\text{SSR}(\beta, w) = \sum^n_{i=1}{\sqrt{MV_i} \cdot (y_i - \hat{x}_i \cdot \beta)^2}$$

where $MV_i$ is the market capitalization of firm $i$. As a result, the weight $w$ still remains the "squared" root and not the "fourth".

In Python, WLS with weights from one to seven is applied as:

```
import statsmodels.api as sm
Y = [1,3,4,5,2,3,4]
X = range(1,8)
X = sm.add_constant(X)
wls_model = sm.WLS(Y,X, weights=list(range(1,8)))
results = wls_model.fit()
results.params
array([ 2.91666667,  0.0952381 ])
```

Whereas in R, you just run:

```
y <- c(1,3,4,5,2,3,4)
x <- 1:7
summary(lm(y ~ x , weights = 1:7))
```

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.