Replicating a Portfolio: Regressing Constituents or Aggregate Returns
Summary
The document compares two ways to estimate hedge positions for replicating a signed, equal-weighted portfolio of stocks with other instruments. One approach runs a separate rolling regression for each constituent and averages the resulting hedge coefficients. The alternative first forms the portfolio’s historical returns and regresses that aggregate series directly on the candidate hedging instruments. The author asks whether the direct regression is methodologically valid and which approach better reduces tracking error.
The document presents the modeling choices but includes no answer, derivation, or empirical results, so it does not establish that either procedure is superior. In general, averaging coefficients from separate regressions need not produce the same coefficients as a regression on the aggregate target; the relationship depends on the data and model setup. The proposed comparison therefore calls for careful validation using the intended rolling estimation and evaluation process. The stated portfolio uses signed stock returns, and the candidate hedge instruments and target construction should be aligned consistently when comparing tracking error.
Key ideas
- The first proposed method regresses each constituent return on candidate hedge instruments and averages the coefficients.
- The second method regresses the aggregate signed portfolio return directly on the hedge instruments.
- The two procedures need not yield identical hedge ratios because regression and averaging need not commute.
- The document poses the methodological question but supplies no answer or comparative evidence.
- Tracking error should be assessed with portfolio and hedge returns defined consistently.
Tags
Full text
# Constructing a Replicating Portfolio : Regression on Individual Constituents or their Average?
# Constructing a Replicating Portfolio : Regression on Individual Constituents or their Average?
I would like to replicate a portfolio of stocks $S_1, \cdots, S_n$ using other instruments, $X_1, \cdots, X_m$. Using the letters above with a subscript $t$ to denote the forward returns over some horizon of the portfolio at time $t$, the return of the portfolio (and the variable I am targeting) is:
$$ P_t = \frac{1}{n} \sum_{i=1}^n \epsilon_{it}S_{it}, \quad \epsilon_{it} \in \{-1,1\} $$ e.g. we hold an equal weighted portfolio of each stock, long and short.
To construct the hedge ratios, I considered estimating $n$ regressions (on a rolling basis, backwards in time):
$$ S_{it} = \sum_{j=1}^m \beta_{ij}X_{jt}, \quad 1 \leq i \leq n $$ and getting coefficients $\beta_{ij}$, representing the number of shares of the hedging instrument $X_j$ that need to be bought/sold-short to replicate stock $i$. The final allocations of the portfolio to each $X_j$ would then be: $$ \beta_j : = \frac{1}{n}\sum_{i=1}^n \beta_{ij}, \quad 1 \leq j \leq m $$ This means - replicate each stock $i$ individually, and average the coefficients across stocks to get the final amount to buy/sell-short for the portfolio.
However, I also considered just reconstructing the $P_t$ backwards in time (averaging the signed past returns of the selected stocks), and then directly estimating: $$ P_t = \sum_{j=1}^m\beta_j X_{jt} $$ This is a (marginally) computationally less intensive procedure, that uses a less noisy left hand side variable.
My question is - from a theoretical/methodological point of view, is there anything incorrect with this second method? Or is it just a question of getting my hands dirty with the data and seeing what works best? Of course I'd be interesting in minimizing my tracking error here, and computational considerations aren't really a problem as these are all OLS regressions.Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.