Skip to content
All library documents

Estimating the Principal Portfolios Prediction Matrix at Scale

Article Quant Q&A · Author: SL133

Summary

The document raises an estimation concern about the prediction matrix in Bryan Kelly’s Principal Portfolios framework. The matrix links each asset’s next-period return to signals associated with other assets, and the cited estimator averages return–signal cross-products over a rolling historical window. The question is whether this approach can be used with a large asset universe.

The concern is dimensionality: the matrix contains a separate estimate for every asset and signal pair, while the cited window is short, creating a potential overfitting problem when the number of assets is large. The author also asks about the Fama–French portfolios used in the paper. No answer, alternative estimator, validation, or empirical result is included, so the document identifies a research issue rather than resolving it. It highlights the need to consider dimensionality and available observations when estimating cross-asset predictive relationships.

Key ideas

  • The prediction matrix represents how signals associated with assets relate to future returns across assets.
  • The cited estimator averages return–signal cross-products over a rolling window.
  • Estimating many matrix entries from relatively few observations can create overfitting concerns.
  • The document asks whether the method scales to large universes but provides no solution or results.
  • It also raises a question about the role of Fama–French portfolios in the referenced study.

Tags

Full text
# Principal Portfolios Prediction Matrix estimation (Bryan Kelly)


# Principal Portfolios Prediction Matrix estimation (Bryan Kelly)












I have recently discovered Bryan Kelly's paper on Principal Portfolios (https://papers.ssrn.com/sol3/papers.cfm?abstract_id=3623983) and had some doubts about the prediction matrix $\Pi$. He defines $\Pi = \mathbb{E}[R_{i,t+1}S_{j,t}]$ as a matrix in which element $\Pi_{ij}$ shows how the return of asset $i$ is predicted by the signal of asset $j$. Kelly writes $\hat \Pi_t = \frac{1}{120} \sum_{\tau = t-120}^{t-1} R_{\tau+1} S_\tau$ as an estimator of $\Pi$, however he is essentially estimating $N^2$ values with 120 observations, which leads to overfitting if used with large stock universes. Kelly uses Fama French Portfolios, which I'm not fully sure what they are either, but my question is, has anyone successfully tried the theory of this paper on a large number $N$ of assets, or has some insight/advice regarding the estimation of $\Pi$? Thank you in advance !

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.