Estimating Covariances When Assets Have Unequal History Lengths
Summary
The document considers rolling portfolio optimization when some assets have shorter histories and therefore lack observations within the full estimation window. It begins from the pairwise structure of a covariance matrix: each asset pair may have a different amount of overlapping return data.
For pairs with enough common observations, it recommends using their empirical covariance. When overlap is absent or too limited, it describes estimating each asset’s exposure to a small set of principal components built from assets observed throughout the period, then inferring their covariance through those exposures. For partial overlap, possible choices include using the shorter sample, using the factor-based estimate, or combining the two. The answer leaves the minimum acceptable sample size, factor count, and blending method to the practitioner; it gives no empirical comparison of these alternatives.
Key ideas
- Covariance estimates are pair-specific, so each asset pair can have a different number of shared observations.
- Use empirical covariance when the pair has enough common return data.
- For pairs with little or no overlap, regress returns on a small set of principal components from assets with complete histories.
- When overlap is partial, choose between the shorter empirical sample, factor-based estimation, or a blend based on confidence in the data.
Tags
Full text
# Portfolio optimization with changing portfolio constituents
# Portfolio optimization with changing portfolio constituents
Say I have time series data for $N$ assets, where for the longest existing asset I have data from $t_0=0$ to $T$, but for several other assets I only have data from say $t_0+k$ to $t_0+l$ for some $0<k<l<T$.
Now, if I want to do portfolio optimization with rolling estimation of the covariance matrix using a window of size $m$, what are some good ways to deal with the assets for which I'm missing data points inside the window?
## Answer by lehalle (score 1, accepted)
https://quant.stackexchange.com/a/40464
Go back to the definition of the covariance matrix: for $N$ stocks this matrix is made of the covariances $C_{i,j}$ of any $i$ and $j$ stocks from $1$ to $N$.
You can face different situations:
- $i$ and $j$ are both present during your reference period of $m$ dates: you have your empirical estimate of $C_{i,j}$
- you never observed $i$ and $j$ the same day during your period... Well this is a problem but if you really need a correlation you can compute a covariance matrix for all the stocks that are there for all the days perform a PCA on it and keep its first $K$ components ($K$ being low) $P_1,\ldots,P_K$ regress the returns of $i$ on $(P_1,\ldots,P_K)$ regress the returns of $j$ on $(P_1,\ldots,P_K)$ compute their covariance thanks to these two regressions (since any $P_{k}$ and $P_{k'}$ are orthogonal, it is not difficult)
- If $i$ and $j$ have few common dates, say $m'<m$: either you decide $m'$ is enough and you use the corresponding empirical covariance either you do not believe $m'$ points is enough: you use the previous approach (as if they had no date in common) alternatively you could mix the estimation on $m'$ points and the interpolation via $(P_1,\ldots,P_K)$Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.