Handling Missing Returns in Covariance-Based Portfolio Optimization
Summary
The discussion explains why missing entries in a return covariance matrix can arise when asset histories have no overlapping observations. Pairwise covariance estimates can use the dates available for each asset pair, but if two series never overlap, their covariance cannot be estimated from the data. The example of securities from entirely different historical periods illustrates why that missing value reflects an ill-posed comparison, not a value that should be filled in mechanically.
For assets with some shared observations, one proposed approach is to estimate each covariance from its available paired returns, then adjust the resulting matrix to make it positive definite. The discussion cautions that pairwise estimates may use different sample sizes and may not form a valid covariance matrix. Large estimates can also be noisy; shrinkage is suggested as a further step. For portfolio weights at a given date, the asset universe should reflect securities actually investable then, and a time-specific covariance estimate is more appropriate than one unconditional matrix.
Key ideas
- A covariance cannot be estimated from returns when two assets have no dates in common.
- Replacing an unavailable covariance with zero falsely assumes the assets have no co-movement.
- Pairwise estimates based on different observations may not produce a positive definite matrix.
- A nearest positive definite adjustment can make an estimated matrix usable, but does not eliminate estimation noise.
- Portfolio optimization should use assets investable at the decision date and a covariance estimate suited to that date.
Tags
Full text
# Variance Matrix with 'nan' values
# Variance Matrix with 'nan' values
I am trying to optimize a simple portfolio using several random weights and choosing the best. When the number of assets is large I get a covariance matrix with 'nan' values because some asset pairs do not have trading days in common.
How should I treat the 'nan' values?
## Answer by Matthew Gunn (score 1, accepted)
https://quant.stackexchange.com/a/34185
- Your estimated covariance matrix includes `nan` entries.
- The current Pandas.cov function already makes a best effort to estimate covariance based upon available data by ignoring nan/null values.
This implies that to obtain a `nan` in the estimate of covariance, you must have at least two return series that have ZERO time periods in common!
### Your question is ill posed
- What's the correlation between returns of the Dutch East India Company (1602-1799) and Google (2004 - now)? It's an unanswerable and non-sensical question.
- And if your portfolio optimizer says to put $\frac{1}{2}$ your portfolio in Dutch East India Company and $\frac{1}{2}$ in Google, how are you going to do that?
### A direction to move in
If you're going to work with securities that enter and leave your sample, you need to do something more sophisticated than estimate some unconditional covariance matrix with `sigma = mydata.cov()` and using that to choose portfolio weights.
- If the point is come up with portfolio weights for time $t$, it doesn't make sense to include securities which one cannot invest in at time $t$!
- You need some notion of $\Sigma_t$, an estimate that's designed for time $t$.
And replacing `nan` with 0 is not a sensible thing to do! The average covariance term is not zero. Systematic aggregate risk exists and this manifests itself in greater than zero covariance terms.
## Answer by nbbo2 (score 4)
https://quant.stackexchange.com/a/34184
This is a common problem in covariance matrix estimation, with several possible solutions. One of the simplest involves two steps:
(1) You compute each element of the covariance matrix on a 'best efforts' basis, meaning you take the covariance of the two time series involved after REMOVING any data pairs having a N/A value. (Note that this means each element of the matrix will be based on a different number of observations, which means the resulting matrix is not a standard covariance matrix, it may not be positive definite for example). I assume that for any two time series there are at least a few common observations, otherwise it is an ill-posed problem as Matthew Gunn pointed out.
(2) You "massage" the resulting matrix to make it positive definite (and thus acceptable for use as a covariance matrix) using the routine nearPD which is available in R link
[Even after all this work, a large covariance matrix will be very 'noisy' and of poor quality. You should consider further steps such as 'shrinkage' link before you use the results to find an optimum portfolio].Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.