Handling Missing Observations in Covariance Estimation
Summary
The document raises a practical problem in estimating covariance for portfolio construction when asset price histories have different gaps. Missing observations arise from holidays, isolated data omissions, and longer stretches without quotes, leaving different effective sample sizes across funds. The goal is to build a covariance matrix for an efficient frontier without losing too much data.
The author reports that computing each pairwise covariance from its own available dates can produce a matrix that is not positive semidefinite, which can imply negative portfolio variance. The question asks whether to align observations by dropping dates or to fill gaps with synthesized values, but it provides no proposed solution or evidence comparing methods. Its main lesson is the constraint that pairwise estimation can violate matrix consistency; the document does not specify an imputation procedure or establish how to address the problem.
Key ideas
- Pairwise covariance estimates can use different date samples for different asset pairs.
- A matrix assembled from pairwise samples may fail to be positive semidefinite.
- A non-positive-semidefinite covariance matrix can imply impossible negative portfolio variance.
- The document frames dropping incomplete dates and synthesizing missing observations as options but does not evaluate a remedy.
Tags
Full text
# Covariance/correlation matrix from data with missing data points # Covariance/correlation matrix from data with missing data points I have a data set with index fund quotes, and I'm trying to compute the efficient portfolio frontier for it. But some data points are missing. In some cases there are few funds that trading even on holidays, while for the rest there is no data on that day. Sometimes there is a single day of data missing here and there. Some funds have no data for 4-5 days or even a month in a row. For 3 years, I would expect 3 * 252 = 756 days of data. But I have from 624 up to 784. To compute covariance, I need to normalise my data - either to strip out the days and/or funds where the some data is missing, or fill the gaps with some synthesised data. I've tried to strip out days for each pair of funds to minimise loss of data, but I've ended up with matrix which is not Positive semi-definite and I end up with negative values for total variance. How can I normalize my data to be able to compute covariance matrix?
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.