Estimating Covariance with Unequal and Missing Stock Return Histories
Summary
The document discusses covariance estimation when equity return histories start at different times or contain gaps, such as holidays. For unequal starting dates, one approach begins with the longest available histories, estimates their relationships, and then uses regressions to incorporate assets with shorter records in successive groups. The method can be extended beyond means and covariances if suitable regression assumptions are available. Multiple imputation is described as repeatedly simulating missing observations from fitted models, while expectation-maximization fills gaps with predicted values during its iterative estimation process.
For gaps between observations, the responses mention EM and also suggest pairwise deletion as a simple option, while warning that discarding data can be costly and imputation may introduce bias. When the number of assets is large relative to the sample, PCA or regressions against sector or country indices can impose structure. The discussion cautions that methods based on independent, identically distributed multivariate normal data do not directly handle nonstationary or lag-dependent systems; Bayesian mixed-frequency Gibbs sampling is mentioned for more complex cases. No single best method is established.
Key ideas
- Unequal history lengths can be handled by adding shorter histories through successive regressions against longer ones.
- Multiple imputation simulates missing observations, whereas EM uses predicted values in its iterative procedure.
- Pairwise deletion avoids modeling missing observations but discards data and may affect later analysis.
- PCA or index-based regressions can impose structure when assets outnumber available time periods.
- Methods relying on independent, identically distributed normal data may not suit lag-dependent or nonstationary series.
Tags
Full text
# Handling Missing values in stocks returns when estimating the co variance matrix # Handling Missing values in stocks returns when estimating the co variance matrix What is the best way to handle missing values when stocks did not exist for the entire historical period?. ## Answer by vanguard2k (score 3) https://quant.stackexchange.com/a/14071 One really nice book that comes to my mind is > Little, Rubin, Statistical Analysis with Missing Data I read part of it but probably it is too much information in your case. For your application, i think you can categorize the problem into two possible subproblems: First, time series that have unequal starting points (when some stocks' history is shorter): > Page, S., 2013, How to Combine Long and Short Return Histories Efficiently, Financial Analysts Journal 69, 45-52 Second, data that misses in between the time series (for example on public holidays): Well, there is the EM algorithm. Take a look at it. The most cited paper here is > Dempster, A. P., M. N. Laird, and D. B. Rubin, 1977, Maximum likelihood from incomplete data via the EM algorithm, Journal of the Royal Statistical Society 39, 1-22 It is an iterative, two-step algorithm. You can also find the concrete formulas in Meucci(2005) "Risk and Asset Allocation". On his (Meucci's) webpage you can find the corresponding matlab code. ## Answer by John (score 2) https://quant.stackexchange.com/a/14075 @vanguard2k and @Theja provide useful information. In my experience, unequal starting points is most common, so I'll try to focus on that. The technique that @vanguard2k mentioned for unequal starting points can be thought of like a regression. You start with the longest available data and get the covariance matrix of that. For the next set of available data, you regress them against the data that is available longer and use the regression coefficients to expand the covariance matrix (and means). You then iterate through each unique group of data, steadily shrinking the amount of data used in each regression. The above approach can be considered more general than simply just for estimating means and covariances. If every step is a regression, then you can make any assumptions you want so long as it fits within the context of a regression (for instance, you could estimate a garch model for S&P 500, then regress Facebook against S&P 500 with some garch process as well for the residual variance). You may not be able to get an analytic formula for the covariance, but you can simulate from the model and calculate the simulated covariance matrix from that. As an alternative, multiple imputation would be like if after the regression you use that model to simulate some missing data points. In the next step, instead of using only the available data, you use the available and the simulated missing. You then re-do the above steps many times (where the multiple comes from) until the parameters settle down. I find it can be helpful to learn about multiple imputation before trying to learn about Gibbs sampling. The EM algorithm is also very similar to multiple imputation, with the exception that EM is filling in the missing data with the predicted value whereas multiple imputation is filling in with simulations. Regardless, you may run into difficulties when the number of stocks you're looking at expands to more than the number of time periods (like any covariance estimation, really). One way to resolve this can be to apply PCA to each of the subsets of data, as appropriate. Alternately, you can limit the regression to something like the country/sector/industry indices. Either of these approaches is like making some assumption that certain correlations are zero. Things become more complicated when you start dealing with data that easily be assumed to be I(0). For instance, suppose you want to estimate a VAR model for the S&P 500 (in log levels), S&P 500 E/P, 10 year U.S. Treasury yields, inflation, and VIX. You want a VAR model in this case because there might be some mean-reversion or cointegrating relationships. You will likely have data over the longest period for S&P 500, but less data for inflation, yields, E/P, and VIX. Since the data is not iid multivariate normal, you can no longer use the techniques mentioned by @vanguard2k. The difficulty is that when you simulate data, it needs to depend on both today's value and however many forward or backward lags are needed. There's a Gibbs sampling approach that can deal with this type of situation (Bayesian Mixed Frequency), but it's rather sophisticated. ## Answer by Theja Tulabandhula (score 1) https://quant.stackexchange.com/a/14069 A simpler question would be the following: suppose you want to find the covaraince between the returns of two stocks and each of their time series has missing values at different places. What is the best way to compute covariance here? One very sensible way to approach this is to throw away the observations where ony one of the stocks has a return value. Of course, you are throwing away observations in this approach. But, using any other approach may introduce bias in the later steps of your workflow that may be undesirable (unless you introduce a knowledge based bias to specifically compensate for the lack of observations). Also, check out some of the questions on Cross Validated related to imputation and missing values (for instance, the first answer for this question).
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.