Avoiding Look-Ahead Bias When Filling Missing Returns
Summary
The document considers whether an empirical covariance matrix calculated from returns imputed by an expectation-maximization procedure can be trusted. Although the questioner reports good covariance estimates in simulations with substantial missing data, the response focuses on a more fundamental issue: an imputation used in an investment strategy must not use information that would have been unavailable at the date being filled.
For a missing return at time t, estimate relationships using only observations available before t. The response suggests a point-in-time factor model, such as a rolling principal component analysis, then a regression of the asset’s past returns on those factors. This reframes the problem as historical covariance estimation rather than relying on a full-sample reconstruction. The advice is cautious: market information can arrive asymmetrically, so time reversibility and stationarity should not be assumed. The document does not establish that one factor approach is best or resolve all covariance-estimation choices.
Key ideas
- Imputing historical returns with future observations can introduce look-ahead bias.
- Estimate the information used to fill a return using data available before that date.
- A rolling principal component model can provide point-in-time factors.
- Past returns can be regressed on those factors to estimate an asset’s relationship to them.
- Market dynamics may not be time-reversible, which limits full-sample inference.
Tags
Full text
# Covariance matrix of Gaussian EM output # Covariance matrix of Gaussian EM output I have a project where i wanted to use Expectation Maximization to fill in missing logreturns. With regards to that I have a question I haven't been able to solve. Logically EM should decreese variance of the data as the estimations will be the "Most likely" results. Covariance effects decreese this, however I believe this is still the case. In other words, can I reliably use a covariance matrix calculated on the full output of the algorithim? I saw that doing this gave me the best estimates of the emperical covariance matrix every simulation, even when truncating up to 60% of my test data and re-estimating it. Still it just doesn't make intutive sense to me... I also considered if I should use the conditional covariance matrix that the algorithm calculates, however that gives worse estimates. TLDR: Is it problematic to use EM to estimate up to 60% of my data and then emperically calculate the covariance matrix of the output using the default formula? ## Answer by lehalle (score 2, accepted) https://quant.stackexchange.com/a/79085 If you replace missing returns (and indeed if you replace anything that can be used as an input of an investment strategy), it is strongly recommended to never use future information (it means: to replace a value at time $t$, do not use any data that have been available after $t$). If you need a covariance matrix to replace returns of a stock $k$ at date $t$ (they are so many difference way), you are right that you thus should not use any data after $t$. Your covariance matrix should be estimated only using past data. It will drive you to the traditional problem of estimating covariance matrices (so many possibilities... have a look on stack exchange only). My advice would be to not really use the full covariance matrix you have in mind but - create a point in time factor model, the "simplest" being a sliding PCA - just estimate (from past data) the coefficients of a regression of the past returns of your stock $k$ with your factors as covariates (ie explanatory variables). [EDIT] following a comment about my "never use future information" recommendation. Unfortunately, there is often not time reversibility on markets, simply because the arrival of information has an asymmetric effect on price formation (see for instance Marcaccioli, Riccardo, Jean-Philippe Bouchaud, and Michael Benzaquen "Exogenous and endogenous price jumps belong to different dynamical classes" Journal of Statistical Mechanics: Theory and Experiment 2022, no. 2 (2022): 023403). This is not really a matter a non stationarity (it is indeed worst than that): they are two effects than are layered, first time-revertible dynamics (when no exogenous information occur), and then one that cannot in general be reverted. So no need to take any risk: just use information in the past. I am not saying it is impossible to make the correct change of variables and projections so that one ends up in a stationary environment, just that it is subtle and so it is better to avoid approaching this kind of problem.
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.