Handling Missing Futures Returns in Historical VaR
Summary
The discussion considers missing daily prices in a futures portfolio, especially where spread legs become misaligned and create artificial profit and loss. It warns that imputing returns changes the observed distribution and can distort historical value at risk. Before filling gaps, determine whether a missing observation represents no trading, a data error, or a market condition that itself signals risk.
Suggested alternatives include using bid and ask quotes to mark positions, examining risk under bid, mid, or ask valuation, and deriving a missing contract’s movement from a more liquid point on its term structure after adjusting the relationship. Statistical proposals include supervised prediction trained by deliberately masking complete observations, benchmark returns scaled by estimated beta, and probabilistic principal component analysis. These are suggestions, not validated solutions; the discussion gives no comparative results. Any imputation should be assessed against the portfolio’s dependence structure and the intended risk measure, with particular caution around illiquid or untradeable markets.
Key ideas
- Imputing returns can alter the historical distribution used to calculate VaR.
- Investigate whether a gap reflects no trading, a data problem, or market illiquidity.
- Bid and ask quotes can provide alternative marks and reveal execution risk.
- A liquid related contract, benchmark model, or statistical predictor may inform an estimate, but proposals require validation.
Tags
Full text
# Imputation of missing returns # Imputation of missing returns I'm trying to calculate a historical VaR for a portfolio of futures, however there are certain days for which some assets are missing prices. Since the portfolio consists of many spread positions, the days that are missing prices have large PnL's when one leg of the spread contains a return and the other leg doesn't. I was wondering what the most appropriate way of imputing this data would be, I've tried using the median/mean for each asset but that ignores the co-variance structure of the portfolio. Would an algorithm such as MICE or EM be appropriate in this situation? ## Answer by PlantFox (score 1) https://quant.stackexchange.com/a/42544 For a VaR calc you might not want to interpolate missing values. By doing that you are inherently editing the returns distribution; potentially this will make a VaR look better or worse. Not good if your goal is an accurate risk distribution. Its worth considering what a missing value signifies. There are two cases. A missing value or unchanged value can reflect a zero trade volume. If you have volume data and it shows that trades occurred, you probably have a data error somewhere. As a result, it is worth considering bid/ask data to help complete this task. Generally in futures markets there is always a bid/ask even if there is no trade volume. You can use this data to compute the mark to market value of each position. Furthermore, it allows you the flexibility to calc VaR based off mark to bid - mid - ask. Now lets consider a case where say bid or ask is missing, that also tell you something about the risk. If there is no bid that means no way to sell and if there is no ask, there is no way to buy. That would signify great risk because that means the the order book is empty for given levels. That is just an example how bid/ask data can inform you. Another method would be to use data from a more liquid contract in the term-structure. You would just have to adjust the price and return to match the target contract. You can do this through a simple ratio and beta calc. One question, are you using daily data or intraday data? Let me know if that helps. ## Answer by Attack68 (score 1) https://quant.stackexchange.com/a/42546 A completely different statistical approach is to pose your own machine learning problem: 1) Collect a set of full data where you have data values available for all instruments on any given day. 2) Propose a machine learning model that will devise its own optimised parameters for the task of regressing any missing data. 3) From your set of good data systematically alter the the data by removing values (which you attempt to predict) under a supervised learning paradigm with a loss function equivalent to the squared error residual. As an example of a possible model: suppose you had the daily changes of 100 stock prices valid on a set of days. You could try implementing a neural network which took as inputs the 100 daily changes and 100 binary flags (0 or 1) stating whether the data was available or not for a particular stock (set an unknown value to zero). The output of the network would be the 100 known stock prices. This is of course completely untested and just an idea but the benefits that this model allows is that data generation is probably plentiful. Even if you had 5y of 100 stock prices you could generate lots of noisy data since for each date you have many billions of combinations of data elements that you could remove and hopefully use to get your neural network weights to converge for a generalist setting. I'm actually tempted to see if this would work... The advantage of this model is that it is non-parametric and agnostic as to which of the 100 stocks are unknown on any given day. It will simply takes the inputs as it sees them and return a vector of 100 expected stock changes. ## Answer by greglama (score 0) https://quant.stackexchange.com/a/69670 More simply, you could use the daily return of a benchmark from the industry of the company. Then you could compute the company's beta versus this benchmark (instead of a market benchmark). You'd replace the missing value by benchmark_return * Beta_regressed. You could also create this benchmark yourself by selecting companies that are highly correlated to the one with missing values... If missing only 1 day at time t, you could use (OPEN_t+1 - CLOSE_t-1)/CLOSE_t-1 OPEN_t+1 is not too far from CLOSE_t, and CLOSE_t-1 not too far from OPEN_t. Those are still a terrible ideas ! Don't invest based on that. ## Answer by Sebastian (score 0) https://quant.stackexchange.com/a/69683 You could use a machine learning technique used for imputing missing data like PPCA (Probabilistic Principle Component Analysis). This is superior to fitting to subset of full data if most rows contain missing data somewhere.
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.