Imputing Missing Histories for Historical Simulation VaR
Summary
Historical simulation VaR requires past market observations, but new securities may lack the price and volatility histories needed to construct stress scenarios. The note presents two broad ways to address the gap: statistical completion and economically informed proxies. One response reports using probabilistic principal component analysis to estimate missing equity returns from patterns shared across assets, then feeding the completed series into historical VaR. It describes an implementation in which returns are standardized, a lower-dimensional factor representation is fitted, and the estimates are transformed back to the original scale.
A second response emphasizes that statistical estimates may lack economic meaning. For an issuer with little history, it suggests starting with its industry or sector as a proxy and adjusting for company information such as leverage. These are alternatives with different strengths: matrix completion can exploit common return structure, while proxies are easier to explain economically. The document offers practitioner experience but no comparative validation, VaR accuracy results, or guidance for missing implied volatility surfaces. Imputed observations remain model estimates, so their assumptions and uncertainty matter.
Key ideas
- Historical simulation VaR needs proxy data when an asset has no sufficiently long history.
- Probabilistic principal component analysis can estimate missing returns by exploiting common patterns across assets.
- Sector or industry returns can serve as an economically interpretable proxy for a new security.
- Issuer information such as leverage may help adjust a proxy series.
- The document does not establish that either approach produces accurate VaR estimates in all settings.
Tags
Full text
# Missing data in historical simulation VaR
# Missing data in historical simulation VaR
A historical simulation approach to VaR estimation relies on the availability of historical data. What do we do when there is no data (say, spot price and implied volatility surface) as, for example, in cases of new equity issues or new bond issues? ("New" in the previous sentence does not necessarily mean "recent" as for SVaR one would probably need data from 2007-2008).
So, what can we do to "fill that missing data"? What are the best-industry-practice methodologies for that? I can envisage approaches like "proxying" and regression, but these seem somewhat crude and primitive. At the other end, I can also envisage a nonlinear autoregressive neural network with exogenous inputs (NARX), but not sure if these are actually used in practice for data-filling.
I would have thought that this is a very common problem and expected to find a lot of literature on this topic, but alas my search did not reveal anything.
## Answer by Jonathan (score 7)
https://quant.stackexchange.com/a/47475
This issue is incredibly important and I agree there is little practical information about it. To me, the key idea is to find the right matrix completion algorithm that best suits your needs. I work mostly with equity time series and there are substantial missing values issues due to, e.g., as you cite, IPOs with limited history. Recently I have had good success with probablistic PCA as a completion step. This is well implemented in the python package pca-magic (appropriately named).
As an example, if your returns are in a dataframe indexed by time and the columns are the (unique) identifiers (and, thus, many NaNs in the data), you can simply do
```
from ppca import PPCA
from sklearn.preprocessing import StandardScaler
scaler = StandardScaler()
scaler.fit(df.values)
returns = scaler.transform(df.values)
n_comp = 20
ppca = PPCA()
ppca.fit(data=returns, d=n_comp, verbose=False, tol=1e-5)
betas = ppca.C
df_beta_returns = pd.DataFrame(index=df.index, data=ppca.transform())
common_returns = np.dot(df_beta_returns, betas.T)
imputed = scaler.inverse_transform(common_returns)
```
And here is one asset where I impute returns that never existed in the past. Magic!
In your use case, you could use the `imputed` dataframe as the input to your historical VaR calculation.
The paper for this algorithm can be found at http://www.robots.ox.ac.uk/~cvrg/hilary2006/ppca.pdf
## Answer by AK88 (score 2)
https://quant.stackexchange.com/a/47480
I think the issue can be addressed in two ways:
- statistical approach;
- economic approach.
While I agree that ML/AI and other statistical tools can enhance missing data in time series, these lack economic meaning. One can implement these techniques and generate some numbers for the simulation. However, the derivation and the end result should also be meaningful from economics and finance perspective. I think that is why many people still approach this problem with proxying and using simple OLS to fill in the missing data.
Although it is very subjective, one can make a reasonable case for choosing a certain proxy. For the newly issued stock, where you basically have no information about the company, I'd say that the best approximation is to mimic that company's industry/sector. This can be augmented using the most recent financial statements, from which one can extract some information such as leverage for example, and adjust the time series accordingly.
Very interesting question. Would love to see comments/inputs from the community.Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.