Why Resampling a Noisy Option-Implied Density Cannot Ensure a Better CDF
Summary
The document addresses whether a probability density inferred from option prices using the Breeden–Litzenberger relationship can be smoothed into a statistically valid CDF by fitting a kernel density estimate to the density values. It argues that resampling points from a noisy estimated distribution and fitting another distribution does not create new information or reliably remove the original noise. Bootstrap methods can help estimate uncertainty in an estimator, but do not automatically improve quantile or CDF accuracy; quantile inference may require specific corrections.
Smoothing may regularize an estimate, but its success depends on assumptions about the underlying distribution and the method used. A parametric fit can work when its assumed family matches the true distribution, but the document gives no general procedure for choosing a valid fit. Because the source is a risk-neutral distribution derived from market prices, adjustments should be justified in terms of their impact on those prices. The response particularly cautions against altering estimates at actively traded strikes based only on a statistical smoothing procedure.
Key ideas
- Resampling from a noisy estimated CDF does not add information about the underlying distribution.
- Bootstrap can estimate estimator variance, but it does not automatically produce accurate quantiles or a better CDF.
- Smoothing can regularize a density estimate, though results depend on the underlying distribution and the fitting assumptions.
- A parametric fit is most defensible when its assumed family is appropriate for the true distribution.
- Smoothing risk-neutral probabilities should be evaluated against the option prices from which they were inferred.
Tags
Full text
# How to fit KDE from existing probability density function values # How to fit KDE from existing probability density function values I am working with options data, and I am using Breeden-Litzenberger formula to derive the risk-neutral terminal stock price PDF. After applying the formula, here is a scatter plot of strike price vs Breeden-Litzenberger pdf: At this stage, I would like to fit KDE using `statsmodels.nonparamteric.KDEUnivariate` and then use `scipy.interpolate.interp1d` to obtain a CDF function, ensuring the area under the fitted PDF = 1 etc. How do I do this with PDF values and not with a sample from the distribution in question? Should I do bootstrapping? Should I fit the CDF with `GammaGam` to have values > 0 and imposing `constraints='monotonic_inc'`? The overall goal is to obtain a CDF function, making sure it is actually statistically correct. Any input would be appreciated! ## Answer by lehalle (score 0, accepted) https://quant.stackexchange.com/a/73711 From a statistical viewpoint, it is not standard to be in a situation to - have a "noisy CDF", - sample points from it, - deduce another CDF that would be "less noisy". You can repeat point (2) and (3) to "bootstrap the CDF", but what would it means? First, you have to know that bootstrap is not magic: it allows to "naively" obtain an unbiased estimate of the variance of an estimator, but nothing more. If you want to obtain quantiles (that is exactly what you aim for, since empirical CDF are made of quantiles), you have to apply some corrections. Efron (the author of bootstrap), has a nice paper on that topic "Bootstrap confidence intervals" by DiCiccio, Thomas J., and Bradley Efron (1996). Qualitatively, it is clear that if you do not have enough sample points, you cannot obtain better estimates of quantiles, just reusing your points. It is only valid at the asymptotic limit (and if you have an infinity of points, you do not need bootstrap). Second, starting from a noisy CDF cannot generate sample points that are not noisy. Your best hope is that if you sample "few enough" points, the method you use at point (3) would "regularize" the "secondary" CDF. The truth is that there is no good reason that for, except if the "true, not noisy, underlying CDF" is of the same family as the one use by method of your step (3). For instance, to make it very simple: - if the underlying CDF is a Gaussian, - and step (3) is assuming it is a Gaussian, hence it just computes its mean and variance. - Then of course it may work. Third, to be very practical, you should keep in mind that you are talking about derivatives, market prices and risk-neutral probabilities: if you change it, it will say something about market price. The cost of modifying your original distribution should be reflected on market prices (i.e. at given strikes), and should not only come from a statistical procedure. For instance you shouldn't modify points are strikes that are heavily traded, because market participants "strongly believe" in them.
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.