Estimating Risky Bond Loss Volatility from Historical Data or Simulation
Summary
The document asks how to estimate the variance of returns on a risky fixed income investment when the default-free yield and the yield shortfall attributed to defaults are known. Its central recommendation is to obtain historical loss yields or specify a data-generating process, rather than infer a full loss distribution from the mean alone. Historical observations can provide empirical moments; a simulation can model losses using pool characteristics and economic drivers such as GDP, leverage, or home values. The answer also suggests combining historical and fundamentals-based estimates.
The discussion cautions that simulation based on an assumed distribution requires a variance assumption, which shapes the result. It advises against assuming defaults are uncorrelated without strong justification and notes that a floating default-free yield may move with default risk. Without sufficient data, estimated variance will largely reflect modeling assumptions and simulation noise. The document gives no fitted model, dataset, or comparative performance evidence, so it offers a modeling framework rather than a calibrated solution.
Key ideas
- Historical loss-yield observations can be used to estimate empirical mean and variance.
- A simulation can link pool losses to economic and borrower characteristics.
- Combining historical data with fundamental drivers can reflect both past behavior and economic conditions.
- Distributional assumptions, including assumed variance, materially affect simulated risk estimates.
- Default independence should not be assumed without justification.
Tags
Full text
# Best simplified way to model volatility in returns of an investment in a risky fixed income asset
# Best simplified way to model volatility in returns of an investment in a risky fixed income asset
I am currently working on a project where I have analyzed a certain category of fixd income instruments, and I now have the gross aggregate yield as well as the theoretical gross-aggregate default-free yield. Taking the difference of these two yields leaves me with what some people call the "loss rate" of the investment. My question is that if I wanted to model these investments as volatile using that information, what is the best way to do it/best distribution? I realize that the probability of default already factors into the risky yield, but I want something a bit better than that. For example, if I could use the mean yield and the loss rate to construct some sort of simple distribution, I would ideally be able to calculate the variance and therefore be able to give some sort of risk-adjusted performance measure.
It does not have to be fancy at all, and I can assume that all defaults are uncorrelated. What should I use? The exponential distribution? Or a gamma distribution? I think the gamma distribution or beta distribution might be best but I can't quite remember how to translate the mean return and loss rate to the parameters of those distributions. Any help would be greatly appreciated.
In other words, I have modeled that actual yield from my investments as follows:
$$Y = Y_{df}-Y_L$$
where $Y_{df}$ is the default-free yield, and $Y_L$ is the loss in yield due to defaults. So the question really amounts to modeling the distribution off losses, i.e. the distribution of $Y_L$.
EDIT: For example, if I assume that the probability of default and the loss given default were independent, I could write the expected loss in yield as the (probability of default)*(loss in yield due to default). Then if I further assume defaults are uncorrelated (an assumtion which I hope to remove but will use until I can get the bare bones model in place), I think I could model the expected losses as some sort of decay process. In this case I think I could simply use an exponential distribution for the number of defaults, which I think I can calibrate if I know the probability of default and the expected loss rate. But I am not sure if this is the best model or which model would allow me to get around those simplifying assumptions, particularly the one where the loss given default is independent of the probability of default.
## Answer by Kyle Balkissoon (score 2)
https://quant.stackexchange.com/a/16037
To Recap:
Your "Note" is a pool a of loans of which are expected to pay Yield Ydf.
You want to estimate the mean and variance of the Loss in yield of non payment.
First and foremost you need to get a historical YL or at least a Data Generating Process for YL.
Some approaches
A) Historical Calculate historically implied loss in yield and then use that time series to extract mean and variance (This works if you have Ydf and Y)
B) Simulation: Depending on what you know about the pool of loans you can simulate it's performance historically and in the future and use that to extract distributional moments.
e.g. Loss in pool_t = Beta*Gdp_t + Beta2*Leverage_t + Beta3*HomeValue_t
This generates a YL
C) MCMC: You assume some distribution + mean + moments and simulate accordingly, problem is you need an assumed variance (which will contaminate your simulation)
D) Mix of A and B You can build an ensemble using A (historical) and B (Fundamental) to ensure that your model reflects economic fundamentals and historical time series.
I would never assume loan defaults are uncorrelated unless you can very much justify this.
Please note:
Some potential twists, if Ydf is dependent on interest rates (e.g. floating) defaults may increase or decrease with it, so model it appropriately.
All in all, without much data your variance calculation will end up being a result of your distributional assumptions + simulation noise.Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.