Estimating Two-Normal Mixtures with Sparse Tail Observations
Summary
The document considers fitting a two-normal mixture to stock returns, with one component intended to represent the distribution's central behavior and another its tail. It highlights the central estimation problem: rare tail events may be absent from a short record, or a small sample may contain an unusually large move that leads to an exaggerated tail estimate.
As a way to reduce the number of free parameters, it proposes assuming equal mixture weights and a shared mean across the components, leaving the two volatilities and common mean to estimate. It suggests fitting the model directly for stocks with substantial representative data, then asks how information from many other stocks could help when an individual stock has little history. A possible constraint linking tail volatility to central volatility is raised, but no estimation method or answer is supplied. The assumptions require validation, and the document offers no data, fitted parameters, or evidence that the proposed constraints improve tail forecasts.
Key ideas
- A two-normal mixture can represent central returns and a separate tail component.
- Sparse histories make tail parameters difficult to estimate because extreme observations are rare.
- Assuming equal mixture weights and a common mean reduces the number of free parameters.
- The document proposes using cross-stock information to constrain tail volatility when an individual history is short.
- The proposed assumptions and constraints are questions for validation, not demonstrated results.
Tags
Full text
# How to estimate mixture of two normal distributions? With limited data
# How to estimate mixture of two normal distributions? With limited data
In finance a mixture of two gaussian models sometimes used. When one normal used to capture the head of the real distribution and another the tail.
$$f(x) = \lambda \cdot \mathcal{N}(x \mid \mu_1, \sigma_1^2) + (1 - \lambda) \cdot \mathcal{N}(x \mid \mu_2, \sigma_2^2) $$
But how to estimate it for specific stock X? There could be not enough data for tail, as by the definition itself tail events are rare.
There are 5 parameters to estimate. Let's try to reduce it and assume that $\lambda=0.5$ and $\mu=\mu_1=\mu_2$ for all stocks (we can validate if our assumption correct by backtesting on lots of various stocks and see how it works).
Now we left with only 3 parameters $\mu,\sigma_1,\sigma_2$. How to estimate it?
For stocks with huge data, that's representative of the real underlying distribution, we can use brute force methods and fit our model.
But, what to do with stocks when there's not enough data, say a) no tail events and we would miss it, or b) the opposite situation - tail event present when there's not much head events, and we may overestimate the tail.
UPDATE: I guess, what I'm asking about, is how to incorporate extra information into the model fitting process. Say we see stock of new tech startup, or new silver mining company, it's been only 3 years on the market, the stock looks nice and smooth. But we have "extra information" - we observed 1000 other stocks, and know that there could be turbulence. So, I guess we need somehow to establish relationship between $\sigma_1$ and $\sigma_2$. We can easily estimate $\sigma_1$ (there are enough head events), and then, use it to somehow set lower bound for $\sigma_2 > K\sigma_1$ or something like that.
P.S. Also references to books etc. explaining it in details appreciated.Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.