Gaussianizing Asset Returns with Volatility Models and Factor Models
Summary
The document examines whether non-Gaussian stock returns can be transformed into a Gaussian series. It warns that forcing observations to match a normal quantile distribution can produce a normal-looking sample while distorting the data and removing information about changing volatility. Preserving observation order may retain some autocorrelation, but that does not make the transformation useful for trading or inference.
As a more principled possibility, it describes fitting a GARCH model and scaling returns by estimated volatility, conditional on returns being drawn from a mixture of normal distributions and the volatility model being correctly specified. The author notes that GARCH cannot anticipate upward volatility shocks and that a suitable forecasting model may be complex. For cross-sectional returns, the post discusses factor models such as CAPM, whose residuals would be normal only under strong specification assumptions. Factor models do not eliminate fat tails in general. The examples illustrate methods, not evidence that transformed returns are truly Gaussian out of sample.
Key ideas
- Quantile-based rescaling can make a sample look Gaussian but may distort useful return information.
- Scaling returns by estimated GARCH volatility is proposed under a mixture-of-normals assumption.
- GARCH scaling depends on correct model specification and cannot predict volatility shocks in advance.
- Factor models can describe cross-sectional returns, but their residual normality depends on strong assumptions.
- Neither volatility scaling nor factor models guarantee that fat tails disappear.
Tags
Full text
# Creating a standard Gaussian model on stock market data
# Creating a standard Gaussian model on stock market data
One of the big problems in creating good statistical models in the stock market is because of the long tails that deviate from Gauss' [regular] bell model, is there a way to create a synthetic Gauss bell on market data, by a random walk model that buys long or short [each time] And so it balances the tails completely [after all: if a person enters the market both long and short all the way, he will find himself with a span of zero at the end of the road, and if so a random walk should create a zero span with a standard Gaussian model] I Do not know if my logic is correct, I would love any answer or comment from the great experts here
Thank you
## Answer by Bob Jansen (score 2)
https://quant.stackexchange.com/a/55728
I think I understand where you are going, please correct me if I'm wrong. I also think whatever you could do to transform the returns to Gaussian will be very complex or not really useful. In short, I'm not convinced this will be a fruitful approach.
If you have a sample of returns you can always apply a scaling to individual observations to make them normal. In the plot below, take every point that's for from the line and scale it so that it's closer to the line, repeat until your convinced the result is normal. I'm unconvinced this is a good idea. If you keep the observations in order you would at least not lose some forms of auto-correlation.
Assuming (big assumption) that returns are drawn from a mixture of normal $N(0, \sigma)$ distributions a better approach would be to model the volatility of the returns using GARCH and scale the returns using the inverse of the predicted volatility. If the model is correctly specified the returns would now be normal. This has two drawbacks, one model specific, one more fundamental:
- GARCH doesn't forecast upward volatility shocks, you can only now-cast which is a limitation for investing;
- The existence of a simple model that forecasts volatility is unlikely. It will almost be certainly be more complicated than what you're trying to achieve right now.
At least, it has a stronger theoretical basis than the scaling described above.
If we drop the assumption of returns coming from a mixture, you would need another model, presumably a lot more complex. I think Cont, R. (2001) Empirical Properties of Asset Returns: Stylized Facts and Statistical Issues. Quantitative Finance, 1, 223-236. is the standard reference for stylized facts you would want to account for.
## Transforming the data to Gaussian
The plot above is created in R with:
```
set.seed(1L)
returns <- rt(1:1e3, df = 10)
qqnorm(returns)
qqline(returns)
```
### Returns normalized by force
These returns can be made Gaussian as follows:
```
# Don't do this
scaling <- qnorm((1:100) / 101)[order(returns)] / returns
qqnorm(returns * scaling)
qqline(returns * scaling)
shapiro.test(returns * scaling)
```
Results in
```
Shapiro-Wilk normality test
data: returns * scaling
W = 0.99774, p-value = 0.9999 # Higher p-value => more evidence of normality
```
This forces `returns * scaling` to be normally distributed and at least retains the original ordering. It removes all information about heteroskedasticity.
### Using GARCH
The quantile matching method above is not great but the calculation of `scaling` can be done in other ways too. If you were to fit a GARCH model on the return series and calculate the volatility for each period you could set
```
scaling <- 1 / volatility
```
and proceed as above.
To mimic Gaussian returns in your portfolio, you can scale your portfolio by `scaling` as well.
### Using factor models
The above methods work on time series data, if your interested in the returns on multiple stocks at one moment you can use a factor model. For example, the CAPM. If the CAPM was correct, returns would be given by: $$R_i = R_f + \beta_i(R_m - R_f) + \varepsilon_i$$ where $R_i$ denotes the return on asset $i$, $R_f$ the risk free rate, $\beta_i = \frac{\mathrm{Cov}(R_m, R_i)}{\mathrm{Var}(R_m)}$ the scaled covariance of the returns of $i$ with the market, $R_m$ the market return and $\varepsilon_i$ the return on $i$ not explained by the model. If the model would be well specified $\varepsilon_i$ is normally distributed with mean $0$ and standard deviation $\sigma_i$. Then $$R_i \sim N(R_f + \beta_i(R_m - R_f), \sigma_i).$$
Other factor models are the Fama & French 3 factor model or the Carhart four-factor model.
However, these factor models don't really get rid of fat tails as they don't capture all there is to know about returns.Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.