Choosing Return Data for Fitting Financial Distributions
Summary
The discussion considers whether distribution fitting for financial returns should use simple percentage returns, log returns, or standardized observations. The accepted response favors percentage returns when fitting distributions such as Student’s t, arguing that this avoids imposing a lognormal price assumption and retains the interpretation of returns in percentage terms. It also notes that log returns have useful properties, including time additivity, which may be relevant for other analyses.
A key caveat concerns combining Student’s t log returns with expected price or money calculations: the distribution’s moment generating function is undefined, which can imply an infinite expected gross return under that model and produce implausibly large risk outputs. The response presents this as a modeling issue and suggests percentage returns when fat tails are desired. It does not compare fit diagnostics or establish that one return convention is best for every purpose; data transformation should match the quantities and assumptions of the intended analysis.
Key ideas
- Percentage returns avoid assuming that prices follow a lognormal distribution when fitting a return distribution.
- Log returns offer properties such as time additivity that may be valuable for some analyses.
- A Student’s t model for log returns can imply an undefined or infinite expected gross return because its moment generating function does not exist.
- Fat-tailed percentage-return models may avoid that particular issue, but the appropriate transformation depends on the intended analysis.
Tags
Full text
# Fitting (marginal/multivariate) distributions to financial return data # Fitting (marginal/multivariate) distributions to financial return data I have calculated the simple arithmetic return on a number of different financial securities and am fitting both a Student-T and Generalised Pareto Distribution. My question is can I just use the simple returns as my input observations or is it better practice to convert the data into log-normal (i.e. take the exponential) or standardise the data (subtract mean and divide by standard deviation). I know this may seem like a very simple question, but I am looking for the soundest statistical practice and why. ## Answer by Robert Szóstakowski (score 1, accepted) https://quant.stackexchange.com/a/20896 I would personally go for a normal returns, because you do not make any assumptions about the data or returns. When we use log returns we assume that prices are distributed log normally (which, usually is very far from the truth). Moreover if you will investigate different distribution you will not use the log returns features like time additivity or approximate raw-log equality. And if you think about T-Student Distribution this is worth considering: > Mathematically there’s a problem: when you assume a student-t distribution (a standard choice) of log returns, then you are automatically assuming that the expected value of any such stock in one day is infinity! This is usually not what people expect about the market, especially considering that there does not exist an infinite amount of money (yet!). I guess it’s technically up for debate whether this is an okay assumption but let me stipulate that it’s not what people usually intend. This happens even at small scale, so for daily returns, and it’s because the moment generating function is undefined for student-t distributions (the moment generating function’s value at 1 is the expected return, in terms of money, when you use log returns). We actually saw this problem occur at Riskmetrics, where of course we didn’t see “infinity” show up as a risk number but we saw, every now and then, ridiculously large numbers when we let people combine “log returns” with “student-t distributions.” A solution to this is to use percentage returns when you want to assume fat tails.
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.