When Normalizing Financial Data Helps—and When It Misleads
Summary
The discussion explains why researchers normalize financial data: transformations can make measurements from different assets or return series more comparable and can simplify some mathematical or computational tasks. It emphasizes choosing a method to fit the objective, since normalization is not one specific operation and transformations such as logarithms or linear regression have limitations.
The answers caution that market returns often depart from normality, with heavy-tailed alternatives such as Pareto distributions discussed as more realistic descriptions. One response invokes the central limit theorem to argue that large samples may tend toward normality, but this is not a universal justification for transforming financial observations or assuming their distribution is normal. Any simplifying assumption should be checked against the problem and its effect on results. The exchange offers conceptual guidance rather than empirical tests or a prescribed normalization procedure.
Key ideas
- Normalization can improve comparability across assets or return series.
- The appropriate transformation depends on the research objective and the data.
- Financial returns may be heavy-tailed and need not follow a normal distribution.
- The central limit theorem does not by itself justify assuming every financial dataset is normal.
- Computational simplifications should be checked for their effect on results.
Tags
Full text
# Why are we obsessed over normalizing financial data? # Why are we obsessed over normalizing financial data? I have recently began work on some high frequency financial tick data. I have been told to 'normalize' the data as much as possible and run linear regressions through them. In fact, the data doesn't seem to be anymore linear after I did transformations on them (box-cox/log etc) I understand the linear regressions bit, but most financial data is non-normal anyway, so why bother normalizing? ## Answer by Jacob Amos (score 3) https://quant.stackexchange.com/a/14216 Short answer: It offers some degree -- and in many cases, a greater degree -- of comparability between two types of data (different assets, returns, etc.) Long answer: You may already know this, but keep in mind that "normalization" can mean different things (see this question). There are various methods and purposes for normalizing data (financial or otherwise) but keep things in perspective. Normalize when doing so would be helpful for what you're trying to accomplish, and use a normalization technique that is appropriate. Linear regression has limitations, taking the logarithm has limitations, and so on. It's great to have a big toolbox of different data transformations, but part of that is knowing what to use. As an aside, you're right that empirically markets have not exhibited normal returns. In fact, Mandelbrot explains in this article that the Pareto distribution is more realistic. It was published in 1963, but more recently he talks in this book about how the data have continued to demonstrate this pattern. The point is that you may read or hear about normalization techniques that rest on assumptions like normality that may not always be suited to the problem at hand. At the risk of editorializing, the assumption of normality has been subtly embedded in a lot of financial research and it may sometimes be misleading, so make sure to check the assumptions underlying what you're being told. That being said, just because such an assumption may fail to hold with generality (either theoretically or empirically) does not mean it isn't useful for effectively navigating some specific problem from a mathematical or computational perspective. For example, if you have a model that's super-expensive computationally, and you can speed it up a bunch by making some simplifying assumption without significantly altering the results, perhaps doing so is worthwhile. (As long as you rigorously check that the results do in fact remain within a reasonable degree of similarity.) ## Answer by evilbiscuit (score 0) https://quant.stackexchange.com/a/14230 got an answer from one of my pals, thought it might be interesting to share it here. The reason why we often use the normal distribution is because the distribution will be stable regardless of the number of samples (central limit theorem). Imagine you had a normal distribution after transforming x amount of samples, and across time, u get more variables and we will want them to stay in the 'normal' distribution shape by exploiting the central limit theorem. However, if you have exponential/Pareto (or any other) distributions, the distribution will tend to a normal distribution once the sample size is large enough due to the central limit theorem. This way, we can have a consistent model regardless of the time/sample size. Hope this helps, and if anyone have different ideas about this, please do comment here!
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.