Skip to content
All library documents

When Normality Assumptions Matter in Machine Learning Models

Article Quant Q&A · Author: Zichu Lee

Summary

The document asks whether a skewed target must be transformed to a normal distribution and which machine learning methods rely on normality. It names Linear Discriminant Analysis and Quadratic Discriminant Analysis as methods derived under a multivariate normal distribution assumption. Decision trees are described as generally needing less preprocessing, though preprocessing may still improve their detection of useful patterns.

The response cautions that a model’s assumptions do not by themselves determine whether it will work: LDA and QDA can perform well even when the data depart from normality, but they can also fail. It frames transformation as a modeling choice to assess empirically rather than an automatic requirement for every algorithm. The discussion is brief and does not give diagnostic procedures, comparative results, or a specific rule for when to apply a Box-Cox transformation; its main lesson is to distinguish model assumptions from guaranteed performance.

Key ideas

  • Linear and Quadratic Discriminant Analysis are derived using a multivariate normality assumption.
  • Decision trees often work without extensive preprocessing, although preprocessing can still help.
  • A violated model assumption does not guarantee poor predictive performance.
  • Normalizing a target is not presented as a universal requirement across machine learning methods.

Tags

Full text
# Which machine learning model rely on the normality assumption?


# Which machine learning model rely on the normality assumption?












In the machine learning project, when the target variable is skewed, we need to use box-cox transformation to turn that into a normal distribution.

- But why do we need to do that? I mean, besides the linear regression, which model has the assumption that the target variable should belong to the normal distribution?

- If we use random forest that don't have any assumption on the data, do we have to transform the data?

thanks

## Answer by Attack68 (score 1, accepted)

https://quant.stackexchange.com/a/45902

Data pre-processing is often a very important (if not the most important) step in a machine learning algorithm. Decisions trees are often an exception that they can work well without any pre-processing. But they may work better if you can identify some processes that might improve the quality of the decision detection.

As an example of other machine learning models: Linear Discriminant Analysis, or Quadratic Discriminant Analysis are both models that are explicitly calculated from the assumption of the distribution being a multivariate normal.

However, all that these models do is create either a 'linear' decision boundary or a 'quadratic' decision boundary to separate classes (in a classification problem)

This has been shown to give good results often when the data is not necessarily normally distributed; so my point being that just because a model assumes one thing that isn't necessarily true does not mean it won't still be a valid and effective way of generating accurate results.

Of course it might also fail miserably - herein lies the art of machine learning.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.