What Linearity Means in Classical Statistics and Estimation
Summary
The discussion clarifies that “linearity” in classical statistics is broader than whether a formula contains only first-degree terms. It refers to properties of relationships and estimators, including linear combinations of random variables and linearity conditions associated with estimation. The answer uses normal variables, whose linear combinations remain normal, to illustrate why linear structure simplifies statistical analysis and filtering.
It also points to the Gauss–Markov result: under conditions including zero-mean, uncorrelated, homoscedastic errors with finite variance, ordinary least squares is the best linear unbiased estimator. Pearson correlation is described as linear in its arguments, and linear regression is connected to projection in a vector space. These examples explain the respondent’s interpretation of the passage, but do not fully survey the historical or philosophical claim about classical statistics.
Key ideas
- Statistical linearity concerns relationships and properties, not merely the visual form of a formula.
- Linear combinations of normally distributed variables remain normal, supporting tractable methods such as Kalman filtering.
- Under specified error conditions, ordinary least squares is the best linear unbiased estimator.
- Linear regression can be understood as a projection in a high-dimensional vector space.
Tags
Full text
# What's the meaning of linearity in classical statistics in Prado's book?
# What's the meaning of linearity in classical statistics in Prado's book?
I am reading Prado's new book, Machine Learning for Asset Managers.
In the page1 of his book, there is this sentence.
> To a greater extent than other mathematical disciplines, statistics is a product of its time. If Francis Galton, Karl Pearson, Ronald Fisher, and Jerzy Neyman had had access to computers, they may have created an entirely different field. Classical statistics relies on simplistic assumptions (linearity, independence), in-sample analysis, analytical solutions, and asymptotic properties partly because its founders had access to limited computing power.
I guess the independence in the quote originated from IID (independent, identical distribution) assumption.
But I'm not quite sure where the linearity of classical statistics comes from. The one place I could think of is about the linear regression. But it is pretty easy to extend linear regression to higher-order using the same Gauss-Markov theorem with the basis function approach.
## Answer by nbbo2 (score 4)
https://quant.stackexchange.com/a/53216
It appears that you are using "linearity" in a litteral sense while De Prado is using it in a broader sense, which is quite common in Statistics.
In Statistics Linearity is not what the formula looks like, it is the properties and assumptions of the system under study. You consider the Normal distribution non-linear because it has an exponential and some squares in the formula. The statisticians frequently use the Normal distribution because it has the nice property that linear combinations of Normal variables are also Normal, which is not necessarily true for other distributions. The whole field of Kalman Filtering for example relies on this property and Non Linear and/or Non Gaussian filtering is very hard to do because you can no longer rely on this basic fact.
An important result of Classical Statistics is the Gauss Markov Theorem: Ordinary Least Squares provides the Best Linear Unbiased Estimator if the errors are (linearly) uncorrelated with mean zero and homoscedastic, finite, variance. Note the word linear appears twice.
Pearson Correlation is also called Linear Correlation, even though the formula has some squared terms in it. It is properties like $\rho(A+B,C)=\rho(A,C)+\rho(B,C)$ which make it linear.
As to the relationship between Classical Statistics and Linear Algebra, it is extensive. Consider for example that the linear regression can be written as $b=(X^T X)^{-1} X^T Y$. Yes, that involves Linear Algebra, it can be interpreted as a projection in a high dimensional linear space.
Like it or not the observations of De Prado you quoted are commonplace, he has not said anything original here (after all, he is still on Page 1 ;) ).Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.