Parametric and Nonparametric Models: Assumptions, Examples, and Tradeoffs
Summary
The document explains how parametric and nonparametric statistical models differ in the structure they impose. A parametric model specifies a relationship between variables and a probability distribution for its random component, with quantities such as regression coefficients estimated from data. A nonparametric approach leaves some part of that structure unspecified, such as the functional relationship or error distribution; rank-based methods are one example. Semiparametric models specify some components while leaving others flexible.
The discussion illustrates these ideas with linear regression, a monotonicity assumption tested using ranks, and a sign test for assessing whether a distribution’s median exceeds a threshold. It describes a general tradeoff: when parametric assumptions are valid, those models can have greater statistical power, while less constrained approaches can be more robust. These are broad tendencies, not guarantees. The document also cautions that nonparametric methods are not necessarily parameter-free or assumption-free, and that the distinctions can vary by modeling framework, including Bayesian methods.
Key ideas
- Parametric models specify a functional relationship and a probability model for random variation.
- Nonparametric models leave some part of the relationship or distribution unspecified.
- Semiparametric models combine specified and flexible components.
- Valid parametric assumptions can improve statistical power, while less constrained models may be more robust.
- Rank-based procedures and the sign test illustrate nonparametric inference.
Tags
Full text
# What is the difference between parametric and non-parametric models?
# What is the difference between parametric and non-parametric models?
I'm reading about volatility modelling and I came across the concept of parametric and non-parametric models. For example, GARCH is a parametric model and Realized Volatility is a non-parametric model.
As far as I can tell, parametric models assume the data has certain shape and have some parameters that need to be estimated/fitted and non-parametric models are rather simple and have no parameters?
## Answer by Dave Harris (score 0, accepted)
https://quant.stackexchange.com/a/54833
It is easier to talk about what a parametric model is than a non-parametric one. Parametric models have a well-defined relationship between the independent variables and the dependent variable, and, as well, use a well-defined probability distribution for the chance or random component of the relationship.
In a non-parametric model, something of the above is not well defined.
For example, in the regression equation $$Y=\beta_0+\beta_1X+\varepsilon,\varepsilon\sim\mathcal{N}(0,\sigma^2),$$ every parameter and every variable map to a fixed number. In addition, $\varepsilon$ has a well defined functional form.
Now let us imagine that we do not know the functional form, only that we believe $Y$ is monotonically increasing in $X$. One way we could test that is to convert the observations to ranks. However, we no longer would have a well-defined relationship between the two variables, even if the ranks follow a well-behaved probability mass or density function.
As well, we could have a well-defined form such as $$Y=\beta_0+\beta_1X+\epsilon,$$ except that we have no idea what $\epsilon$ is drawn from.
To complexify the matter a bit, there is also a category called semiparametric models. In a semiparametric model, some parts are very well defined and others are not.
Generally, parametric models have higher statistical power if the model assumptions are actually valid assumptions. Non-parametric models tend to be more robust.
While I spoke of independent and dependent variables, that isn't actually required. There could be only one variable, for example. You could have variable $X$ where you do not know its distribution and you believe it is ill-behaved regardless.
If you wanted to know if the center of location is greater than five, you could use the sign test, splitting them at the median to see if the median is greater than five.
There tends to be a mistaken phrasing with regard to Frequentist non-parametrics that comes up a lot. It is that the data determines the model, or that the data is used more. Neither phrase is true. If the first part were true, then it would be a Bayesian non-parametric model. If the latter were true, then non-parametric tests would be more powerful than parametric ones.
What has happened is that the data is subject to less well-defined relationships which are roughly like loosening constraints. If you go from a linear relationship to a monotonically increasing (decreasing) relationship, then you are making weaker statements.
Generally, with Frequentist models, you are getting less information about how the world works. You are also less straight-jacketed. to your models. That is not necessarily a true statement for Bayesian models because of how Bayesian model selection works and that the likelihood function is minimally sufficient.
Distribution-free and parameter-free models take advantage of other properties of a problem other than the direct relationship between the variables. For example, in all circumstances, a median exists for a distribution. Likewise, even tied variables can be ranked, it is just that they all have the same rank.
## Answer by alexprice (score 1)
https://quant.stackexchange.com/a/54804
In general "parametric" models make a strong assumption (dynamics equation, like Garch, parametric Dupire local vol) about underlying process. Coefficients (parameters) of these equations usually need to be estimated (calibrated).
In "non-parametric" models there's usually less assumptions , and they are estimated directly from data. They do have assumptions to justify the formulas used, but these are usually very general (i.e. Gaussian distribution). Examples are deep hedging, and most of machine learning models.Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.