Parametric Regression, Linear Projections, and Nonparametric Estimation
Summary
The document compares two ways to relate an outcome to predictors. A parametric approach specifies a functional form, such as a linear model, and studies whether an estimator recovers its parameters under suitable error assumptions. A second approach avoids specifying the full relationship and instead targets the best linear approximation to the conditional mean, with ordinary least squares converging to that projection under stationarity and other regularity conditions.
The included answer describes nonparametric methods as more flexible, but notes that estimating them reliably usually requires more observations and can be harder to interpret than a linear coefficient. Another answer points toward total least squares and principal components, which minimize distance in both variables rather than only vertical residuals. These are related estimation ideas, but they do not directly answer the same question as estimating a conditional mean; the document gives no derivations or comparisons of finite-sample performance.
Key ideas
- Parametric regression assumes a specified functional form, often a linear relationship.
- A linear projection can summarize the conditional mean without claiming that the true relationship is linear.
- Nonparametric flexibility can require substantially more observations for precise estimates.
- Ordinary least squares coefficients are usually easier to interpret than local nonparametric estimates.
- Total least squares and principal components address errors in both dimensions rather than only vertical residuals.
Tags
Full text
# Is there a relation between these two forecasting/estimation approaches? # Is there a relation between these two forecasting/estimation approaches? When learning econometrics I have usually seen stuff from the following perspective: - Assume $Y_t = f(X_t) + e_t$, where f is some function of $X_t$ (typically linear). For example, assume $Y_t = X_t * \beta + e_t$. Then if $e_t$ satisfies certain properties the OLS estimator will converge to beta. However I have also seen, but less frequently: - Make no assumption on the function relationship between $Y_t$ and $X_t$. Without any assumptions we know there exists an optimal linear approximation of $E[Y_t|X_t]$ (the alpha such that $X_t*\alpha + e_t$ minimizes MSE, for example). Now if we assume that $(Y_t,X_t)$ is covariance stationary, the OLS estimator converges to alpha. To me it seems like the perspective of 2. is more interesting because the analysis is not predicated on assuming that Y and X have a specific functional relationship. Instead, assumptions like "covariance stationary" seem more general than assuming that $Y = a + bX + e$. Is there a reason why there seems to be more of a focus on 1.? Are the two perspectives related in some way? ## Answer by Jianxun Li (score 1) https://quant.stackexchange.com/a/18414 Approach 1 is parametric regression, whereas approach 2 is non-parametric regression. How are they related: non-parametric regression models the entire distribution of all possible function forms, and then do the integration to calculate a single value E[Y|X]. It is function-form free. In contrast, parametric linear regression ASSUMES that the function form can be well described by a simple linear relation between Y and X. So, yes, approach 2 is more flexible than approach 1. However, such flexibility does come at a cost. To reach an acceptable standard error in estimates, Non-parametric regression typically requires a much much larger amount of observations than the simple linear regression. Also, linear regression has a straightforward interpretation (1 unit increase in X would drive Y up by alpha units), but non-parametric regression result is not so intuitive (you can think of it as an weighted average of Y taken around the neighbourhood of X=X0). ## Answer by Guys Math (score -1) https://quant.stackexchange.com/a/18397 look at the econometrics literature on "total Least squares" (van huffel has a text out by that name)...or more generally think about what principal components does (hint: it's minimizing the distance to a regression line as in #2..it's not minimizing just the "vertical" distance)
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.