Why OLS and Maximum Likelihood Can Differ in Autoregressive Models
Summary
The document considers whether maximum likelihood and ordinary least squares estimates coincide for an autoregressive process, focusing on the role of Gaussian errors. Its answer uses an AR(1) illustration that can be extended conceptually to AR(2). Although a current innovation may be uncorrelated with the prior observation, the current dependent variable includes that innovation and becomes a regressor for later observations. Consequently, conditioning on the full sample’s lagged values can make the innovation dependent on regressors used in the regression.
The answer writes the OLS estimate as the true coefficient plus a ratio involving lagged observations and innovations, then notes that finite-sample unbiasedness would require a conditional mean-zero property that fails in the illustrated setup. It points readers to an example involving Gaussian errors. The discussion raises a distinction between Gaussianity, unbiasedness, and estimator equivalence, but it does not fully derive the maximum likelihood estimator or establish a general comparison for AR(2). It also does not spell out the precise assumptions under which asymptotic equivalence may hold.
Key ideas
- In an autoregression, an innovation enters the current observation, which can serve as a regressor for future observations.
- A current innovation can be uncorrelated with the prior observation without being conditionally mean zero given all sample regressors.
- The answer uses this dependence to explain why finite-sample OLS unbiasedness may fail.
- The document does not provide a full derivation of when OLS and maximum likelihood estimates coincide.
Tags
Full text
# Time Series Multiple Choice
# Time Series Multiple Choice
1) Consider a standard AR(2) process. When is the maximum likelihood estimator identical to the OLS estimator? (a) when $\varepsilon $~ (N 0,$\Sigma^2) (b) always?
I'm thinking (a), but that I also need to add in large samples this would be correct.
## Answer by phdstudent (score 1, accepted)
https://quant.stackexchange.com/a/38861
In an simpler AR(1) case (you can generalize to AR(2)) we have that:
\begin{equation*} y_{t}=\beta y_{t-1}+\epsilon _{t}, \end{equation*} Even under the assumption $E(\epsilon_{t}y_{t-1})=0$ we have that \begin{equation*} E(\epsilon_ty_{t})=E(\epsilon_t(\beta y_{t-1}+\epsilon _{t}))=E(\epsilon _{t}^{2})\neq 0. \end{equation*} But, $y_t$ is also a regressor for future values in ain AR model, as $y_{t+1}=\beta y_{t}+\epsilon_{t+1}$.
Or in other words:
$$\hat\beta =\beta + \frac{\sum_{t=2}^Ty_{t-1}\varepsilon_t}{\sum_{t=2}^Ty_{t-1}^2}$$
For unbiasedness we need
$$E\frac{\sum_{t=2}^Ty_{t-1}\varepsilon_t}{\sum_{t=2}^Ty_{t-1}^2}=0.$$
But for that we need that $E(\varepsilon_t|y_{1},...,y_{T-1})=0,$ for each $t$. For AR(1) model this clearly fails, since $\varepsilon_t$ is related to the future values $y_{t},y_{t+1},...,y_{T}$.
Take a look of to this example with gaussian errors: http://www.alexchinco.com/bias-in-time-series-regressions/Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.