Skip to content
All library documents

How ARIMA Prediction Interval Width Relates to Sample Size

Article Quant Q&A · Author: Ericcheng

Summary

The document explains why a one-step-ahead prediction interval from a fitted ARIMA model need not automatically narrow as the estimation sample grows. In a simple ARIMA specification with Gaussian, homoskedastic errors, the conditional forecast variance is the innovation variance. If model parameters are treated as known after substituting their maximum-likelihood estimates, that interval calculation does not directly include uncertainty from estimating those parameters.

The example is a conditional one-step forecast given the latest observation and residual. The stated interval therefore reflects assumed error variance, rather than parameter uncertainty. In practice, estimated parameters can change when observations are added or removed, so the resulting interval may move. The explanation is limited by its simplifying assumptions: other forecast horizons, model specifications, error distributions, changing volatility, and intervals that account for parameter uncertainty may behave differently.

Key ideas

  • For a one-step ARIMA forecast under Gaussian homoskedastic errors, the conditional forecast variance is the innovation variance.
  • Plugging in estimated parameters while treating them as known omits parameter estimation uncertainty from the interval calculation.
  • Under that simplified calculation, interval width is not directly a function of sample size.
  • In practice, changing the sample can change fitted parameter estimates and make the interval vary.

Tags

Full text
# Relationship between Data Size and Arima Prediction Interval Width?


# Relationship between Data Size and Arima Prediction Interval Width?












When we use Arima model to acquire Interval Predictions, will the width of prediction intervals decrease if we use more data (longer history) to fit the model?

## Answer by Stéphane (score 1)

https://quant.stackexchange.com/a/51329

Take an ARIMA(1,0,1) for simplicity: \begin{equation} y_t = \phi_0 + \phi_1 y_{t-1} + \theta_1 \epsilon_{t-1} + \epsilon_t. \end{equation} Typically, this is estimated by maximimum likelihood which requires us to make an assumption about the distribution of $\epsilon_t$. Most of the time, people pick a Gaussian distribution and impose homoskedasticity, i.e., they say $\epsilon_t \sim N(0,\sigma)$.

For simplicity, we'll do the one-step ahead prediction interval, conditional on $(y_T, \epsilon_T)$: \begin{align} y_{T+1} | (y_T, \epsilon_T) &\sim N(\mu_T, \Sigma_T) \\ \mu_T &= E_T(y_{T+1}) = \phi_0 + \phi_1 y_T +\theta_1 \epsilon_T \\ \Sigma_T &= var_T(y_{T+1}) = \sigma^2. \end{align}

In practice, you will substitute the MLE estimates for the parameter values. Given the known distribution, you can build prediction interval. In other words, you will neglect the uncertainty due to the fact that parameters are estimated and not known. In other words, the prediction interval isn't a function of sample size, although in practice the fact that you rely on an asymptotic argument to replace parameter values with their MLE estimates does mean that it would bounce around if you added or subtracted observations from your sample.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.