Skip to content
All library documents

Diagnosing Explosive ARFIMA–EGARCH Forecasts as Sample Size Grows

Article Quant Q&A · Author: alex337d

Summary

The document presents a time-series modeling problem: an ARFIMA(12,0,12) mean with an eGARCH volatility specification produces plausible one-step forecasts on a sample of 2,000 observations, but extreme mean and variance forecasts after fitting the same orders to 6,000 observations. Reducing the ARMA orders removes the explosion, while residual autocorrelation tests then suggest remaining serial dependence.

The reported output includes parameter estimates and weighted Ljung–Box diagnostics for standardized residuals and their squares, but the post does not include an answer or a confirmed cause. It therefore illustrates a model-diagnosis question rather than a resolved method. The output alone does not establish that the larger-sample fit is adequate or identify why filtering becomes unstable. Readers would need to investigate parameter and root stability, optimization and filtering behavior, and model specification before interpreting the forecasts or selecting a lower order.

Key ideas

  • A high-order ARFIMA–eGARCH fit can produce dramatically different forecasts as the sample grows.
  • Lowering the ARMA order removes the reported forecast explosion but leaves residual autocorrelation concerns.
  • Ljung–Box diagnostics are reported for standardized residuals and standardized squared residuals.
  • The document gives no resolution, so it does not identify a validated cause or remedy.

Tags

Full text
# Exploding forecast when increasing sample size in ARIMA in R/Python


# Exploding forecast when increasing sample size in ARIMA in R/Python












I am fitting ARFIMA - eGARCH to time-series log returns in Python using R Rugarch package. The sample size is 2000, and I am doing 1-step ahead forecast. The following is the output

```
 Conditional Variance Dynamics  
-----------------------------------
GARCH Model : eGARCH(2,1)
Mean Model  : ARFIMA(12,0,12)
Distribution    : sstd 

Optimal Parameters
------------------------------------
        Estimate  Std. Error     t value Pr(>|t|)
mu      0.021287    0.003462     6.14839 0.000000
ar1     0.542787    0.000082  6659.75884 0.000000
ar2    -0.392375    0.000064 -6101.43666 0.000000
ar3     0.731225    0.000102  7197.18369 0.000000
ar4    -0.386443    0.000063 -6167.48200 0.000000
ar5     0.150615    0.000037  4116.94426 0.000000
ar6    -0.288939    0.000054 -5348.16765 0.000000
ar7     0.232796    0.000044  5282.09465 0.000000
ar8    -0.809897    0.000113 -7199.04032 0.000000
ar9     0.582303    0.000085  6816.71469 0.000000
ar10   -0.210702    0.000044 -4751.43759 0.000000
ar11    0.795572    0.000111  7190.20153 0.000000
ar12   -0.369857    0.000062 -5948.72174 0.000000
ma1    -0.698025    0.000324 -2154.84576 0.000000
ma2     0.434028    0.000132  3293.28328 0.000000
ma3    -0.792481    0.000242 -3279.96665 0.000000
ma4     0.465008    0.000260  1789.43730 0.000000
ma5    -0.165260    0.000253  -652.30069 0.000000
ma6     0.311199    0.000093  3354.92929 0.000000
ma7    -0.306110    0.000091 -3349.74284 0.000000
ma8     0.844868    0.000145  5830.41599 0.000000
ma9    -0.687115    0.000123 -5607.67847 0.000000
ma10    0.246369    0.000068  3618.79384 0.000000
ma11   -0.797917    0.000128 -6236.95580 0.000000
ma12    0.438962    0.000088  4986.13572 0.000000
omega  -0.009447    0.002659    -3.55262 0.000381
alpha1  0.047739    0.047770     0.99935 0.317627
alpha2 -0.007567    0.047962    -0.15776 0.874643
beta1   0.989998    0.000194  5112.96496 0.000000
gamma1  0.263583    0.064669     4.07591 0.000046
gamma2 -0.160119    0.067340    -2.37776 0.017418
skew    1.035755    0.026110    39.66916 0.000000
shape   3.139950    0.236605    13.27084 0.000000

Weighted Ljung-Box Test on Standardized Residuals
------------------------------------
                          statistic p-value
Lag[1]                        4.243 0.03941
Lag[2*(p+q)+(p+q)-1][71]     35.810 0.62628
Lag[4*(p+q)+(p+q)-1][119]    56.029 0.75254
d.o.f=24
H0 : No serial correlation

Weighted Ljung-Box Test on Standardized Squared Residuals
------------------------------------
                         statistic p-value
Lag[1]                     0.03815  0.8451
Lag[2*(p+q)+(p+q)-1][8]    0.16537  0.9999
Lag[4*(p+q)+(p+q)-1][14]   0.42526  1.0000
d.o.f=3

Adjusted Pearson Goodness-of-Fit Test:
------------------------------------
  group statistic p-value(g-1)
1    20     7.405       0.9917
2    30    13.111       0.9950
3    40    19.097       0.9969
4    50    32.981       0.9615

Mean forecast -0.020318754215676776
Var forecast 1.0401001687924032
```

As you can see the model fits the data pretty well and forecast is reasonable. However, if I increase the sample size to 6000 and fit the same model coefficients I get exploding forecasts:

```
Filter Mean forecast -148845285.88045484
Filter Var forecast 30725154.575601663
```

The problem disappears if I reduce the ARMA order (eg ARIMA (3,0,3)), but then my standardized residuals serially correlated according to Ljung-Box test, indicating there is valuable information in residuals. So my question is why does the forecast explode and how to solve it without hurting the residual autocorrelation?

Thanks!

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.