Skip to content
All library documents

Why Simulated GARCH Returns May Be More Volatile Than Historical Returns

Article Quant Q&A · Author: ayamathss1

Summary

The document compares historical daily log-return standard deviations for NASDAQ, gold, and silver with average standard deviations measured across simulated GARCH(1,1) paths. The simulations use rugarch, fit separate models with a zero-mean specification, and generate many paths over a 200-step horizon. The reported simulated path volatilities exceed the corresponding historical estimates, and the discrepancy persists across the initialization and innovation choices described.

This is a useful diagnostic question about what a fitted conditional-volatility model implies for future paths and how path-level standard deviations should be compared with historical sample estimates. The document gives example figures and code, but no explanation or resolution. It also does not provide the dataset or fitted parameters in the text, so the cause cannot be established from the reported comparison alone; the exponential conversion shown is not needed to compare standard deviations of log returns.

Key ideas

  • The author compares historical log-return dispersion with dispersion measured across simulated GARCH paths.
  • The stated simulated path standard deviations are higher for all three series.
  • The discrepancy is reported under multiple initialization and innovation settings.
  • The document poses the difference as an open diagnostic question and supplies no confirmed explanation.

Tags

Full text
# standard deviations of the dataset


# GARCH parameters generating simulations that have a lot higher standard deviation's than the historical standard deviation












Below is code using the rugarch package for the daily returns of the NDX, XAU and XAGm. The issue is that when I simulate the data 200 steps ahead for 50k sims, calculate the standard deviation of the returns (and take the average for each sim) it's way higher than the standard deviation of log-returns of the original. This happens no matter if I use `startMethod="sample"` or `startMethod="unconditional"` or use student-t innovations or normal.

I understand there will be a discrepancy, but the discrepancy is quite large. For the below example output:

```
# standard deviations of the dataset
 NASDAQ         XAU         XAG 
0.011372519 0.008126473 0.016545057 

# standard deviations of the simulated paths
Series: NASDAQ Mean Std Dev of Simulated Paths: 0.01497786 

Series: XAU Mean Std Dev of Simulated Paths: 0.01182768 

Series: XAG Mean Std Dev of Simulated Paths: 0.02394441
```

NDX's simulated path standard dev is 0.01497786, whilst the data set is only 0.011372519, which, when taking their exponential's: $$\sigma_{sample} = 1 - e^{0.011372519}=0.011437 = 1.14\%$$ $$\sigma_{sims} = 1 - e^{0.01497786}=0.01509059= 1.51\%$$

Which is obviously a massive difference on the daily. Does anyone know why this is?

csv data set: https://github.com/mathmathqq/garch_data

My code:

```
library(rugarch)

# Read and preprocess data
data <- read.csv("mc_close_price.csv")

# Remove unnecessary column
data$UTM <- NULL

# Calculate log returns for each column
log_returns <- apply(data, 2, function(x) diff(log(x)))

# Calculate standard deviation for each column
std_log_returns <- apply(log_returns, 2, sd, na.rm = TRUE)

# Convert to a data frame for easier handling
log_returns <- as.data.frame(log_returns)

# Define the GARCH(1,1) specification
uspec <- ugarchspec(
  variance.model = list(model = "sGARCH", garchOrder = c(1, 1)),
  mean.model = list(armaOrder = c(0, 0), include.mean = FALSE),
  distribution.model = "norm"
)

# Fit the GARCH(1,1) model for each series individually
garch_fits <- lapply(log_returns, function(series) {
  tryCatch({
    ugarchfit(spec = uspec, data = series)
  }, error = function(e) {
    message("Error fitting GARCH: ", e)
    NULL
  })
})

# Simulate 50,000 paths for 200 steps into the future
simulations <- lapply(garch_fits, function(fit) {
  if (!is.null(fit)) {
    tryCatch({
      # Simulate
      ugarchsim(fit, n.sim = 200, m.sim = 50000, startMethod = "unconditional")
    }, error = function(e) {
      message("Error in simulation: ", e)
      NULL
    })
  } else {
    NULL
  }
})

# Access simulated paths for each series
simulated_paths <- lapply(simulations, function(sim) {
  if (!is.null(sim)) {
    fitted(sim)  # Extract simulated series
  } else {
    NULL
  }
})

# Calculate standard deviation of returns for each path and the mean of the standard deviations
mean_std_dev <- lapply(simulated_paths, function(paths) {
  if (!is.null(paths)) {
    # Calculate log returns for each path
    path_returns <- apply(paths, 2, diff)
    
    # Calculate standard deviation of returns for each path
    path_std_dev <- apply(path_returns, 2, sd, na.rm = TRUE)
    
    # Calculate mean of the standard deviations
    mean(path_std_dev, na.rm = TRUE)
  } else {
    NA  # Return NA if the simulation failed
  }
})

print(std_log_returns)
# Print mean standard deviations for each series
for (i in seq_along(mean_std_dev)) {
  cat("\nSeries:", colnames(log_returns)[i], "Mean Std Dev of Simulated Paths:", mean_std_dev[[i]], "\n")
}
```

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.