Skip to content
All library documents

Machine Learning Methods for Multi-Step Volatility Forecasting

Article Quant Q&A · Author: develarist

Summary

The document asks how machine-learning methods such as random forests, support vector regression, gradient boosting, and nearest neighbors compare with GARCH when forecasting volatility beyond a single step. It also asks whether these models use lagged volatility in an autoregressive way and what theoretical rationale supports applying machine learning to volatility series. The replies do not provide a survey or quantitative head-to-head evidence.

Instead, they emphasize that model performance depends on the setting and available features. Historical returns at several frequencies and derived measures of volatility or mean reversion can be supplied to classification or regression models. One respondent is skeptical of K-nearest neighbors and random forests, arguing that tree partitions may miss smooth relationships, while another points to recurrent neural networks as possible candidates. These are opinions rather than established comparative findings; the discussion gives no testing design, forecast horizon results, or general performance ranking.

Key ideas

  • The exchange does not establish a reliable ranking of machine-learning models against GARCH.
  • Forecast performance depends on the prediction setting and the features supplied.
  • Lagged returns and derived volatility measures can serve as model inputs.
  • Skepticism about nearest-neighbor and random-forest models is presented as opinion, not empirical evidence.

Tags

Full text
# Forecasting volatility farther ahead with autoregressive machine learning


# Forecasting volatility farther ahead with autoregressive machine learning












ARIMA and GARCH are old news for predicting volatility time series of asset returns. I am aware of papers that replace ARIMA and GARCH with machine learning algorithms to predict financial volatility more accurately, and so this question is a reference request for a survey of what's out there:

How do the performances of the following, and other, machine learning algorithms compare with one another, and GARCH, at forecasting volatility for horizons longer than just 1-day/step ahead?

- random forest

- support vector regression (SVR)

- gradient boosting

- K nearest neighbors, etc)

And are the machine learners also implemented in an autoregressive formula like GARCH (i.e. they use preceding historical volatility observations to estimate the current volatility)?

Also what do the research articles say regarding the theoretical reason for applying machine learning to volatility time series in particular?

## Answer by confused (score 3)

https://quant.stackexchange.com/a/57069

Just based on my understanding of the ML models themselves, I have a hard time believing KNN or RF are useful in anyway. They wouldn't be the first models I try and tend to just be ML models taught in class for who knows what reason honestly - maybe because they are easy to understand? From what I have read about ML in general (not in relation to time series), all of the ones you have listed have been outclassed by neural networks. Gradient boosting might be one that is still somewhat useful.

KNN looks to predict a value based on K observations that are most similar and then takes the average. Do you think tomorrow's volatility really is equal to the most similar days in your data set, even if the days are from 3 years ago? If so KNN, may be helpful.

RF is just a less good version of Gradient Boosting Trees. It predicts based on thresholds of features and as a result just partitions your data and then predicts based on the average of the average of many trees. So tomorrows volatility is equal to days when yesterday's vol is greater than x but less than y, snp moved by more than z but less than a, etc... Does that make sense? Maybe, but due to the partitioning nature of RF, it can never truly replicate any mathematical function. Meaning, if the true relationship between x and y is linear, a linear regression will always do better than a random forest.

## Answer by Andreas (score 2)

https://quant.stackexchange.com/a/53307

This largely depends on your setting and the available features.

You can include further information into classification or regression algorithms by providing the model with additional features such as the daily, weekly, monthly returns of previous periods and eventually also use these to create more features such as measures of volatility, mean-reversion or other aspects you would include in a "traditional" approach.

Benefits of machine learning algorithms (especially neural networks such as LSTMs or other RNNs) are, that they tend to be really fast and still offer a comparably good performance to many sophisticated option pricing models.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.