Approaches to Forecasting Volatility from Five-Minute Data
Summary
The discussion addresses volatility forecasting for financial stocks observed at five-minute intervals and notes that standard ARCH and GARCH approaches may not work as well at this frequency. One response reports that, in a thesis and related literature, nonparametric methods such as support vector regression and random forests outperformed classic econometric specifications. Nonlinear relationships and fat-tailed data are offered as possible reasons, while the suggested practical step is to experiment with lag choices.
Another response recommends beginning with an autoregressive model, then adding variables such as trading volume and liquidity measures in a vector autoregression. Such inputs may capture information that reduces the number of lags needed. It also mentions LSTM networks, which were found more useful for volume than volatility in that experience. These are recommendations and reported findings, not a controlled comparison supplied in the document. Performance may depend on the dataset, forecast target, features, and validation design; no specific best model or Python implementation is established.
Key ideas
- Five-minute volatility forecasting may require approaches beyond daily ARCH and GARCH specifications.
- One answer reports better performance from support vector regression and random forests in its research context.
- An autoregressive baseline can be expanded with volume and liquidity variables in a vector autoregression.
- Additional predictors may reduce the number of useful lags.
- The cited LSTM experience was more favorable for volume than for volatility, and the discussion does not establish a universal best model.
Tags
Full text
# Volatility forecast for 5-minute frequency data # Volatility forecast for 5-minute frequency data I have high frequency data for financial stocks (5-minute periodicity) and I want to forecast volatility. I'm familiarized with the usual ARCH/GARCH models and their variants for daily data but after doing some research I've learnt that these models don't work well with high frequency data. Which model is best for volatility forecasting when I have one data point every 5 minutes? Are there any known Python implementations of that model? ## Answer by Trader2B (score 1) https://quant.stackexchange.com/a/71744 I did my MSc thesis on this topic and found nonparametric methods such as SVR and RF outperform classic econometric specifications (ARCH, GARCH, EGARCH ect.). Similar results are found in the literature and nonlinear relations as well as fat tails are often cited as possible culprits. Might be worth it to take a look at what happens when you plug your desired lags in a model like that! ## Answer by lehalle (score 0) https://quant.stackexchange.com/a/71768 We study exactly this in Endogenous Dynamics of Intraday Liquidity, by Bińkowski and L (the preprint is there). My recommendation is to start with a AR (auto-regressive) model, and then to expend it. You should see that is you introduce other variables in a VAR (vectorial auto-regressive) version of this model, the needed lags will decrease, like if the recent traded volumes (and other liquidity variables) contain information on the less recent volatility. In the past I used LSTM (a neural net version of ARMA) too, that was more useful for volumes than for volatility (especially if one want to learn the correct lags to be used).
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.