Cleaning High-Frequency Data for Volatility Estimation
Summary
The document concerns noisy, incomplete one-minute price and volume data for several stocks, with the goal of forecasting volatility for the month after the sample ends and investigating whether volatility is related to liquidity. The response recommends an R package designed for high-frequency financial data, describing it as a source of procedures for cleaning observations and estimating volatility.
The answer offers a tool-oriented starting point for analysis, but it does not specify which cleaning rules, volatility estimator, forecast design, or liquidity measure to use. It provides no empirical results or comparison of methods, and it does not explain how to test a relationship between liquidity and volatility. The suggested package and tutorials may help a researcher begin, but selecting and validating an approach for this dataset remains unresolved.
Key ideas
- High-frequency price and volume observations may contain missing values, noise, and spikes that need treatment before analysis.
- The response recommends an R package with tools for cleaning high-frequency data and estimating volatility.
- The stated forecasting objective is next-month volatility based on a multi-stock sample.
- The document does not provide a specific estimator, forecasting evaluation, or method for testing liquidity dependence.
Tags
Full text
# Estimating volatility from high frequency price volume data of multiple stocks # Estimating volatility from high frequency price volume data of multiple stocks I have price volume data of five stocks, sampled at 1 minute interval for six months. The data is quite noisy, lots of missing data and also some weired spikes. Can someone suggest me how to clean this dataset? The main objective is to estimate the volatility for the next month following the end of samples, what is the best method for this? Also how do I verify whether the volatility is dependent on market liquidity, based on this data. Can someone point me to easy tutorials/books that allows this kinds of data analysis. I have a good background on statistical data analysis but I am a complete noob to financial data. Any help will be much appreciated ## Answer by Malick (score 1) https://quant.stackexchange.com/a/22615 If you are familiar with programming (which is required to deal with HF data), I would strongly recommend you to use the "HighFrequency" R package. It includes a lot of procedures to clean HF data and to estimate volatility. You can find here a very good tutorial about the package. If needed you can find here some good tutorials for R. Credits: The package has been developed by Kris Boudt , Jonathan Cornelissen , Scott Payseur,Giang Nguyen and Maarten Schermer .
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.