Choosing Price Data for Realized Volatility Estimation
Summary
The document explains that the right price series for realized volatility depends on the quantity being measured, the sampling frequency, and the estimator. Trade prices, bid or ask quotes, and mid-prices can produce different volatility estimates. Quotes may move without trades and create erratic observations; mid-prices can smooth those quote movements, while traded prices may better represent actual price variation when available. Some markets, including over-the-counter currencies, may lack accessible trade data.
Time aggregation also matters: methods designed for intraday, high-frequency, or daily returns may behave differently. The choice can be constrained by the estimator itself; for example, Garman–Klass uses open, high, low, and close prices. The discussion gives selection principles rather than a single preferred input or empirical comparison. It does not resolve the choice for particular time windows or prove that one price measure better satisfies a Brownian-motion assumption.
Key ideas
- Choose the price series according to the volatility concept being measured.
- Trade, quote, and mid-prices can produce different estimates and have distinct limitations.
- The suitable sampling frequency depends on the target horizon and volatility estimator.
- Some estimators prescribe their inputs, as Garman–Klass does with open, high, low, and close prices.
Tags
Full text
# Which prices to use to compute realized volatility? # Which prices to use to compute realized volatility? For computation of realized volatility, especially range based volatility, deal prices are commonly used. If Level I data available should the deals data still be used or another measures of spot price would be preferable(for example, mid-price)? Such measures could diminish the impact of bid-ask spread, but what would be the consequences for the volatility measures? Would the assumptions of the process be violated or quite the opposite - the price process would become closer to theoretical one? Assuming for the ranged based models, price follows Brownian motion with zero drift: $dp_t = \sigma dW_t$, where $p = ln(P)$. What would the better price measure for them? Does the answer change if we use different time windows( 1, 5, 30, 60 minutes)? ## Answer by Matt Wolf (score 1) https://quant.stackexchange.com/a/8071 Which realized volatility are you attempting to measure is highly important in order to determine which prices and return series to utilize to compute realized volatility. Here couple ideas: - What do you attempt to measure: Bid/Offer spread volatility, traded price variations,...Even if you attempt to measure asset price variations it can make a difference whether you use bid, offer, mid-prices, or traded prices. Using bids or offers for this particular purpose can sometimes result in erratic moves (because prices may fluctuate widely even though nobody trades on those bids and offers), mid prices somewhat smooth that out, and traded prices are maybe the most preferable to use in this regards. However, there are asset classes where you do not easily get a hold of traded prices, such as OTC currency price series. So, depending exactly what you attempt to measure will help to narrow down which price series to use. - Which time compression do you target with your volatility measure? Are you dealing with tick based data, with compressed intraday data, daily, weekly data. It can make a difference because there are realized volatility models out there that shine on capturing intraday variations, while others are better in measuring high frequency price return variations or daily return volatility. - What specific volatility measure are you targeting. Garman-Klass, for example specifies exactly which price series to use (such as Open, High, Low, Close). Combined, you should be able to exactly determine which price series to use.
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.