Skip to content
All library documents

Estimating Daily ETF Bid-Ask Spreads from OHLC Data

Article Quant Q&A · Author: zetrherae

Summary

The document asks how to use the Edge estimator, which derives bid-ask spreads from open, high, low, and close prices, to approximate daily closing spreads for ETFs. It contrasts fitting one estimate across a long history, using expanding historical data, and estimating from intraday bars. The answer frames this as inferring a high-frequency trading cost from daily measurements and notes that the estimator is intended for in-sample estimation rather than forecasting.

The response says the paper's estimates need roughly a month of daily observations and reports that reliability varies with liquidity. Since spreads also depend on market volatility and vary within the trading day, it recommends comparing the three approaches while checking how liquidity, volatility, and intraday seasonality affect results. Its initial suggestion is a rolling window combining intraday bars with daily OHLC, so the estimate reflects recent conditions while retaining information from opening and closing auctions. This is guidance to investigate rather than a demonstrated accuracy result; the document does not establish that the proposed window reproduces broker quotes.

Key ideas

  • The Edge estimator infers bid-ask spreads from OHLC observations and is described as an in-sample estimator.
  • A long-history estimate, an expanding estimate, and an intraday estimate answer different questions and should be compared.
  • Spread estimates vary with liquidity, volatility, and intraday trading patterns.
  • A rolling window of recent intraday bars combined with daily OHLC is suggested as a practical starting point.
  • Intraday bars may omit opening and closing auction information, limiting their usefulness for estimating closing spreads.

Tags

Full text
# Ho to use "Edge" bid-ask estimator correctly for daily bid-ask spreads?


# Ho to use "Edge" bid-ask estimator correctly for daily bid-ask spreads?












I found that you could use the bidask Python module, which is based on the paper "Efficient estimation of bid–ask spreads from open, high, low, and close prices" by Ardia et al (2024), to estimate bid-ask spreads from OHLC prices. If I understand the paper correctly, the model can be used for any time series interval to estimate the average bid-ask spread within that time-series.

Now my goal is to use this model to approximate daily close bid-ask spreads of etfs. This means I want to add the daily close bid-ask spread to the OHLC prices of for example the S&P 500, and this should be as close as possible to actual broker quotes. Now I'm wondering what kind of time series I should use for this task. I figure the main options are:

- I could pass the complete daily time-series from 2000 to today. This gives me a single value and I could use a constant bid-ask spread for any day

- I could pass the daily time-series from 2000 until day i, to get the close spread for day i

- I could pass a intraday timeseries with minute or hour interval for each day i, to get the spread for day i

Could anyone tell me what will give me the most realistic result?

## Answer by lehalle (score 2)

https://quant.stackexchange.com/a/83628

Your question is indeed very generic and goes beyond the estimation of bid-ask spread with daily data.

Let me reformulate it as:

> If I have an estimate of an high-frequency variable using daily measurements, how can I use it to estimate an high-frequency statistic?

In your case

- the bid-ask spread is the high-frequency variable,

- the OHLC are the daily measurements,

- the "broker quote" is the high-frequency statistic.

First a remark: I am not sure what you mean by "broker quote" but I guess it is the markup you pay by trading a decent quantity using a mixed execution strategy (i.e. you don't know in advance which fraction of the desired quantity will be traded at open and close auction and in intraday). In any case you assume that this markup will be a fraction of the bid-ask spread.

You may have seen that in the paper you cite (Ardia, David, Emanuele Guidotti, and Tim A. Kroencke. "Efficient estimation of bid–ask spreads from open, high, low, and close prices" Journal of Financial Economics 161 (2024): 103916):

- Their "unit of computation" is one month. It means that for the authors, you need at least 20 daily observations to have a not very bad estimate of the bid-ask spread.

- The reliability of their estimate is between 17% (for liquid stocks) and 75% (for illiquid stocks).

Your proposals are the following

- use the estimator from all the dates you have 2000 to 2025, and apply it in sample for any of these days,

- use the estimator on data from 2000 (the first available date for you) and date $d$ and use it on date $d$, this is a progressive estimate,

- use the estimator on intraday bars (with half an hour bars you can observe around 16 bars a day) and use it for this day.

My answer has three premisses:

- it is clear in the paper that the proposed estimator is nowcasting and not forecasting the bid-ask spread: it is an in-sample estimate.

- we know that the bid-ask spread is a function of the volatility of the market (cf Dayri, K. and Rosenbaum, M., 2015. Large tick assets: implicit spread and optimal tick size Market Microstructure and Liquidity, 1(01), p.1550003).

- we know that there is an intraday seasonality fo the bid-ask spread (cf. L, C-A, and Sophie Laruelle. Market microstructure in practice. World Scientific, 2018).

First my recommendation is to compare your 3 proposals having in mind my remarks: how different they are (to be compared to the 17% to 75% intrinsic accuracy)? what is the influence of liquidity, and of volatility on your estimates? can you recover the bid-ask spread intraday seasonality with your third proposal?

My prior advice would be to use a sliding window, using around 15 days of intraday bars plus the OHLC of the day (because intraday bars will not get the Open and Close auctions, would they?). For instance +/-7 days of your date. It would allow you to use around $15\times 8$ data to estimate one bid-ask spread, and would capture the "volatility ambiance".

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.