Skip to content
All library documents

Estimating Bid-Ask Spreads and Slippage from OHLC Data

Article Quant Q&A · Author: user49573

Summary

The document outlines ways to estimate trading costs when order-book data is unavailable. For spread estimation, it describes Roll’s approach based on serial covariance of returns, a correction related to microstructure noise, and methods that infer spreads from daily highs, lows, and closes. The Corwin–Schultz method uses one-day and two-day high-low ranges, under assumptions about where trades occur and how volatility scales over time. Abdi–Ranaldo also uses highs, lows, and closes to separate efficient-price variation from spread effects.

These estimates can inform subsequent price-impact models, but OHLCV data alone does not reveal execution details or intraday price paths. The response treats latency as especially difficult to estimate because it depends on infrastructure and system response times, and suggests limiting analysis to comparisons between software or brokers. A second answer emphasizes that slippage depends on order size, trading frequency, and intraday behavior, so simple assumptions may be more appropriate for small, infrequent trades in liquid equities.

Key ideas

  • Return autocovariance can be used to estimate a bid-ask spread under Roll’s model.
  • High-low ranges across adjacent days support spread estimates that separate volatility from spread effects.
  • Using close prices alongside highs and lows offers another way to estimate spread and efficient-price variation.
  • Estimated spreads and volatility can serve as inputs to price-impact models.
  • OHLCV data cannot fully capture execution slippage or latency, which depend on trading and infrastructure details.

Tags

Full text
# Modeling Slippage without Order Book data


# Modeling Slippage without Order Book data












I am building a portfolio simulator and finding ways to make it more 'realistic'. For example, giving the option to reinvest dividends, include capital gain taxes, commission/fees (fixed for now) etc. Regarding slippage/latency, however, I'd like to make a more dynamic model. Have you ever had experience for modeling slippage? For example, as a function of volume and volatility (already embedded in OHLCV data) rather than using the order book for the spread?

Thank you for your guidance :)

## Answer by kurtosis (score 4)

https://quant.stackexchange.com/a/57328

I'm going to assume that by analyzing "slippage" you mean transactions costs.

### Analyzing Latency

First, I will say that analyzing latency issues is incredibly hard. You probably do not even know where your strategy will be located: colocated? not colo but close by? You also do not know how fast your algorithm will respond to a signal: milliseconds? micros? nanos? For example, colocation at the CME of Trading Technologies software on a shared server can sometimes achieve response times in the 100-300 microsecond range. I know other software people have created that responds in the millisecond range.

I would not get too deep on analyzing latency apart from (maybe) comparing different software or brokers.

### Estimating Bid-Ask Spreads

It might seem hopeless to analyze slippage, but that is not so. There are some excellent papers on estimating bid-ask spreads from daily close or OHLCV data.

Roll (1984)

First, you could use Roll's (1984) work on bid-ask spreads and estimate spreads as $\sqrt{-\textrm{cov}(r_t,r_{t-1})}$.

Zhang, Mykland, and Aït-Sahalia (2005)

You could also look at Zhang, Mykland, and Aït-Sahalia's (2005) TSRV work which estimates variances but has to correct for the "microstructure noise pollution" caused by bid-ask bounce. They have a subtractive correction: their adjusted "fast scale" estimator $\frac{\bar{n}_k}{n-\bar{n}_k}\sum_{i=1}^n r_i^2$. You could use that as something similar to the $2c^2$ in Roll's model.

Corwin and Schultz (2012)

Another approach would be to use Corwin and Schultz's (2012) method for estimating volatilities and bid-ask spreads from OHLCV data. Their method is a bit more involved, but has some economic reasoning behind it: they assume high prices are likely executed at the offer and low prices were likely executed at the bid.

They then look at highs and lows for one- and two-day periods. They estimate an average squared daily one-day "log-return" from low to high ($\log(H_t/L_t)$) and the squared two-day "log-return" from the two-day low to high. $$ \begin{align} \hat\beta &= \frac{1}{n/2}\sum_{j=1}^{n/2}\sum_{i=2j-1}^{2j} [\log(H_i/L_i)]^2, \\ \hat\gamma &= \frac{1}{n/2}\sum_{j=1}^{n/2} \left[\log\left(\frac{\max(H_{2j-1},H_{2j})}{\min(L_{2j-1},L_{2j})}\right)\right]^2. \end{align} $$ That lets them solve a system of equations since variance scales linearly with time while the bid-ask spread is assumed to be constant across both days: $$ \begin{align} \beta &= 2k_1\sigma^2 +4k_2 \sigma \alpha + 2\alpha^2, \quad \text{and}\\ \gamma &= 2k_1\sigma^2 +2\sqrt{2}k_2 \sigma \alpha + \alpha^2 \quad \text{where} \\ \alpha &= \log\left(\frac{2+S}{2-S}\right), \quad S = \text{spread}, \\ k_1 &= 4\log(2), ~\text{and} \quad k_2 = \sqrt{\frac{8}{\pi}}. \end{align} $$

Abdi and Ranaldo (2017)

Finally, you could try Abdi and Ranaldo's (2017) method. They assume, like Corwin and Schultz, that highs are at the offer and lows are at the bid. However, they also use close prices and presume there is some efficient price for lows, highs, and close prices $l_t^e, h_t^e, c_t^e$. They then assume the average of the efficient lows and highs $(l_t^e+h_+t^e)/2$ is a fair estimate of the efficient close (albeit with some noise of the efficient price process). Also, they point out the the observed high and low prices may be averaged since the plus-and-minus of a half spread cancels out. Thus $$ \eta_t = \frac{l_t^e + h_t^e}{2} = \frac{l_t + h_t}{2}. $$

They next note that $E(\frac{\eta_t + \eta_{t+1}}{2}) = E(c_t^e)$. Therefore, the variance of $\eta$ changes estimates the efficient price variance $\sigma_e^2$ and the variance of $c_t$ versus the average of $\eta$'s depends on both $\sigma_e^2$ and the spread $S$. That gives a system of equations which is easily solved (since it is already triangular): $$ \begin{align} E[(\eta_{t+1}-\eta_t)^2] &= \left(2-\frac{k_1}{2}\right)\sigma_e^2, \quad \text{and} \\ E\left[\left(c_t-\frac{\eta_t+\eta_{t+1}}{2}\right)^2\right] &= \frac{S^2}{4} + \left(\frac{1}{2} + \frac{k_1}{8}\right) \sigma_e^2 \end{align} $$ where $k_1=4\log(2)$, as in Corwin and Schultz's method.

### Analyzing Slippage

Once you have estimates of bid-ask spreads and volatilities, you can easily try fitting your trading or returns to various price impact models. While I could write up plenty on those, I'll just self-plagiarize and suggest the answer here to guide you on using your spread and volatility estimates.

## Answer by Quantoisseur (score 0)

https://quant.stackexchange.com/a/57321

OHLCV data is not sufficient to estimate slippage as it depends on execution and intraday price action.

How to simulate slippage

Based on the size of your orders and trading frequency you can make some assumptions on the overall impact though. For example, if you are not trading frequently, don't have a huge portfolio, and trading liquid equities, don't worry about it.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.