Quantile Regression for Execution Costs in Sparse Order Books
Summary
The document considers how to estimate bid and ask execution prices at different order volumes when historical order books are sparse. Its proposed approach separates mid-price from spread and fits piecewise linear quantile regressions to spread changes across volumes. The example targets a high spread quantile, enforces monotonicity across volume levels, and allows exogenous variables. Historical order book snapshots, with some time slippage, provide the calibration data.
The question highlights gaps that matter for practical execution estimates: a requested volume may not be displayed, hidden liquidity can alter fills, and all-or-nothing orders may leave smaller sizes unavailable while larger ones are present. The proposed regression does not represent these cases. The author seeks rough probability-based estimates without the complexity of a queueing model, but the document contains no answer, backtest, or evidence that the proposed approach works. It is therefore a problem statement and modeling proposal, not a validated execution method.
Key ideas
- Piecewise linear quantile regression can estimate spread costs across order sizes.
- Monotonicity constraints encode the expectation that larger volumes should not improve displayed execution prices.
- Sparse books require modeling the chance that a requested volume is unavailable.
- Hidden liquidity and all-or-nothing orders are identified as important effects absent from the proposed model.
Tags
Full text
# Model suggestions price per volume sparse order book # Model suggestions price per volume sparse order book TLDR Any suggestions for industrial methods to model spreads per volume of sparse order books? Full question I am trying to directly model execution spreads per volume (e.g., 1, 2, 3, ... stocks) of a subsequent snapshot of an order book. I am decoupling price into mid-price and spreads, looking only at spreads here. The order book is, however; sparse, such that the modelled volume range is not always present (on both bid/ask sides). Full historic order book data (considering some time slippage) is attainable though. In particular, I am interested in industrial methods avoiding modelling full order book mechanics through queuing-like models. Essentially, I want to model bid/ask prices for a range of volumes, such that I can assume the order is filled directly (or within some time interval) with a certain probability. Furthermore, I am interested if such a model can also incorporate the following dynamics: - Missing volumes (modelled spreads have probability of missing) - Invisible liquidity (in terms of iceberg orders, but also market responses) - All-or-nothing (AON) orders (allows higher volumes to be present, while smaller ones are missing) My current, approach is a piece-wise linear quantile regression model. This models the, e.g., 95% quantile of the spread difference between each volume through linear regression forcing monotonicity and allowing the inclusion of exogenous variables. This model is calibrated on the past snapshots of the limit order book data within some time interval. Unfortunately, this approach ignores missing volumes (in general or due to AON orders) and invisible liquidity. Finally, I am not looking for approaches with same complexity, performance or accuracy as modelling full order book mechanics; however, if possible, would like to have rough approximations of all above features to be included (in some way). I am interested in discussing any comments, thoughts, questions or suggestions!
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.