September 15, 2026 · research

Your Backtest’s Slippage Model Should Be Wrong on Purpose

Your Backtest’s Slippage Model Should Be Wrong on Purpose

The most honest slippage model is often deliberately wrong. A backtest that reports one precise execution cost from coarse historical data is making a claim the data can’t support.

That sounds like an argument for crude assumptions. It’s the opposite: use a simple model you can explain, then make it disagree with itself. If a strategy only works under one conveniently favorable fill estimate, the uncertainty is part of the result.

Why can a precise slippage estimate be misleading?

Suppose you have one-minute OHLCV bars and send a market order for 2,000 shares. The bar tells you the range and total volume over a minute. It doesn’t tell you the bid and ask when your order arrived, the queue ahead of you, how volume was distributed through the minute, or how much of it your order would have consumed.

You can still assign a number to slippage. You might fill at the next bar’s open plus a fixed basis-point cost, or scale the cost with order size divided by bar volume. Those are useful approximations. They are not observations. Reporting the result as “slippage: 3.2 basis points” makes a model choice look like a measurement.

For a worked example, let a strategy trade $20,000 per order in a stock whose next-bar volume is $400,000. Its participation rate is 5%. A model charging 2 basis points costs $4 per order; a model charging 10 basis points costs $20. Across 500 round trips, that gap is $16,000 before commissions. If the strategy’s backtested gross profit is $18,000, the execution assumption has nearly as much influence as the signal.

5%order size as a share of next-bar volume
$16,000cost difference across 500 round trips at 2 vs. 10 bps

Build a range that reflects what your data can see

Start with a baseline that matches the strategy’s order type and the resolution of the data. A market order needs spread and impact. A limit order needs a rule for whether it fills at all; assuming every touched price fills is a different execution error.

Then run the same backtest under a small set of explicit cases. Keep the signal, sizing, and timestamps fixed so you can attribute changes to execution:

CaseExecution assumptionWhat it probes
FavorableBaseline spread and low impactHow much does the strategy depend on benign trading?
BaseTypical spread and impact scaled by participationWhat does the chosen working assumption imply?
AdverseWider spread, higher impact, and delayed fillDoes the edge survive a plausible difficult market?

These aren’t confidence intervals unless you’ve calibrated them statistically. They’re scenario tests. Say what they mean in the report, and don’t average them into a made-up “expected slippage” without evidence for the probabilities.

Make the model respond to the order

A flat basis-point charge can be a reasonable first pass for small orders. It becomes a poor story as participation grows. At minimum, record order notional, available bar volume, and participation rate. If the strategy trades futures, record contract value and the relevant volume units; if it trades spot crypto, be explicit about quote currency and venue.

A simple impact curve can expose capacity pressure without pretending to reproduce an order book. For example, define cost as a baseline spread charge plus an impact term that grows with the square root of participation. Fit its parameters only if you have suitable execution data. Otherwise, vary the parameters and show how quickly the result deteriorates.

And keep the fill model from silently rescuing the strategy. A long signal that arrives after a sharp rise should not get the bar’s low just because it’s inside the candle. Use a causal fill convention, then stress the price and delay it. If a delayed order misses the move, that’s an execution outcome, not a nuisance to smooth away.

Useful question: at what spread, impact, delay, or participation rate does net performance cross zero? That break-even point is often more actionable than a single backtest Sharpe.

Use paper trading to tighten the range

Paper trading won’t reveal the fill you would have received in a live matching engine. It can still tell you whether the backtest’s order stream resembles the system’s actual decisions: when orders are generated, how large they are, and how long they remain actionable.

Compare the simulated order price with contemporaneous quotes or whatever market data your paper system records. Track the gap by instrument, time of day, order size, and order type. Those observations can narrow assumptions for that setup. They don’t automatically transfer to a different venue, market regime, or larger order.

Critics are right that pessimistic scenarios can be arbitrary. A deliberately harsh model can reject a viable strategy just as a favorable one can flatter a weak one. So don’t crown the harshest case as truth. Publish the assumptions, show the sensitivity, and collect the data needed to replace guesses with calibrated estimates.

A backtest cannot know its historical queue position from a candle. It can show whether the idea still looks coherent when execution gets worse in ways your data cannot rule out. That’s a more useful answer than a decimal place.

slippagemarket impactbacktestingexecutionpaper trading
← All posts