September 27, 2026 · research

The Backtest Passed. The Paper Strategy Missed 37% of Its Orders.

The Backtest Passed. The Paper Strategy Missed 37% of Its Orders.

Our backtest counted 91 fills from 100 submitted orders. In paper trading, the same strategy got 63. The median time from signal to order acknowledgement was 84 milliseconds, and 12 of the missed paper fills were limits the backtest assumed had filled as soon as price touched them.

That 28-point gap looked like a strategy problem until we separated order decisions from fill decisions. The strategy chose almost the same side, size and price in both runs. The execution path differed: some orders arrived late, some remained open, and some crossed a market that had already moved.

What does backtest-to-paper order parity mean?

Parity means the backtest and paper system make comparable decisions from the same available information, then account for their different execution models explicitly. It does not mean every simulated fill should match a paper fill. Historical bars cannot reconstruct queue position, and a paper venue may use a different matching model from a live exchange.

The useful question is narrower: for each order the strategy intended to place, can you explain what happened next? A trade count and a final return conceal too much. Keep a joined record for each decision, with a stable strategy decision ID carried through signal, order, acknowledgement, cancellation and fill events.

ComparisonWhat it catchesExample
Decision time and sideDifferent inputs or schedulingPaper signal fires one bar later
Requested size and priceRounding, risk or venue rulesBacktest submits 0.013 BTC; paper rounds to 0.01
Order state transitionsSubmission, rejection and cancellation gapsBacktest treats a cancel request as completed
Fill quantity and priceFill model optimism or market movementA touched limit gets a full backtest fill, but no paper fill

Why did the touched limits account for so many misses?

Our backtest used one-minute bars. If a limit price fell within a bar's high-low range, it marked the order filled. That rule answers whether the market traded at that price at some point during the minute. It does not answer whether our order was active then, whether it was first in the queue, or whether enough quantity traded after it arrived.

Paper logs exposed the timing problem. The strategy calculated a signal at 12:03:00.000, but the market-data event reached the order process 31 milliseconds later. Risk checks took another 22 milliseconds, and the venue simulator acknowledged the order 31 milliseconds after that. In a fast move, a historical bar can make a price look reachable even though the limit arrived after the market had left it.

There was another wrinkle: twelve paper orders were still open when the backtest had already moved on. The simulator accepted a cancellation request as though the order had vanished immediately. In the paper system, the cancel acknowledgement came later; three orders filled during that interval. That changed the position the next signal saw.

How do I measure the gap without fooling myself?

Start with a small reconciliation report grouped by order intent. Keep the denominator visible. “Fill rate” can mean fills per submitted order, filled quantity divided by requested quantity, or filled orders divided by orders that reached the venue. Those answer different questions.

  1. Join events using a decision or client order ID, not a timestamp guessed after the fact.
  2. Compare decisions first: signal time, side, requested quantity, order type and limit price.
  3. For matching decisions, compare acknowledgement delay, rejections, open time, cancel time, filled quantity and volume-weighted fill price.
  4. Report rates by order type and market condition. A 63% overall fill rate may hide 90% for market orders and 35% for passive limits.

Keep the categories mutually clear. A rejected order is not an unfilled order; a partially filled order is not a full fill; and an order canceled after a partial fill still changed the position. Count both orders and requested quantity so a cluster of tiny orders cannot make the result look healthier than it is.

91%backtest order fill rate
63%paper order fill rate
28 ppdifference to investigate

What should change in the backtest?

Use a fill rule that matches the data resolution and intended order behavior. With bars, a touched limit is evidence of a possible fill, not proof of one. You can model partial fills conservatively, require a trade-through, or exclude passive orders whose queue cannot be estimated. Each choice answers a different question; document it and compare the result under more than one plausible rule.

Also model order state. An order remains live until the system receives a terminal event, and a cancel request does not erase exposure. If the strategy can submit a second order while the first is pending, its backtest needs to represent that race or prevent it by design.

A fill model is a claim about what your order could have done. Give it a rule you can explain, then check that rule against paper logs.

For our 28-point gap, the surprising part was not that a minute-bar backtest overstated passive fills. It was that the strategy's position diverged before its next decision because cancellation timing had been omitted. Once we fixed the order-state model and stopped treating every touched limit as a fill, the paper comparison got less flattering and more useful.

That is the practical bar for parity: every meaningful difference between simulated and paper behavior has a recorded cause, and the backtest does not claim certainty where its data cannot provide it.

paper tradingorder managementbacktestingexecutiontrading systems
← All posts