October 7, 2026 · research

Your Strategy Passed Walk-Forward. Why Did the Paper Results Drift?

Your Strategy Passed Walk-Forward. Why Did the Paper Results Drift?

Four numbers can tell an uncomfortable story: a walk-forward Sharpe of 1.4, 312 historical trades, 46 paper trades, and a paper-to-backtest fill-price gap of 18 basis points. The Sharpe gets the headline. The 18 basis points may explain why the strategy feels different before there’s enough paper history to judge its performance.

Walk-forward testing asks whether a research process could have selected a strategy using past data and evaluated it on later data. Paper trading asks whether a running system, with its current data, clocks, order logic, and fills, behaves like that tested process. Those are related questions, but the second is not automatically answered by the first.

What does walk-forward validation actually establish?

Say you train or select a strategy on six months of data, then evaluate it on the next month. Roll that window forward and repeat. If each test month stays out of the corresponding selection process, the exercise gives you evidence about performance under a series of historical conditions.

It does not establish that paper trading will reproduce the same decisions or fills. The historical test may use completed bars, a fixed fee schedule, and a simplified market order model. The live paper process consumes arriving data and submits simulated orders through a different path. One can be sound while the other has a mismatch.

That’s why 46 paper trades versus 312 in the walk-forward sample is not, by itself, a verdict. Check the comparison window, signal frequency, open positions at the boundaries, and whether both counts use the same definition of a trade. A six-week paper run cannot match a year of rolling tests on trade count.

Why is the fill gap so revealing?

An 18-basis-point average gap deserves to be unpacked before interpreting a paper equity curve. Is it measured from the backtest’s assumed fill to the paper simulator’s fill, or from decision-time midpoint to paper fill? Those measure different things. Record the decision price, order arrival price, simulated fill, fees, and any funding separately.

Observed differenceFirst checkWhy it matters
Fewer paper tradesSignal timestamps and eligibility rulesA missing or delayed input can suppress decisions
Same trades, worse fillsOrder type, spread, impact, and fee tierThe backtest may assume a cheaper execution path
Different position sizesCash, contract multiplier, and roundingSmall unit mismatches compound across orders
Drift after a restartRecovered position and pending ordersThe paper process may resume from a different state

For example, a backtest that buys at the next bar’s open may look close to a paper market order in a liquid instrument. But if the paper order arrives after a sharp move, the gap includes timing and market movement as well as execution cost. Calling all 18 basis points “slippage” hides the cause.

How do I compare paper trading with a walk-forward test?

Start with decisions, then orders, then fills. For each paper decision, replay the same strategy version over the same input history and compare what it knew at that instant. A useful parity record includes:

If the signals differ, inspect data timing and code versions. If signals agree but orders differ, inspect sizing, venue rules, and portfolio state. If orders agree but fills differ, inspect the execution assumptions. Keep those layers separate; otherwise a single performance number sends you hunting in the wrong place.

When is the drift evidence of a broken strategy?

After the pipeline matches, compare outcomes over a meaningful shared period and inspect the strategy’s intended behavior. A handful of paper trades can reveal a broken timestamp or order path quickly. It cannot reliably tell you that a strategy’s average return has changed. That takes enough observations, and the required number depends on how noisy and clustered its returns are.

Market conditions can change, too. A spread widening or funding turning persistently adverse may hurt the strategy even when every decision matches. That’s a real change in trading conditions, not a validation bug. Keep a record of costs and exposures alongside returns so the distinction is visible.

Walk-forward validation is evidence about a selection process on historical periods. Paper trading is evidence about the running system under observed conditions. When their results diverge, first find the earliest point where their records disagree. The surprising number is often not the Sharpe. It’s the small execution or data difference that makes the two experiments answer different questions.

walk-forwardpaper tradingbacktestingexecutionstrategy research
← All posts