Four numbers can tell an uncomfortable story: a walk-forward Sharpe of 1.4, 312 historical trades, 46 paper trades, and a paper-to-backtest fill-price gap of 18 basis points. The Sharpe gets the headline. The 18 basis points may explain why the strategy feels different before there’s enough paper history to judge its performance.
Walk-forward testing asks whether a research process could have selected a strategy using past data and evaluated it on later data. Paper trading asks whether a running system, with its current data, clocks, order logic, and fills, behaves like that tested process. Those are related questions, but the second is not automatically answered by the first.
What does walk-forward validation actually establish?
Say you train or select a strategy on six months of data, then evaluate it on the next month. Roll that window forward and repeat. If each test month stays out of the corresponding selection process, the exercise gives you evidence about performance under a series of historical conditions.
It does not establish that paper trading will reproduce the same decisions or fills. The historical test may use completed bars, a fixed fee schedule, and a simplified market order model. The live paper process consumes arriving data and submits simulated orders through a different path. One can be sound while the other has a mismatch.
That’s why 46 paper trades versus 312 in the walk-forward sample is not, by itself, a verdict. Check the comparison window, signal frequency, open positions at the boundaries, and whether both counts use the same definition of a trade. A six-week paper run cannot match a year of rolling tests on trade count.
Why is the fill gap so revealing?
An 18-basis-point average gap deserves to be unpacked before interpreting a paper equity curve. Is it measured from the backtest’s assumed fill to the paper simulator’s fill, or from decision-time midpoint to paper fill? Those measure different things. Record the decision price, order arrival price, simulated fill, fees, and any funding separately.
| Observed difference | First check | Why it matters |
|---|---|---|
| Fewer paper trades | Signal timestamps and eligibility rules | A missing or delayed input can suppress decisions |
| Same trades, worse fills | Order type, spread, impact, and fee tier | The backtest may assume a cheaper execution path |
| Different position sizes | Cash, contract multiplier, and rounding | Small unit mismatches compound across orders |
| Drift after a restart | Recovered position and pending orders | The paper process may resume from a different state |
For example, a backtest that buys at the next bar’s open may look close to a paper market order in a liquid instrument. But if the paper order arrives after a sharp move, the gap includes timing and market movement as well as execution cost. Calling all 18 basis points “slippage” hides the cause.
How do I compare paper trading with a walk-forward test?
Start with decisions, then orders, then fills. For each paper decision, replay the same strategy version over the same input history and compare what it knew at that instant. A useful parity record includes:
- Strategy version and parameters, including the selection window that produced them.
- Input values with both event time and availability time.
- Signal, target position, current position, and proposed order.
- Order constraints, fees, funding, and simulated fill details.
- Cash and position state before and after execution.
If the signals differ, inspect data timing and code versions. If signals agree but orders differ, inspect sizing, venue rules, and portfolio state. If orders agree but fills differ, inspect the execution assumptions. Keep those layers separate; otherwise a single performance number sends you hunting in the wrong place.
When is the drift evidence of a broken strategy?
After the pipeline matches, compare outcomes over a meaningful shared period and inspect the strategy’s intended behavior. A handful of paper trades can reveal a broken timestamp or order path quickly. It cannot reliably tell you that a strategy’s average return has changed. That takes enough observations, and the required number depends on how noisy and clustered its returns are.
Market conditions can change, too. A spread widening or funding turning persistently adverse may hurt the strategy even when every decision matches. That’s a real change in trading conditions, not a validation bug. Keep a record of costs and exposures alongside returns so the distinction is visible.
Walk-forward validation is evidence about a selection process on historical periods. Paper trading is evidence about the running system under observed conditions. When their results diverge, first find the earliest point where their records disagree. The surprising number is often not the Sharpe. It’s the small execution or data difference that makes the two experiments answer different questions.
← All posts


