September 26, 2026 · research

We Tried to Reproduce a Strategy’s Backtest. Here’s Where the Numbers Drifted.

We Tried to Reproduce a Strategy’s Backtest. Here’s Where the Numbers Drifted.

We reran a saved backtest and got a different equity curve. Same strategy code, same date range, same symbol. The ending balance was off by 1.8%, and three trades had moved by a bar.

That’s enough to make a research result hard to trust. If a teammate, a future version of you, or a paper-trading service can’t reproduce the run, you can’t tell whether a change improved the strategy or just changed the experiment. Here’s the sequence we followed to find the drift.

Day 1: We wrote down what “the same run” meant

Our first mistake was treating a strategy file as the experiment. It wasn’t. The run also depended on the input data, the engine version, the calendar, instrument metadata and execution settings. The code only described one part of the calculation.

We made a run manifest before changing anything. It recorded the strategy commit, data snapshot identifiers, date range, venue, fee schedule, funding source, fill model and software versions. We also saved the resulting orders and fills, because an equity curve alone can’t show where two runs first diverged.

ArtifactWhat to recordWhy it matters
Market dataSnapshot ID, schema version, adjustmentsVendors repair history and revise corporate actions
ExecutionFee tier, funding series, fill and impact settingsDefaults and account assumptions change results
RuntimeCode commit, engine and dependency versionsLibraries can change ordering, rounding or indicators
OutputOrders, fills, positions and metricsShows where the runs begin to disagree

Day 2: We compared the trades, not the Sharpe

The summary metrics were a distraction. Both runs had nearly the same Sharpe, but their fill logs showed the first mismatch on a funding settlement. One run charged the rate to the position open at the settlement timestamp; the other used the position after that timestamp’s rebalance.

The strategy code hadn’t changed. The engine’s event ordering had. A small version update had made the sequence explicit where it used to depend on how two events happened to sort.

We fixed the run contract to state the ordering: apply funding to the position carried into settlement, then process strategy decisions for that timestamp. The exact convention can vary by venue and engine. Leaving it implicit is the bug.

Day 3: A “same” data file turned out to be different

After pinning event order, the remaining mismatches were concentrated in a few equity trades. The vendor had corrected a historical split adjustment. Our file had the same name and row count as before, which had made it look unchanged.

Now we fingerprint each immutable data snapshot and keep the adjustment policy alongside it. A hash tells us whether bytes changed; it doesn’t explain why. So the manifest also carries source, retrieval time and transformation version. For data that gets revised, those details are part of the result.

A reproducible backtest needs an answer to “which version of the past did it see?”

Day 4: We found one quiet default

The last difference was a maker fee set to zero because the strategy’s config omitted the field. A newer engine applied the account’s default fee. That single default changed the marginal trades enough to account for most of the ending-balance gap.

We made economically meaningful settings explicit and had the engine print the resolved configuration into the run record. Defaults are convenient while exploring. They’re poor evidence when you’re comparing results across time.

3sources of drift found
1.8%initial ending-balance difference
0value in matching Sharpe alone

What we’d skip next time

We spent half a day comparing aggregate metrics before looking at the first differing fill. Don’t start there. Sort both event logs by timestamp and compare the first divergence; later differences often follow from that one cause.

We’d also skip the idea that a container image makes a run reproducible by itself. It pins much of the software environment, but not an external data file, a fee schedule fetched at runtime or a vendor’s revised history.

When a backtest changes, preserve both run manifests and logs, then fix one source of drift at a time. The useful result isn’t just a curve you can rerun. It’s a record that explains which data and assumptions produced it, and why the next run might differ.

reproducibilitybacktestingdata engineeringpaper trading
← All posts