An AI strategy agent can produce a valid backtest while using information that did not exist when its trades were supposed to happen. The code runs. The equity curve looks plausible. The signal may even be a sensible one. Yet a timestamp, a join, or a revised data field has quietly given it tomorrow's answer.
The fix is to make availability time part of the research contract. Every input needs a clear answer to two questions: when does this value describe, and when could the strategy first have known it?
What does lookahead bias look like in an AI-generated strategy?
The obvious version is a signal calculated from the same bar's close, followed by a fill at that close. If the strategy needs the closing price to decide, it cannot also trade at that price. But agents often create subtler versions while joining data sources or choosing convenient defaults.
Suppose a model calculates a 20-bar moving average at the close of each minute and takes a position when the close crosses it. If the backtest fills at that same close, it has used the last trade of the bar before that trade was available to the strategy. Shifting the order to the next bar's open may be a reasonable approximation, though a market order still needs fees and impact.
Now make the feature a daily equity fundamental or a crypto open-interest snapshot. The row's date might say Monday, but the value may have been published after Monday's close, or corrected later. A date is not an availability timestamp.
How should I timestamp a strategy's inputs?
Keep at least three times where the source allows it: the period a value describes, the time the publisher released it, and the time your system received it. The strategy's information set at decision time can include only values available by then.
| Field | What it answers | Typical trap |
|---|---|---|
| Event time | When did the market event happen? | Using a bar's close before the bar is complete |
| Release time | When did the source publish this value? | Treating an end-of-day label as an opening-time release |
| Ingest time | When could this research system have consumed it? | Ignoring vendor or pipeline delay |
| Revision time | When was this version recorded or corrected? | Backfilling revised history as though it were original |
For a 1-minute strategy, a one-second delay is not automatically harmless. Whether it matters depends on when the decision is made and what the signal uses. If the input is a completed hourly statistic, it may barely matter. If it is a book imbalance sampled near the order, it can reverse the trade.
Can a point-in-time data store prevent leakage?
It helps, if “point-in-time” means you can retrieve the value known at a historical decision time, including its then-current revision. A table that merely has historical dates can still contain today's corrected values for those dates.
For each record, preserve an effective interval and an availability timestamp, and keep revisions instead of overwriting them. Then make historical queries explicit: return the latest version available as of the simulated decision time. This is especially important for fundamentals, index membership, economic releases, and vendor-cleaned datasets.
There is an unglamorous operational wrinkle: a perfect release timestamp is useless if the ingestion job ran twenty minutes late. If the historical store does not capture ingest time, use a conservative delay and say so. Precision that the source never recorded is just decoration.
What checks catch lookahead bias before paper trading?
Ask the research agent to emit a feature and order timeline alongside the backtest. For each decision, record the latest source availability time for every feature, the decision time, the order time, and the modeled fill time. Reject any row where an input arrived after the decision.
- Shift signals forward by one bar and compare results. A large collapse can reveal close-to-close timing dependence, though it is a diagnostic rather than proof of leakage.
- Truncate every source at a historical cutoff, rerun the pipeline, and compare the resulting features with the stored historical features.
- Replace a suspected feature with a constant. If performance remains nearly identical, inspect whether the code actually used the intended series.
- Run a deliberately impossible future feature through the pipeline. The validation should fail loudly when its availability time exceeds the decision time.
These checks do not certify a strategy. They make specific timing assumptions visible and catch common ways those assumptions get violated.
Does paper trading prove the backtest had no leakage?
No. Paper trading can expose a live data path that is late, missing, or differently aligned from the historical one. It cannot establish that old training features reflected what was known at the time. A model can also stop benefiting from leakage simply because the future it accidentally saw is now the present.
Use paper trading as a parity check: compare the live feature values, decision timestamps, order generation, and modeled fills with the backtest's definitions. When they differ, trace the exact input and clock. An agent team that can explain each decision's information set is doing useful research. An agent team that can only show you a smooth curve has skipped the hardest audit.
← All posts


