Skip to content
All library documents

Perturbation Tests for Look-Ahead Leakage in Feature Pipelines

Article Quant Q&A · Author: Theo Howard

Summary

The document proposes detecting temporal leakage by changing all input rows after a chosen cutoff and rerunning feature construction. If outputs at or before the cutoff change, the feature pipeline has used future data. This empirical check can reveal alignment errors that static source inspection may miss, such as an unshifted rolling statistic. The discussion reports one futures example in which correcting such an error reduced the model’s AUC from 0.957 to 0.675.

The method depends on deterministic feature construction and a suitable corruption strategy; some changes may preserve the statistic being tested. It also cannot detect problems already embedded in the historical inputs, such as revised records or inaccurate timestamps. The reply recommends data provenance practices, including preserving hashes, marking backfilled records, and preventing backfills from replacing live data. It suggests cross-process and cross-version determinism checks and points to metamorphic and mutation testing as potentially relevant research vocabulary, while offering no specific publication on this exact test.

Key ideas

  • Changing data after a cutoff and comparing earlier feature outputs can expose future dependence in feature construction.
  • The test requires deterministic pipeline behavior, including across process restarts and library versions.
  • The choice of data corruption matters because some perturbations can preserve the feature statistic.
  • Perturbation tests do not detect historical revisions or timestamp errors already present in the inputs.
  • Data provenance, backfill labels, and immutable records can help address upstream data leakage.

Tags

Full text
# Is there existing work on detecting look-ahead bias by perturbing future data and re-running the feature pipeline?


# Is there existing work on detecting look-ahead bias by perturbing future data and re-running the feature pipeline?












Standard advice for avoiding look-ahead bias in time series ML is procedural — use TimeSeriesSplit, apply .shift(1), don't fit scalers on the full series. Tooling that checks this automatically appears to be static: Yang et al. (ASE 2022) analyze source code for preprocessing, overlap, and multi-test leakage, and the IDE plugins built on their work cover the same three categories.

Static analysis seems structurally unable to catch look-ahead in feature construction. Whether df.rolling(60).mean() leaks depends on how the output is aligned to the target, which is often determined elsewhere in the codebase. The syntax alone doesn't settle it.

An alternative would be an empirical test:

Run the feature-construction function on unmodified data and store the output as a baseline Choose a cut point t Corrupt all rows strictly after t (permute, or overwrite with noise), leaving rows ≤ t byte-identical Re-run the same function Compare outputs at rows ≤ t

Since only post-t data changed, any difference in a feature value at or before t implies that feature depends on future information. This requires the pipeline to be deterministic, which can be verified beforehand by running the baseline twice.

I encountered the underlying problem in my own work on CME futures: a missing .shift(1) on a rolling window produced an AUC of 0.957 that fell to 0.675 once corrected.

Questions:

Is there published work or existing tooling that uses this perturbation approach for feature-level look-ahead detection? Everything I have found is static analysis of source code. Are there known failure modes? I can construct cases where permutation preserves the aggregate statistic a feature depends on, so the corruption strategy clearly matters.

I am aware of Gao et al. (2025) on Lookahead Propensity, but that addresses memorization in LLM forecasts rather than leakage in a constructed feature pipeline, so it seems a different problem.

## Answer by Adol (score 0)

https://quant.stackexchange.com/a/85741

The test is sound and I have not seen it packaged, so I will answer the part I can speak to from practice, which is what it catches and what it structurally cannot.

It catches temporal leakage inside your feature construction. That is real and it is the common case, same for your AUC 0.957 to 0.675 example.

The class it cannot catch is leakage from the input data being restated. Your procedure holds rows at or before t byte-identical and perturbs everything after. That tests your function. It assumes the rows themselves are what you would have had at time t. If your source revises history, or re-timestamps it, the pipeline stays deterministic, the perturbation test passes, and the backtest still moves. The corruption happened upstream of your cut point, so nothing downstream of it can see it.

This is not hypothetical in news and alternative data. In one source I ingest, document timestamps arrive stale relative to when the document was actually retrievable. A point in time join against those timestamps looks clean by construction, because the field you are joining on is the field that is wrong. No amount of perturbing rows after t surfaces that.

What addresses that class is provenance rather than testing, and it is three separate things:

- Commit a hash of each period's data at the moment you claim to have had it, and publish the commitments somewhere append only. Then "the history has not been rewritten" becomes a checkable statement instead of an assumption. We do this per hourly bucket, and the reason is not marketing. It is that we could not otherwise prove it to ourselves six months later.

- Label reconstructed or backfilled rows distinctly from live collected ones, and keep the boundary queryable, so you can re-run any result on live only data. We found the two populations do not behave the same, which is a result we would have missed entirely if they had been stored in one undifferentiated table.

- Never let a backfill overwrite a live collected window. If it can, the hash log is the only thing that would ever tell you it happened.

On your determinism precondition: it is doing more work than the question gives it credit for, and running the baseline twice in one process is the weak version. Check it across process restarts and across library versions too. Dict and set iteration order, thread pool scheduling, and non associative float reduction in grouped aggregations will all give you a pipeline that is deterministic within a run and not across them. If determinism is only intra-process, a diff at rows less than or equal to t is ambiguous between leakage and noise, which is the one outcome your test cannot afford.

For literature, I do not have a specific paper for the perturbation formulation as you describe it and I would rather say so than guess. The framing I would search under is metamorphic testing and mutation testing applied to data pipelines, since your construction is a metamorphic relation. That vocabulary sits outside the finance literature, which may be why the leakage work you found is all static analysis.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.