Assessing Mutation Adequacy in Backtest Leakage Tests
Summary
The document frames look-ahead leakage checks as mutation testing. At a historical cutoff, a valid mutation preserves information available at that time while changing only future data; if the historical artifact changes, this suggests the pipeline used unavailable information. The question proposes measuring adequacy through planted-fault mutation scores, coverage across leakage classes, multiple adversaries, and confidence bounds on the probability of detecting planted faults.
The response recommends targeted perturbations over random noise, arguing that random changes may fail to activate a hidden dependency. It describes the conventional killed-mutant score adjusted for equivalent mutants and suggests threshold-crossing changes such as sign flips or clamping. However, it overstates that targeted adversaries can guarantee detection whenever a dependency exists: activation depends on the fault, data, and artifact being checked. It offers no formal statistical procedure or treatment for equivalent mutants beyond bypassing them, so its suggestions are a starting point rather than a complete adequacy standard.
Key ideas
- A future-data mutation should preserve all information legitimately available at the historical cutoff.
- A changed historical artifact after such a mutation is evidence of temporal leakage.
- Random perturbations may fail to activate some hidden dependencies.
- Targeted adversaries can improve detection when designed for a known leakage mechanism.
- Mutation scores need a defensible account of equivalent or non-activating mutants.
Tags
Full text
# How can mutation adequacy be quantified for look-ahead leakage tests in backtests?
# How can mutation adequacy be quantified for look-ahead leakage tests in backtests?
Suppose a backtest pipeline is evaluated at a historical cutoff $T$. Let $F_T$ be the information legitimately available by $T$. I am testing temporal causality by applying a mutation $M_T$ that leaves $F_T$ unchanged but perturbs only information that should still be unavailable, then rerunning the pipeline and checking whether the historical artifact $A_T$ changes.
Formally, the test is:
$$M_T(D)\,|\,{F_T} = D\,|\,{F_T} \\ A_T(\,M_T(D)\,) = A_T(D)$$
A violation is evidence that some supposedly historical state depends on unavailable information.
The problem is mutation adequacy. In a controlled fault-injection benchmark I planted six leakage classes: centered rolling windows, full-sample normalization, global feature selection, forward labels that had not matured by the training cutoff, completed bars treated as available at their open timestamp, and same-open retroactive execution.
For the first five classes, targeted mutations exposed all 3,000 planted faults, while causal controls had 0 violations in 3,600 trials. But for the execution fault, a random perturbation of the unavailable close exposed only 286/600 planted faults. A targeted sign-flip adversary exposed 600/600.
So a statement such as 'the pipeline passed a future-mutation test' seems weak unless the mutation family itself is characterized. A weak mutation may simply fail to excite the hidden dependency.
What is a defensible way to quantify mutation adequacy for this kind of look-ahead-bias test?
My current candidates are:
- a mutation score over deliberately planted temporal faults;
- operator coverage by leakage class (normalization, selection, labels, bars, execution, etc.);
- multiple independent or targeted adversaries per cutoff; and
- confidence bounds on the planted-mutant kill probability.
How should equivalent or non-activating mutants be treated in this setting, and is there a more standard statistical or software-testing formulation that maps cleanly to backtest leakage detection?
## Answer by Russlan Ramdowar (score 0)
https://quant.stackexchange.com/a/85795
This is a really solid question. I think it's great you're actually trying to break your backtest systematically instead of just trusting a high Sharpe ratio and getting burned in live trading.
Thinking through this, you’ve basically hit a classic mutation testing problem: random noise is just a weak adversary. I think your sign-flip example proves this perfectly. If you just add random noise, the test might pass even if there's a look-ahead leak. It’s kind of like testing a fire alarm by lighting a single match; if it doesn't go off, you haven't actually proven the system works against a real fire.You need a meaningful trigger.
Looking at your proposed candidates, a combination of #1 and #3 is prob the most practical way forward. In standard software testing, you'd calculate a Mutation Adequacy Score using $Killed / (Total - Equivalent)$. But since financial data is highly structured, random perturbations are going to generate way too many equivalent mutants. For instance, if a leaked indicator triggers when volume > 1M, and your random mutation just shifts the future volume from 1.5M to 1.6M, the leak is physically there, but the mutant fails to activate it.
Honestly, imo, the easiest way to handle these equivalent mutants is just to bypass them. It's probably worth trying to ditch random perturbations entirely and exclusively use extreme, targeted adversaries—like your sign flips. That guarantees that if a dependency exists, it will break the logic, and then you can just confidently track your kill rate.
Btw, it might be worth looking into non-linear thresholding for your mutations—similar to how ReLU activations (you tagged machine-learning so I guess you are familiar with the concept) act as hard cutoffs in neural networks. Clamping future values (like zeroing out volume or capping price changes) instead of just adding noise could be a really clean way to guarantee you force a state change and avoid equivalent mutantsShown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.