Skip to content
All library documents

Walk-Forward Backtesting and Data-Mining Bias in Strategy Selection

Article Quant Q&A · Author: tmakino

Summary

The document reviews a strategy-selection pipeline that fits models on training data, tunes hyperparameters using validation-period simulated profit and loss, selects top-performing instruments or strategies, and evaluates the resulting portfolio on a held-out test period. The question arises because strategies selected on validation data failed to meet expectations in live trading, motivating a final test stage.

The response questions whether a simple training, validation, and test split is sufficient for trading research. It proposes repeatedly retraining on historical data at each day or bar, predicting the next period, and then measuring data-mining bias. Strategies with lower measured bias are suggested as candidates for live trading. This is presented as an alternative methodology and personal view, not as a demonstrated guarantee of live performance. The source gives no precise definition or calculation for the bias measure and does not specify transaction costs, portfolio construction rules, or other implementation details, so those require separate treatment.

Key ideas

  • Selecting strategies by validation-period performance can overfit the selection process.
  • A final test period can reveal failures that validation results conceal.
  • Repeatedly retraining on past observations and predicting the next period is a proposed alternative to a fixed split.
  • The response recommends measuring data-mining bias when evaluating a large set of strategies.
  • Low estimated bias is suggested as a selection criterion, but it does not guarantee live success.

Tags

Full text
# How to properly set strategy parameters and select portfolio


# How to properly set strategy parameters and select portfolio












I have the following strategy pipeline which is a function of several hyperparameters and execution parameters:

```
for each instrument {
    1. Calculate features
    2. Split the data into (training, validation, test) // test not used yet

    for each hyperparameter permutation {
        3. Fit regression model on training, calculate predictions on validation
        4. Use validation predictions to simulate pnl
    }

    5. Choose hyperparameters that maximize f(pnl) // Sharpe, etc
}    

6. Add best performers to portfolio
7. Evaluate portfolio performance on test set
```

Can anyone find anything wrong with this process? In the past, I did a training / validation split without a final test set, and ended the process at step 6. The consequence was that my optimized portfolio did not perform as expected during live testing. In other words, the portfolio selection process didn't generalize, and the top performers during the validation period did not perform well during live testing. By adding step 7, I aim to know about such failures before going into live testing.

## Answer by bernacek (score 1)

https://quant.stackexchange.com/a/37832

Some people claim (and I tend to agree with them) that splitting data into [training, validation, test] subsets may not be a right approach at all when backtesting trading strategies. The alternative way is to backtest by retraining your model at every historical day/bar using past data and taking a prediction for the next day/bar (kind of walk-forward but retraining as often as possible) and after all measure DMB (data mining bias - very important!) instead of performing the out-of-sample tests. Then for live trading you select only strategies having low DMB values. There is an old, but very nice thread at FF which gives more insights into this methodology: https://www.forexfactory.com/showthread.php?p=7927988#post7927988

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.