Use Out-of-Sample Data to Validate Regression-Based Trading Models
Summary
The document addresses whether historical data used to estimate regression models can also serve as the backtest. Reusing the same observations creates data-mining bias: the models may appear strong because they were selected to fit patterns in that sample, rather than because those patterns persist. As a result, in-sample performance alone offers weak evidence about future results.
The proposed procedure splits the history into development and evaluation periods. Fit and rank candidate models using the first portion, then test those fixed models on a later portion that was not used to choose them. If the relative ranking is roughly consistent across the two periods, confidence in the models for future use increases. The example uses a ten-year history divided into two five-year blocks, but this is illustrative rather than a universal split rule. The note gives no specific regression design, performance metric, or guidance on repeated model selection, and one held-out period cannot ensure future performance.
Key ideas
- Evaluating a regression model on the data used to select it introduces data-mining bias.
- Reserve a later segment of historical observations for testing models chosen on earlier data.
- Compare candidate-model rankings across the development and held-out periods as one check on robustness.
- A stable ranking out of sample can increase confidence, but it does not guarantee future success.
- The appropriate split and evaluation metric depend on the modeling context.
Tags
Full text
# What data should be used for regression-based model backtesting? # What data should be used for regression-based model backtesting? I ran regressions using historical valuation data and now want to backtest the models I came up with. Are there any issues with using the same historical data set for the backtest that I need to be aware of? ## Answer by GNUser (score 3) https://quant.stackexchange.com/a/14783 Yes, this is an issue. There will be datamining bias. The best practice is to hold enough of your data out-of-sample to test your models. For example, if you have 10 years of data, use the first 5 years to come up with your models. You could then rank them from best to worst, based on whatever metric you prefer. Then use the second 5 years of data to test your models. If they have approximately the same ranking, then you should have increased confidence in using those models for future data.
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.