Skip to content
All library documents

Resampling Methods for Estimating Prediction Error from Limited Data

Article MQL5 articles

Summary

The article explains how resampling can estimate predictive model performance when a separate validation dataset is unavailable. It defines apparent error on the training data, prediction error for a fitted model on new observations, population error across possible training sets, and excess error as the gap between prediction and apparent error. These distinctions clarify why training performance tends to understate generalization error.

It presents cross-validation as a way to reuse observations: repeatedly hold out one or more observations, train on the remainder, and average the validation errors. The article notes that cross-validation is broadly applicable and often nearly unbiased, but its estimate can have substantial variance and that variability may be underestimated. It introduces other resampling methods as alternatives intended to estimate error and its variability, while acknowledging their added computational cost and complexity. The techniques can make fuller use of scarce data, but do not remove uncertainty about performance on the broader population.

Key ideas

  • Apparent error measures fit on training data and is typically optimistic about future performance.
  • Prediction error concerns a fitted model’s expected error on new observations, while population error averages across possible training sets.
  • Cross-validation estimates future error by repeatedly withholding observations and averaging their validation errors.
  • Cross-validation can be nearly unbiased yet highly variable across samples.
  • Resampling can reduce the need for a separate validation dataset but may require more computation and complexity.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.