Skip to content
All library documents

Bias–Variance Tradeoff and Model Selection for Regression

Article QuantStart

Summary

The document explains how model flexibility affects prediction error in supervised regression and why the lowest training error does not necessarily identify the best model. It distinguishes training mean squared error from test error, which measures performance on unseen observations and is the more relevant target for generalization. A simulated example with noisy observations from a sinusoidal relationship compares a linear fit with polynomial fits of differing flexibility. Training error falls as flexibility rises, while test error can first improve and then worsen as overfitting increases.

The article frames expected squared prediction error as the sum of irreducible noise, squared bias, and variance. Simpler models may have greater bias, while more flexible models can become more sensitive to the particular training sample. Model choice therefore requires balancing these effects rather than minimizing training error alone. Cross-validation is proposed as a way to estimate test performance, but the discussion focuses on regression and does not provide a complete validation workflow or trading results. Its relevance to quantitative research is conceptual: forecasts on financial data also need out-of-sample evaluation.

Key ideas

  • Training error can decline as model flexibility increases even when unseen-data performance deteriorates.
  • Expected squared prediction error comprises irreducible noise, squared bias, and variance.
  • Simple models may underfit, while highly flexible models may overfit training observations.
  • Cross-validation can help estimate generalization performance when test data are limited.
  • Model selection should target out-of-sample error rather than the smallest training error.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.