Skip to content
All library documents

Overfitting, Model Validation, and Choosing Simpler Models

Article BigQuant

Summary

This article distinguishes overfitting from overtraining and explains why cross-validation does not, by itself, prevent a model from overfitting. Validation estimates predictive performance, including through out-of-sample data and averaged errors across folds. Overfitting is instead framed as a relative model-selection concern: when models fit comparably, the more complex one may be unnecessarily elaborate. Complexity can involve more parameters or greater nonlinearity, though the article acknowledges there is no universally definitive measure.

A constructed example uses noisy observations from a sine function and compares polynomial fits of different degrees. Their reported cross-validation errors are similar, illustrating that predictive error alone cannot identify which model is overfit; the simpler fit is preferred under the stated simplicity principle. The example is illustrative rather than a trading study, and its complexity assumptions are informal. It offers conceptual guidance, not a complete procedure for measuring complexity or selecting models in every setting.

Key ideas

  • Cross-validation estimates predictive error but does not guarantee that a model is free of overfitting.
  • Overfitting is assessed by comparing models with similar fit and different complexity.
  • Model complexity can reflect parameter count and functional form, though neither gives a universal measure.
  • The article’s noisy sine example shows that similar validation errors can coexist with different model complexity.
  • When fit is comparable, the article favors the simpler model under an Occam-style principle.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.