Recognizing and Reducing Overfitting in Predictive Models
Summary
This introductory article explains overfitting as a model learning quirks in its training data instead of patterns that generalize. It contrasts strong training performance with weaker results on unseen observations, using a résumé screening example, and frames the issue as fitting noise rather than signal. It also relates overfitting to underfitting and the bias-variance trade-off: overly simple models can miss structure, while highly flexible ones can become sensitive to noise.
To detect the problem, the article recommends comparing performance on training and held-out test data and using a simple model as a baseline. Suggested controls include K-fold cross-validation for tuning while preserving a final test set, collecting more relevant data, removing unhelpful features, stopping iterative training when validation performance declines, and applying model-specific regularization. It distinguishes bagging, which combines complex models to smooth predictions, from boosting, which combines simpler learners that focus on prior errors. The discussion is a broad overview rather than a trading-specific guide; it does not compare methods empirically or address time-ordered financial data, leakage, or regime changes.
Key ideas
- Overfitting occurs when a model learns training-set noise and performs poorly on unseen data.
- Comparing training and test performance can reveal a generalization gap.
- Cross-validation can support model tuning while a separate test set remains untouched.
- Relevant data, feature selection, early stopping, and regularization can help control overfitting.
- Bagging smooths predictions from complex base models, while boosting combines simpler learners focused on earlier errors.
Tags
Cited by
- Strategies Earnings-Gap Continuation (PEAD Proxy), Long-Short Event-Driven Basket on 20 Liquid US Large Caps (USEQ 1-DAY — enter at the CLOSE of a >=3-sigma high-volume overnight gap, ride the post-announcement drift for ~15 sessions, 3-parameter)
- Hypotheses Earnings-Gap Continuation (PEAD Proxy), Long-Short Event-Driven Basket on 20 Liquid US Large Caps (USEQ 1-DAY — enter at the CLOSE of a >=3-sigma high-volume overnight gap, ride the post-announcement drift for ~15 sessions, 3-parameter)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.