Understanding the Bias–Variance Tradeoff in Trading Models
Summary
This tutorial uses simulated data generated from a second-order relationship with normally distributed noise to explain model underfitting, overfitting, and the bias–variance tradeoff. It divides the observations into training and test sets, fits polynomial regression models of several complexities, and compares their errors. A simple model may miss the underlying pattern and perform poorly on both sets; a highly complex model can fit training observations closely while generalizing poorly to unseen data. Intermediate models can balance these errors.
The article presents mean squared error as a combination of squared bias, variance, and irreducible noise, and uses the simulation to build intuition for those terms. Its evidence is illustrative rather than market-based: it does not test price prediction or a trading strategy in this first part, and the small simulated sample, noise, and model choices affect the observed comparison. The promised application to market prediction is outside the document’s scope.
Key ideas
- A model with excessive simplicity can underfit and have high bias.
- A highly flexible model can overfit training noise and show high variance.
- Training and test errors help assess whether model complexity generalizes beyond fitted data.
- Mean squared error can be decomposed into squared bias, variance, and irreducible error.
- The examples use simulated polynomial data and do not establish trading performance.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.