Skip to content
All library documents

Measuring and Reducing Overfitting in Quantitative Models

Article BigQuant

Summary

The document compares machine-learning approaches to strategy development with more subjective manual parameter tuning. It recommends measuring overfitting by comparing performance on training and test data, using cross-validation where appropriate. For time-series strategies, it advises preserving time order through chronological splits and rolling training, since random splits can give a misleading assessment. It also identifies limited data and changing market distributions as sources of model fragility.

A further risk arises when researchers repeatedly inspect test results while developing a strategy, allowing future information to influence choices. The text argues that restricting model fitting to training data can reduce this human-driven leakage, though it does not remove it automatically. To control model complexity, it cites limiting the number of trees and leaves and requiring a minimum number of samples per node. These are general recommendations, not a documented experiment: the post provides no dataset, performance comparison, or detailed validation protocol, and its claims about machine learning’s advantage over manual tuning should be treated as context rather than empirical proof.

Key ideas

  • Training and test performance gaps can help quantify overfitting.
  • Time-series validation should respect chronology, with rolling training offered as one approach.
  • Repeatedly consulting test results during strategy development can leak future information into decisions.
  • Market data are limited and their distributions can shift, reducing model generalization.
  • Reducing tree complexity and setting minimum node sample sizes can constrain overfitting risk.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.