Skip to content
All library documents

Time-Series Cross-Validation to Reduce Overfitting in Quantitative Models

Article BigQuant

Summary

The report compares ordinary K-fold cross-validation with time-series cross-validation for tuning machine-learning models on sequential data. Random folds can place later observations in training while earlier ones are used for validation, creating leakage that violates the independence assumptions of conventional cross-validation. Time-series folds instead train on earlier periods and validate on subsequent periods, preserving temporal order.

Experiments cover public datasets and a Chinese A-share stock-selection strategy. The report finds that time-series validation often produces weaker in-sample scores but stronger test performance, especially for more complex learners such as XGBoost; results for simpler models and non-time-series data are less distinct. In the stock-selection tests, time-series validation generally improved model metrics and strategy results, though drawdown benefits varied. The authors note limitations including fold construction that could still mix some monthly observations, possible underfitting, dependence on the underlying learner, and the risk that historical relationships may fail under changed market conditions.

Key ideas

  • Random K-fold validation can leak future information into training when applied to time-ordered data.
  • Time-series cross-validation trains on earlier observations and validates on later periods.
  • The reported advantage is clearest for complex learners, while simpler models show smaller differences.
  • Tests on A-share selection found stronger out-of-sample model performance and generally better strategy results with time-ordered validation.
  • The method can underfit and cannot protect a model from future market regime changes.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.