Skip to content
All library documents

Time-Series Cross-Validation for Reducing Trading Model Overfit

Article BigQuant

Summary

The document explains why conventional random or K-fold cross-validation can mislead when tuning models on time-series data. Because observations are ordered and related, a random split may train on later periods and validate on earlier ones, allowing future information to influence model selection. Time-series cross-validation keeps validation periods later than their training data while comparing model performance across successive time-based folds. The report compares the approaches using public machine-learning datasets and an all-China equities selection dataset. It says time-series validation tends to choose simpler models, may show weaker training performance but better test performance, and has greater impact for complex learners than for simpler models. Its reported strategy tests show higher returns and some drawdown improvement, though the supplied text gives no detailed numerical evidence. The approach can also guide parameter searches for other quantitative strategies. Results depend on the underlying learner and historical patterns; changing market conditions can make those patterns fail, and the method may underfit.

Key ideas

  • Random cross-validation can leak future-period information into model selection for time-series data.
  • Time-series folds preserve chronological order by validating on periods later than the training data.
  • The report finds less overfitting and stronger test performance, especially for complex models.
  • Chronological validation can also be applied to tune parameters in non-machine-learning strategies.
  • The method may underfit and remains vulnerable to changes in market conditions.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.