Skip to content
All library documents

Time-Based Machine Learning Backtests, Rolling Training, and Live Prediction

Article BigQuant

Summary

The article answers practical questions about machine-learning stock strategies. It says the platform separates training and test periods by time, with the test interval serving as the backtest, and describes a setup without a separate validation set. It also explains that rolling training is not guaranteed to improve results: minimum training history and model update frequency affect sample size, while short histories can lead to underfitting and different market regimes can change model performance.

For live prediction, the article says a model can generate forecasts from newly available data, such as information after the market close, which may then inform trading signals. It discusses two choices after a successful test period: keep using the tested model or retrain it with the intervening data. More recent data may better reflect current conditions, while the older model has a period of backtest evidence. These are conceptual explanations, not independent validation; the claims depend on correct data timing and leakage controls, and no strategy performance statistics are supplied.

Key ideas

  • Training and test data should be separated in time to reduce leakage in backtests.
  • Rolling training performance depends on training-window length, update frequency, and market regime.
  • A short rolling training sample may leave a model underfit.
  • Newly available post-close data can be passed to a model to generate signals for future decisions.
  • Retraining with recent data may improve relevance, while retaining a tested model preserves its prior backtest record.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.