Skip to content
All library documents

Diagnosing XGBoost Replay Differences and Simulation Data Gaps

Article BigQuant

Summary

This Chinese-language discussion describes problems encountered when reusing a saved XGBoost model in a separate strategy and moving a feature-based strategy from research into simulated trading. Keeping visible strategy settings unchanged did not reproduce similar results, and a feature requiring historical data produced no usable rows after missing-value removal in simulation. The post raises model persistence and access to saved user-space models as unresolved platform questions.

The author reports a practical diagnosis for the missing-data issue: simulation had not passed the stock-list date range into automatic labeling and feature extraction, causing an incompatible cache to be used. Matching those date ranges and disabling caches resolved that particular error. The post does not establish why the separate run produced different model results, nor does it provide a general method for loading saved models into simulation. It is therefore most useful as a troubleshooting example about keeping data windows and cache settings aligned across research and deployment, rather than as a complete guide to model reproducibility or deployment.

Key ideas

  • A saved model may produce different results in another strategy run even when visible settings appear unchanged.
  • Simulation can yield no rows after missing-value removal if feature extraction lacks the intended date range.
  • The reported fix was to align the date ranges used for the stock list, labeling, and feature extraction.
  • Disabling caches helped resolve the reported data mismatch.
  • The discussion leaves saved-model access paths and differences between model runs unanswered.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.