Applying Filters Consistently to Training and Prediction Data
Summary
This discussion explains when stock universe filters should apply to both model training data and prediction data, and when they might be applied only at trading time. Its general recommendation is to use the same data processing in both sets. Filters that express stable universe rules or risk constraints—such as excluding certain exchanges, weak fundamentals, or stocks locked at limit up—are presented as candidates for consistent application. Prediction-only filtering is framed as a deliberate trading policy when a model is meant to learn broader market patterns but live trading has additional constraints.
The response defines overfitting as strong training performance that fails to generalize to test or prediction data, and asserts that differing sample distributions do not cause overfitting. That claim is too broad: distribution shifts can undermine out-of-sample performance, and choices made after inspecting prediction results can bias evaluation. The note gives conceptual guidance rather than an empirical comparison. In practice, filters should reflect the intended deployment universe, and model evaluation should use data that preserves a realistic time sequence and avoids leakage.
Key ideas
- The discussion generally recommends applying equivalent data processing to training and prediction samples.
- Stable universe rules and risk constraints may justify filtering both sets.
- Prediction-only filters can represent additional live trading constraints beyond the model’s learning universe.
- A difference in sample distributions can still harm out-of-sample performance, despite the document’s broad assertion.
- The response offers guidance but no empirical test comparing filtering choices.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.