Troubleshooting Missing Factor Results and Machine-Learning Leakage
Summary
This guide lists reasons a submitted quantitative factor may fail to produce an evaluation score. It recommends checking factor-analysis output on the designated dataset before submission and verifying that a machine-learning model completed training and prediction. Missing values in training data may make some models fail, while excessive cross-sectional missingness in a daily factor can trigger an evaluation error. A forward-fill approach is shown for carrying each instrument’s last available factor value forward.
The article also warns against future information and overlapping training and prediction data. It recommends separating the periods, using rolling training, or training on a designated historical dataset before predicting on the evaluation set. Example workflows illustrate date windows and returning date, instrument, and factor fields. These are platform-specific troubleshooting suggestions, not evidence that a factor has predictive value. Forward filling can also preserve stale observations, and the guide does not discuss model validation, leakage controls beyond time separation, or the suitability of its example windows for every strategy.
Key ideas
- Check factor-analysis output on the target dataset before submitting a factor.
- Missing training values may prevent some machine-learning models from training or predicting successfully.
- High cross-sectional missingness can cause the evaluation framework to reject a factor.
- Forward filling can replace missing observations with each instrument’s last available value.
- Keep training and prediction periods separate and consider rolling training to reduce leakage risk.
- The guide addresses evaluation failures, not whether a factor will be profitable.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.