Avoiding Data Leakage in Factor Model Evaluation
Summary
This competition FAQ explains how submitted equity factors are reviewed and how model training should be separated from evaluation. It warns that training on the test dataset, or fitting a model on a portion of data and then using that same portion to produce factor values, can make apparent performance fail to demonstrate generalization. It describes two permitted approaches: train on a designated historical dataset, or use a rolling process that trains only on data preceding the prediction date. The organizers state that submissions are evaluated on a later, separate time period and ranked using factor-analysis measures such as information coefficient.
The post also lists operational reasons a submission may fail, including missing dependencies, an incorrect output type, or factor values that are entirely missing. It gives the training and evaluation periods used in the competition and says a test dataset is available for checking work. These instructions are tied to that platform and contest, and the text does not provide factor results or a full account of scoring. Its main research lesson is to preserve temporal separation between training and prediction and verify that factor outputs are valid.
Key ideas
- Training data must be separated in time from the data used to generate evaluated predictions.
- Rolling models should train on observations preceding each prediction date.
- A later test period is used for factor analysis and ranking, including information coefficient measures.
- Submissions can fail if dependencies are missing, outputs have the wrong type, or factor values are all missing.
- The evaluation rules and dates described apply to the specific competition.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.