Avoiding Overfit and False Narratives in Quantitative Factor Research
Summary
The article examines how large datasets and complex models can make factor discovery more productive while increasing the risk of overfitting. It contrasts hypothesis-led research, where an economic or behavioral rationale is tested, with data-led mining, where models identify patterns first. A statistically strong in-sample factor can be a coincidence, and researchers may invent a persuasive explanation after seeing the result, mistaking association for a durable cause.
It also warns that complex models can be difficult to interpret when their performance deteriorates, making it harder to diagnose failure or anticipate changing conditions. Its proposed research discipline is to set plausible theoretical constraints, favor simpler models when their out-of-sample behavior supports them, and investigate why existing factors fail as carefully as researchers search for new ones. The piece is a conceptual argument rather than an empirical comparison: it supplies no dataset, validation results, or specific procedure for measuring overfit, so its recommendations require operational choices by the researcher.
Key ideas
- A strong in-sample factor can reflect noise, especially when many datasets and model variations are explored.
- Researchers risk inventing an explanation after discovering a statistical pattern and treating correlation as causation.
- Complex models may reduce interpretability and make factor failures harder to diagnose.
- Use economic reasoning to constrain searches, evaluate out-of-sample behavior, and favor simplicity where performance supports it.
- Studying why established factors fail can improve research judgment as much as mining for new signals.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.