Explicit, Implicit, and Tacit Overfitting in Trading Research
Summary
The document distinguishes explicit, implicit, and tacit overfitting in trading research. Explicit overfitting comes from fitting too many parameters to historical data; suggested controls include reducing degrees of freedom, using robust fitting, and testing with expanding or rolling windows that preserve the order of time. Examples include constraining parameter relationships or sharing parameter choices across instruments.
Implicit overfitting can arise through repeated manual choices: trying alternative parameters, rules, fitting settings, or secondary settings and retaining the backtest that looks best. Tacit overfitting comes from prior beliefs and design decisions that shape the model before formal fitting, including choosing momentum because of earlier experience or limiting the candidate rules. The document notes that even broad machine learning models embody assumptions about which inputs might predict returns. It offers conceptual cautions rather than empirical results or a validated procedure; its humorous remarks about hiring uninformed researchers are not a practical method.
Key ideas
- Explicit overfitting can be limited by reducing model flexibility and applying robust fitting methods.
- Backtests should preserve chronology by fitting on past data and evaluating on later periods.
- Repeatedly changing parameters or rules after viewing results creates implicit fitting.
- Prior beliefs about strategies and data inputs can shape a model through tacit fitting.
- More flexible models require greater care because they can encode assumptions and fit noise.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.