Evaluating Quantitative Strategies Through Falsifiability, Testing, and Simplicity
Summary
This essay applies ideas from philosophy of science to judging trading strategies. It contrasts testable claims, which make predictions that evidence could disprove, with claims that cannot be meaningfully tested. For strategies whose edge is inferred from historical frequencies rather than derived from a formal theory, it argues that confidence should rise with more observations, longer histories, and additional out-of-sample checks. It recommends building a strategy on an earlier period and assessing it on later data, and varying test start dates when trades or rebalancing are infrequent.
The essay also invokes Occam’s razor: when strategies explain results similarly, prefer the simpler one with fewer conditions. It warns that apparent edges can disappear as market regimes change, and gives examples of trend, moving-average, small-cap, and valuation effects that may have limited or uneven histories. These are methodological arguments and illustrative examples, not a systematic empirical study. The piece does not quantify how much testing is sufficient, and its probability analogies do not address all issues such as selection bias, transaction costs, or changing market structure.
Key ideas
- Treat strategy claims as hypotheses that should make predictions capable of failing empirical tests.
- Historical frequencies become more informative when supported by many observations and later-period validation.
- For infrequently rebalanced strategies, vary the test start date to expose sensitivity to timing.
- Prefer simpler strategies with fewer conditions when competing approaches perform similarly.
- An edge that worked historically may weaken or fail under different market regimes.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.