Backtesting as a Scientific Test of Trading Hypotheses
Summary
The document explains why trading institutions use backtests despite the inability of historical performance to guarantee future results. One response frames backtesting as part of a scientific process: start with an economic, behavioral, technical, or other intuition, formulate a hypothesis, translate it into a strategy, and test it against historical data. A failed test can falsify the idea; a successful one calls for further robustness checks and refinement.
A second response points to persistent market behaviors, including volatility clustering, factor effects, correlations, and order-book patterns, as reasons researchers study historical data for potentially recurring signals. It also notes that firms report past performance because future performance cannot yet be observed. The discussion treats backtesting as imperfect evidence, not a guarantee. It warns against data mining and does not specify particular validation procedures or establish that any cited market behavior will persist.
Key ideas
- Backtesting can test a strategy derived from an explicit market hypothesis.
- A failed backtest can count against the intuition that motivated the strategy.
- A successful historical result needs robustness checks and does not guarantee future performance.
- Historical analysis is motivated by the possibility that some market behaviors persist.
- Searching data indiscriminately for profitable patterns risks mistaking noise for signal.
Tags
Full text
# Why do institutions backtest?
# Why do institutions backtest?
I see that institutions still use backtesting by computing P&Ls over historical data and then compute some aggregating ratios to see whether a trading strategy is good or not even though it is not a rigourous approach at all. I mean, how can a trading strategy that happened to perform well in one sample path be guaranteed to perform as well out of sample ?
## Answer by vonjd (score 6, accepted)
https://quant.stackexchange.com/a/24622
> How can a trading strategy that happened to perform well in one sample path be guaranteed to perform as well out of sample?
I think you are having it backwards - this is how I do it:
- Intuition about some economic, psychological, behavioral, technical etc. phenomenon.
- Trying to make my intuition precise in the form of a hypothesis.
- Trying to translate my hypothesis into a trading strategy.
- Backtesting the trading strategy. In case it does not work: Falsification of my intuition in case it does work: More tests (robustness), and with some refined intuition back to step 1.
Basically this is how the scientific method works when doing research on the stock market. At least this is how it should be, so I somewhat agree with your insinuation that just data mining stock market data to find something is bad science ("Torture the data until they confess" ;-)
So, yes, it is not perfect - but it is the best we have to try to find the signal in the noise (and there is a lot of noise...)
A good starting point to understand more about this approach is this book: Evidence-based technical analysis by David Aronson
It explains the whole process (including the complete statistical background).
See for a short summary of important points here: CXO Advisory
See for a comprehensive review here: Automated trading system
## Answer by madilyn (score 9)
https://quant.stackexchange.com/a/24609
Mostly because of convention and tradition. As Student T mentioned earlier, part of this is that it is common practice. You report to your clients or managers how well something performed in the past; you cannot report to them how well it performed in the future. You may have thought of some useful forward-looking measures, but unfortunately the adoption rate in finance for these things is extremely slow. We still teach CAPM as the forefront in the leading business schools, even though this was introduced in the 1960s. We still credit novelists like Taleb for "discovering" non-Gaussian and black swan behavior in 2000s, even though the authors of the models that he critiques had themselves introduced jump diffusion models in the 1970s.
That said, I think you are making the implicit assumption here that regime shifts prevent you from applying your models out-of-sample:
> How can a trading strategy that happened to perform well in one sample path be guaranteed to perform as well out of sample ?
There are persistent phenomena across all time horizons: The low volatility anomaly, the strong dominance of certain factors in explaining returns, volatility clustering, sector correlations, arbitrage between specific symbols, order book properties. It's the presence of these persistent behaviors that motivate people to attempt to extract insight from past data.Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.