Skip to content
All library documents

Why Stable Backtest Parameter Regions May Not Generalize

Article Quant Q&A · Author: vonjd

Summary

The document considers how to assess robustness when a trading strategy has several tunable parameters. It gives moving-average crossover lengths as an example and describes plotting backtest returns across two parameter dimensions as a heatmap. Broad areas of positive performance may merit further investigation, while scattered, irregular results can suggest sensitivity to parameter choices and possible data-snooping risk. The original question asks how to make this analysis rigorous when there are too many dimensions to visualize directly.

The response challenges the assumption that a stable parameter region is necessarily desirable. A constructed example shows that a signal can appear to have a favorable parameter setting under brute-force search and smoothing while still failing to generalize. The key lesson is to test the hypothesis behind any chosen stability measure. The document does not propose a specific multidimensional method or provide empirical trading results, so it offers a caution about interpretation rather than a complete robustness procedure.

Key ideas

  • A heatmap can reveal how backtest results vary across two strategy parameters.
  • Broad positive regions may warrant further study, while irregular performance can indicate parameter sensitivity.
  • A stable-looking parameter region does not by itself establish that a strategy will generalize.
  • Any proposed stability measure should be evaluated against the hypothesis it is intended to test.
  • The response offers a caution rather than a specific multidimensional robustness method.

Tags

Full text
# Finding robust regions of multidimensional parameter combinations in trading strategies


# Finding robust regions of multidimensional parameter combinations in trading strategies












Trading strategies often have many degrees of freedom. As a toy example let's say you have two moving averages (MA) which trigger a trade each time they cross each other: There are at least two parameters which could be optimized, namely the length of each MA.

Now, you would not want to bet your money on a strategy that just by chance found that you have to take 57 days for one and 243 days for the other MA (see data snooping bias). What you want to see is that the strategy as such is sound per se, i.e. robust und not dependant on the exact parameter settings.

One way to go for two parameters is to plot a heatmap with the two parameters as axes and a colour coding for the respective return of the strategy in a backtest. If you only see noise and many convoluted regions of different return levels this is a good sign that this is not a robust strategy. If there are bigger regions of positive returns these regions merit further investigation.

My question What are established methods to find robust regions of multidimensional parameter combinations in trading strategies? The challenge here: You obviously cannot visualize more than three degrees of freedom and you have to make these ideas mathematically rigorous.

Multivariate kernel density estimation comes to mind but this is just my first idea. I am thankful for every lead, reference and code example (preferably in R).

## Answer by madilyn (score 1, accepted)

https://quant.stackexchange.com/a/32892

Not exactly the answer you're looking for: It's not obvious that a region of stability is a desirable property.

One can trivially construct an example where this is true: suppose the actual generation function of your target is $f: x_t \mapsto 2x_t+1$ over $\mathbb{R}$, and you have a signal $s$ of one parameter $s\left(p,x_t\right)=\left(p^{e} \mod 3\right) \cdot \left(2x_t+1\right)$ where $s$ is defined over the domain of $\mathbb{R}^{+} \times \mathbb{R}$. If you try to learn some $\hat{f}\left(p,\cdot \right)=s\left(p,\cdot \right)\approx f(\cdot)$ by brute force grid search of $p$ that minimizes the RMSE followed by some smoothing transformation of your grid, you will easily end up with a model that doesn't generalize well, even though a naive optimizer might converge to $p=1$ easily.

So whatever fanciful stability function that you choose, I would test your hypothesis first.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.