Mean-Centering Returns in Block Bootstrap Tests
Summary
The document considers whether to subtract the sample mean from log returns before applying a block bootstrap. Its main answer says centering is useful when the bootstrap is meant to represent a null hypothesis of zero expected excess return. Resampling contiguous blocks preserves some serial structure, while subtracting the mean creates a distribution centered on the null for evaluating whether observed strategy profits are statistically significant.
It connects this use to performance evaluation and data-snooping tests, describing how a sample statistic can be compared with a centered bootstrap distribution to obtain a p-value. The answer argues that using the full sample mean for this test does not itself create forward-looking trading bias because it is used for statistical inference, not a live decision. Another response cautions that centering removes information about the original mean, so the appropriate treatment depends on the estimand. The discussion does not provide implementation details or settle every bootstrap design choice.
Key ideas
- Block bootstrapping resamples contiguous blocks to create synthetic return sequences while retaining some dependence structure.
- Centering returns can construct a bootstrap distribution under a zero-mean null hypothesis.
- A centered bootstrap distribution can support significance tests of observed strategy performance.
- Subtracting the mean discards information about the sample mean, so centering should match the question being tested.
Tags
Full text
# Block Bootstrapping Relative Returns
# Block Bootstrapping Relative Returns
I want to run a block bootstrap on the relative returns, but I'm not sure if subtracting the mean is important.
A bootstrap sequence is a synthetic sequence generated using the original sequence. If you're backtesting or estimating certain measures, such as volatility or the mean, then you can use the bootstrap to get a confidence interval on these values. The issue is in the way the bootstrap sequence is generated. You have to cut the original sequence into blocks, for the block bootstrap, and select blocks uniformly at random and place them in the sequence they were selected until you have a new bootstrap sequence of length n. You can't use the original prices, so you have to use the relative prices. These relative prices are the same as 1 day returns. I am wondering if mean centering is important in practice. I currently take the difference of their logs,
$\log\left(\frac{r_t}{r_{t-1}}\right)$
but it seems that some posts have suggested subtracting the mean return from the difference of logs
$\log\left(\frac{r_t}{r_{t-1}}\right) - \mu_r$
Most of the papers have an assumption of zero mean sequences. It's easy to zero-mean, but I'm afraid that this introduces some lookahead bias as the mean of the entire sequence is only known with knowledge of the full sequence. Why is subtracting the mean useful/important in practice?
## Answer by user5399 (score 3, accepted)
https://quant.stackexchange.com/a/8917
It obviously depends on what you're trying to do but since we're speaking about returns zero centering is what's usually done because of the null hypothesis claiming that expected excess returns are zero. You zero center the distribution because you want to obtain a distribution satisfying the null hypothesis. In this distribution you then plug your sample mean and get a p-value.
This comes in handy in performance evaluation. It's what Aronson does in Evidence Based Technical Analysis when measuring the significance of the observed profits. It's also what White does in A Reality Check for Data Snooping when calculating the p-value for each model. White calculates for example those two V-values. For a single model you have
$$\bar V_1 = n^{1/2} \bar f_1$$
which is basically the sample mean (see the paper if you don't get the $n^{1/2}$ value) and you also have the bootstrapped distribution
$$\bar V_{1,i}^* = n^{1/2} (\bar f_{1,i}^* - \bar f_1)$$
which as you can see is zero centered through mean subtraction. The p-value is then obtained by plugging in $\bar V_1$ in the $\bar V_{1,i}^*$ distribution.
I also wouldn't say this adds any forward looking bias as you're not using the mean information to make any kind of decision. You're simply trying to determine whether the observed returns over a certain period beat the expected returns of a random strategy.
## Answer by user7056 (score 0)
https://quant.stackexchange.com/a/8225
When you do bootstrapping, the actual problem is that you have too few events to really get an actual distribution. There are X methods to get the original set enriched, reinserting the events based not only on the uniform distribution when selecting the nth event out of the N total ones.
As the mean is one of the first characteristics of a distribution (the 1st moment), substracting a mean (whatever reasons you have), you implicitly throw information away. You might well calculate whatever mean you want at the very end and do with it the intended stuff.Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.