Testing Sequential Bootstrap for Bagging Correlation
Summary
The article tests whether sequential bootstrapping reduces correlation among bagged trees trained on overlapping triple-barrier labels. It separates two changes often made together: reducing the number of sampled rows to reflect average uniqueness, and changing the sampling rule to favor less-overlapping rows. Four regimes use the same decision-tree base learner and are evaluated with purged cross-validation, avoiding feature subsampling as a confound.
The reported results show that reducing the draw count lowered between-tree prediction correlation, while sequential sampling at that same count added little or did not help. Sequential sampling did increase the uniqueness of sampled rows, but this did not translate into further decorrelation. Across the tested bar representations, out-of-sample AUC remained near chance, and a permutation test did not establish a meaningful performance difference. The findings are limited to the studied EURUSD data and experiment design; the article also describes calibration and out-of-bag diagnostics and proposes further holdout evaluation.
Key ideas
- The study separates bootstrap draw size from the sequential sampling rule to identify their effects.
- All four regimes use the same decision-tree learner to isolate row sampling as the experimental variable.
- Reducing the draw count lowered between-tree correlation more consistently than sequential sampling did.
- Sequential sampling increased draw uniqueness without producing corresponding out-of-sample gains.
- The measured AUC differences were statistically inconclusive and remained close to chance.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.