Skip to content
All library documents

Sequential Bootstrapping for Financial Machine Learning Ensembles

Article Hudson & Thames

Summary

The article explains why ordinary bagging can be problematic for financial labels. In event-based datasets, labels may share underlying returns, so observations are not independent. It introduces concurrency to describe overlapping information and uniqueness as a measure of how much a label is distinct from others. An indicator matrix records which return periods contributed to each label, providing the basis for calculating uniqueness.

Sequential bootstrapping samples labels one at a time to favor selections that increase the average uniqueness of a training sample, while still allowing overlap and repetition. The article describes using this sampling procedure with bagging classifiers and regressors, including ensemble tree methods. It reports implementation timing comparisons and says repeated experiments found higher average uniqueness than standard random sampling, but the excerpt does not give the histogram results or a full out-of-sample trading evaluation. The method aims to make training samples closer to independent and reduce inflated out-of-bag scores; it does not eliminate dependence or guarantee better predictive performance.

Key ideas

  • Financial event labels can overlap in the returns used to construct them, violating the independence assumption behind standard sampling.
  • Concurrency measures shared information among labels, while uniqueness reflects how distinct a label is.
  • An indicator matrix maps each label to the return periods used in its construction.
  • Sequential bootstrapping selects samples to raise average label uniqueness while retaining the possibility of overlap and repetition.
  • The method can be integrated with bagging models, but it does not guarantee improved out-of-sample performance.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.