Skip to content
All library documents

Sequential Bootstrapping for Less Redundant Financial ML Samples

Article MQL5 articles

Summary

The article explains why ordinary bootstrap sampling can be a poor fit for financial machine learning when labels cover overlapping time intervals. Repeatedly sampling with replacement may select observations that encode much of the same market event, reducing the sample’s effective independent information. It introduces sequential bootstrapping, which adjusts each candidate observation’s draw probability according to its average uniqueness relative to observations already selected.

The article illustrates the calculation with an indicator matrix and a worked example, then outlines how to build the matrix, calculate uniqueness, and use sequential sampling in a broader financial ML workflow. It also describes Monte Carlo comparisons and application to bagging models. The stated rationale is that sampling less redundant labels may improve training and inference, but the supplied excerpt does not include the simulation or trading results needed to assess the size or reliability of any improvement. The method also depends on correctly defining label intervals and overlap; it does not by itself establish that a model’s signals are predictive or profitable.

Key ideas

  • Standard bootstrap sampling can repeatedly select labels that overlap in time and carry redundant information.
  • Sequential bootstrapping recalculates draw probabilities as the sample grows, favoring observations with greater average uniqueness.
  • An indicator matrix represents the time intervals covered by each label and supports the overlap calculation.
  • The article presents sequential sampling as an input-sampling method for financial ML and bagging workflows.
  • The excerpt explains the method but does not provide enough reported results to verify trading performance gains.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.