Correcting Sequential Bootstrap in Cross-Validated Model Calibration
Summary
The article diagnoses why sequential-bootstrap and standard-bootstrap models produced nearly identical out-of-fold Brier scores. It traces the problem to two linked implementation defects: a pipeline helper replaced the sequential bagging classifier with a standard bagging shell, so calibration refits lost the sequential sampling behavior; retaining the original classifier alone then exposed a mismatch between the full event metadata and each smaller training fold.
The proposed correction preserves the classifier’s identity and has cross-validation locate it, slice its event information to each fold, and inject that data before refitting. The article also recommends full-size draws for the sequential sampler because its uniqueness-aware selection already addresses overlap, while the standard sampler continues to use a reduced sample fraction. It describes a comparison module for inspecting sampling and predictive behavior and reports the initial near-identical scores as evidence of the defect. These changes address the described pipeline; the article’s results do not by themselves establish general predictive gains across datasets.
Key ideas
- Replacing the sequential bagging classifier with a standard classifier shell erased its sampling behavior during calibration.
- Keeping the original classifier exposed a mismatch between full-length event metadata and fold-sized training data.
- Cross-validation must pass fold-sliced event metadata to each sequential-bootstrap refit.
- The article recommends full-size draws for sequential sampling and retains reduced sampling for the standard bootstrap arm.
- The reported Brier-score similarity is used to diagnose an implementation problem, not to prove broad model performance.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.