Correcting Triple-Barrier Label Overlap with Sample Uniqueness Weights
Summary
The article explains how triple-barrier labels can overlap in time, making financial machine learning observations dependent rather than independent. When many samples encode information from the same market interval, a model can effectively learn repeated patterns, inflating in-sample performance and weakening generalization. The proposed correction is to weight each observation by its average uniqueness, reducing the influence of labels that overlap heavily with others while retaining their information.
The method counts active events across bars, then computes each event's average reciprocal concurrency over its lifespan. It describes functions for calculating concurrency, average uniqueness, and associated sample weights, including parallel processing and integration with scikit-learn classifiers. The article frames the method as a remedy for a specific dependence problem; it does not make overlapping labels independent or guarantee improved live results. Robust evaluation still depends on appropriate cross-validation and careful event interval construction.
Key ideas
- Triple-barrier labels can share time intervals, violating the independence assumption used by many standard learning methods.
- Repeated information across concurrent labels can make in-sample results look stronger than out-of-sample performance warrants.
- Average uniqueness estimates an event's distinct information by averaging reciprocal concurrency over its lifespan.
- Sample weights lower the contribution of heavily overlapping observations while preserving them in the dataset.
- Concurrency-aware weighting addresses label overlap but does not guarantee generalization or replace sound validation.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.