Skip to content
All library documents

Sample Uniqueness, Sample Weights, and Sequential Bootstrapping

Article Quant Q&A · Author: Tsz Chun Leung

Summary

This explanation distinguishes sample uniqueness from sample weights in financial machine learning. A model’s sample weights tell it how strongly to emphasize observations during fitting. Uniqueness measures how much an observation overlaps with others, and can be used when calculating weights; the answer gives return-based weighting as an example, dividing the absolute return by sample uniqueness.

The answer also describes sequential bootstrapping for bagging ensembles. Instead of selecting training observations at random, the method favors samples that add uniqueness to the set already selected. This changes how ensemble estimators receive their training data, whereas sample weights affect the emphasis given to samples during fitting. The discussion is conceptual and points to an implementation and further reading, but supplies no comparative performance evidence or details for evaluating when the method improves results.

Key ideas

  • Sample weights control how much training observations influence model fitting.
  • Sample uniqueness measures overlap among observations and can inform their weights.
  • Return-based sample weights can scale absolute returns by sample uniqueness.
  • Sequential bootstrapping selects samples to increase uniqueness in bagging ensembles.
  • Weighting observations and changing the sampling procedure are distinct techniques.

Tags

Full text
# Sample uniqueness and sample weight in AFML book


# Sample uniqueness and sample weight in AFML book












With reference to AFML ("Advances in Financial Machine Learning" book by Marcos Lopez de Prado). Are sample uniqueness and sample weight pointing towards to the same thing? I am confused on the term here. Thanks if anyone could help.

## Answer by Alexandr  Proskurin (score 5)

https://quant.stackexchange.com/a/50950

That is a very good question.

If you look at sklearn fit() method parameters, you can find sample_weights parameter which tells the model which samples it should give more attention/weight when the model is fit.

Sample uniqueness is a bit different. Firstly, sample uniqueness is used to calculate sample weights (for return based sample weights, we divide abs(return of sample) / sample uniqueness). However, the rest of Sample Weights chapter tells the reader how to improve bagging algorithm (used in ensemble models like BaggingClassifier, RandomForestClassifier) is such a way that instead of randomly choosing samples used for training estimators in ensemble model, the algorithm chooses the most unique one. This is called Sequential Bootstrapping. So this algorithm (if implemented) changes how the model is fit.

NB: I am the contributor of an open-source package mlfinlab (https://github.com/hudson-and-thames/mlfinlab) which implements the concepts described in AFML book. We have an ensemble model SequentiallyBootstrappedBaggingClassifier/Regressor which extend sklearn's BaggingClassifier model with Sequential Bootstrapping instead of random sampling.

Our research team also has a blog where we discuss various financial machine learning problems and how mlfinlab can be used to solve them. Regarding your question, there is the blog post on sample uniqueness and Sequential Bootstrapping (https://hudsonthames.org/bagging-in-financial-machine-learning-sequential-bootstrapping-python/)

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.