Skip to content
All library documents

Active Learning Strategies for Selecting High-Value Labels

Article BigQuant

Summary

This overview explains active learning as a way to reduce the cost of building supervised or semi-supervised models when expert labels are scarce. A model repeatedly identifies candidate examples for human review, incorporates the resulting labels through retraining or incremental updates, and then selects another batch. The article describes sequential and pool-based workflows and gives applications such as fraud detection, anomaly detection, and message classification.

It surveys query strategies that prioritize uncertain or informative examples. These include least-confident predictions, small margins between the top class probabilities, predictive entropy, disagreement among a committee of models, expected model change, expected error or variance reduction, and density weighting to avoid overemphasizing isolated outliers. A probability example illustrates how a near-even class prediction can be more useful to label than a confident prediction. The discussion is conceptual and cites survey literature, rather than presenting a trading experiment or comparative results. It does not prescribe one universally best query strategy; the appropriate choice depends on the model, data, and labeling setting.

Key ideas

  • Active learning directs expert labeling toward examples expected to improve a model, then feeds those labels back into training.
  • Uncertainty sampling ranks candidates by low confidence, a narrow gap between leading class probabilities, or high entropy.
  • Query-by-committee uses disagreement among multiple models, including voting entropy or divergence between predictions.
  • Other strategies target expected model change, error reduction, variance reduction, or representative dense regions.
  • The article describes a general machine-learning workflow and does not establish which query method performs best in a trading task.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.