Skip to content
All library documents

Meta-Labeling to Filter False Positives in Trading Signals

Article Quant Q&A · Author: PyRsquared

Summary

The document discusses a two-stage classification approach to finding trading opportunities. A primary model is tuned for high recall, accepting false positives so that it identifies most potential positive cases. A secondary model then attempts to filter those candidates. The proposed workflow adds the primary model’s predictions to the original features and labels each case according to whether the primary prediction matches the true outcome.

The author asks whether this correctness-based label is appropriate and what objective the secondary model should optimize. The document does not answer either question or report empirical results; it presents the workflow as an assumption based on a secondary source and an interpretation of a book’s description. It also leaves open how to train and evaluate the filtering model. Thus, it introduces the purpose of meta-labeling and a possible label construction, while its central methodological choices remain unresolved.

Key ideas

  • A primary classifier can prioritize recall to identify a broad set of candidate opportunities.
  • A secondary model is intended to reject false positives from the primary model’s predictions.
  • The proposed meta-label marks whether the primary prediction matches the observed outcome.
  • The document leaves the secondary model’s optimization metric and label construction open to discussion.

Tags

Full text
# Meta Labeling for trading opportunities


# Meta Labeling for trading opportunities












In Advances in Financial Machine Learning, Lopez explains how we should build a primary exogenous model (binary classifier) to identify trading opportunities and a secondary meta model to filter out the false positives from the exogenous model. The exogenous model should have high recall to identify most trading opportunities at the expense of a low precision (high number of false positives). The idea of the meta model is to "increase your F1-score by filtering out the false positives, where the majority of positives have already been identified by the primary model."

It is clear we should train the primary model such that it maximises recall and / or choose a probability threshold that yields high recall at the expense of low precision. However, Lopez does not explain which metric we should maximise / minimise when training the secondary model to "filter out the false positives".

I assume a scheme outlined by Hudson and Thames:

- Train a primary model on the original data features

- Choose a probability threshold which yields high recall

- Use the primary model to make predictions on the training data with the threshold from step 2.

- Append the predictions from step 3. to the original data features. This is a new feature for the secondary model

- Define the meta labels for the secondary model as follows: If the primary model’s predictions matches the actual values, then we label it as 1, else 0. This part is important in how we define labels for the secondary model

- Train the secondary model using the features and labels constructed from the steps above

I have 2 questions

i) Is the above method in step 5. of generating meta labels correct? Lopez did not make it explicit, but in the book he does say "... we correct for the low precision by applying meta-labeling to the positives predicted by the primary model" which seems to indicate step 5. above is correct

ii) What metric should we maximise / minimise when training the secondary model so that indeed false positives are filtered out? Again, Lopez is not explicit here; he only says that the primary model should have high recall.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.